Socio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability Distillation
Organizations: School of Data Science, Fudan University · Shanghai Innovation Institute · Tencent
Abstract
Simulating individual behavior requires large language models (LLMs) to preserve persona traits while adapting to dynamic social contexts. However, general-purpose LLMs often flatten distinct personas, while task-specific tuning suffers from fragmentation and generalization. To overcome these challenges, we organize individual simulation into the \textbf{FONTS Taxonomy}, comprising five complementary capability dimensions: \emph{persona fidelity} (\textbf{F}), \emph{outcome realization} (\textbf{O}), \emph{behavioral naturalness} (\textbf{N}), \emph{trajectory coherence} (\textbf{T}), and \emph{social grounding} (\textbf{S}). Grounded in this taxonomy, we curate a standardized training corpus library of approximately 10 million instances across 14 representative datasets and present \textbf{Socio-Foundation}. Socio-Foundation decouples specialization from integration via a three-stage pipeline: learning task experts via DAPO, consolidating them into capability experts via off-policy distillation, and unifying them via multi-teacher on-policy distillation (MOPD). We also establish \textbf{IndiEval}, consolidating 29 metrics across the FONTS dimensions. Experiments show that Socio-Foundation outperforms its \textit{Qwen3-8B} base by 11.0 points and approaches frontier models such as \textit{GLM-5.2}, with ablations and out-of-distribution evaluations further demonstrating the effectiveness and generalization of our model.
Figures & tables
| F | O | N | T | S | Avg. | |||||||||||||||||||||||||
| Model | LifeChoices Acc. | AlignX Choice acc. | HumanLLM Hit@5 | BehaviorChain Node | CoSER Char. fid. | HUMANUAL Response | HUMANUAL State | UserLM-S3 Intent adh. | SOTOPIA Believ. | SOTOPIA Goal | SOTOPIA Relation. | SOTOPIA Knowledge | SOTOPIA Benefits | AgentSense Goal maj. | UserLM-LiC Intent cov. | UserLM-LiC Task score | UserLM-S3 Term. F1 | MirrorBench GTEval | CoSER Anthrop. | UserLM-S3 Human lik. | UserLM-S3 Diversity | UserLM-S3 Inv. overlap | UserLM-S3 Role adh. | CoSER Story cons. | BehaviorChain CumScore | CoSER Story qual. | FANToM Item acc. | Social-R1 Acc. | AgentSense Private acc. | Average |
| Frontier models | ||||||||||||||||||||||||||||||
| GPT-5.5 | 92.0 | 61.0 | 64.1 | 89.7 | 49.6 | 21.8 | 44.3 | 98.0 | 84.7 | 59.0 | 63.7 | 29.3 | 55.0 | 96.7 | 99.0 | 94.4 | 0.0 | 41.3 | 41.9 | 51.2 | 91.6 | 75.6 | 32.0 | 62.6 | 58.7 | 65.6 | 79.3 | 69.0 | 82.4 | 67.2 |
| Open-weight general-purpose models | ||||||||||||||||||||||||||||||
| DeepSeek-V4-Flash | 81.8 | 59.6 | 62.0 | 69.3 | 43.2 | 15.5 | 39.5 | 96.0 | 82.7 | 54.5 | 60.8 | 27.4 | 51.9 | 95.5 | 93.3 | 57.9 | 34.1 | 48.5 | 44.5 | 29.8 | 86.8 | 87.5 | 5.0 | 62.6 | 26.3 | 75.7 | 61.5 | 51.0 | 82.4 | 59.7 |
| GLM-5.2 | 84.2 | 62.2 | 63.4 | 76.7 | 43.9 | 20.0 | 45.5 | 94.0 | 86.1 | 55.5 | 72.9 | 37.2 | 50.5 | 96.7 | 97.3 | 76.8 | 0.0 | 43.2 | 40.8 | 79.7 | 91.0 | 78.4 | 6.0 | 59.7 | 34.1 | 68.0 | 74.3 | 69.0 | 80.8 | 63.2 |
| F | O | N | T | S | Avg. | |||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Variant | LifeChoices Acc. | AlignX Choice acc. | HumanLLM Hit@5 | BehaviorChain Node | CoSER Char. fid. | HUMANUAL Response | HUMANUAL State | UserLM-S3 Intent adh. | SOTOPIA Believ. | SOTOPIA Goal | SOTOPIA Relation. | SOTOPIA Knowledge | SOTOPIA Benefits | AgentSense Goal maj. | UserLM-LiC Intent cov. | UserLM-LiC Task score | UserLM-S3 Term. F1 | MirrorBench GTEval | CoSER Anthrop. | UserLM-S3 Human lik. | UserLM-S3 Diversity | UserLM-S3 Inv. overlap | UserLM-S3 Role adh. | CoSER Story cons. | BehaviorChain CumScore | CoSER Story qual. | FANToM Item acc. | Social-R1 Acc. | AgentSense Private acc. | Average |
| Task-expert SFT | 72.8 | 58.6 | 58.6 | 83.9 | 26.3 | 14.5 | 41.2 | 100.0 | 77.6 | 45.3 | 55.6 | 28.3 | 48.7 | 91.2 | 70.0 | 31.6 | 45.3 | 53.9 | 30.1 | 34.0 | 80.5 | 97.9 | 100.0 | 53.4 | 44.2 | 48.1 | 78.2 | 61.0 | 79.6 | 59.5 |
| Task-expert OPD | 70.8 | 52.6 | 56.6 | 77.4 | 19.1 | 13.2 | 35.5 | 99.0 | 72.6 | 44.4 | 54.8 | 23.0 | 50.0 | 91.4 | 74.6 | 36.8 | 0.0 | 42.8 | 21.1 | 34.7 | 85.4 | 79.2 | 91.0 | 51.7 | 43.2 | 38.8 | 60.2 | 48.0 | 82.1 | 54.2 |
| Task-expert averaging | 68.2 | 57.2 | 55.8 | 76.5 | 20.6 | 12.2 | 35.0 | 97.0 | 77.4 | 45.7 | 55.5 | 27.1 | 50.7 | 89.5 | 72.0 | 32.3 | 12.1 | 40.2 | 25.4 | 26.4 | 86.1 | 74.5 | 71.0 | 49.6 | 44.8 | 42.2 | 63.9 | 50.0 | 83.5 | 54.5 |
| Capability-expert averaging | 70.2 | 55.4 | 58.2 | 77.6 | 20.4 | 12.6 | 35.4 | 100.0 | 76.5 | 46.5 | 55.9 | 25.3 | 50.5 | 89.1 | 79.2 | 37.9 | 58.1 | 42.1 | 26.5 | 27.1 | 82.9 | 92.2 | 100.0 | 53.7 | 45.3 | 44.2 | 65.8 | 51.0 | 79.6 | 56.8 |
| Capability-expert off-policy distill. | 75.2 | 55.0 | 53.6 | 87.3 | 23.6 | 13.2 | 38.1 | 100.0 | 70.3 | 45.5 | 52.4 | 25.0 | 49.8 | 89.5 | 55.5 | 20.4 | 54.8 | 50.6 | 29.1 | 40.4 | 79.8 | 96.1 | 100.0 | 56.0 | 54.1 | 47.7 | 72.3 | 65.0 | 85.1 | 59.3 |
| Variant | F | O | N | T | S | Avg. |
|---|---|---|---|---|---|---|
| Socio-Foundation | 63.93 | 63.85 | 54.41 | 50.50 | 75.77 | 61.69 |
| w/o F | 60.72 | 66.99 | 53.33 | 49.38 | 72.81 | 60.65 |
| w/o O | 62.57 | 63.28 | 52.46 | 49.32 | 73.95 | 60.32 |
| w/o N | 61.66 | 64.12 | 46.66 | 49.19 | 74.59 | 59.24 |
| w/o T | 62.44 | 65.20 | 51.57 | 49.36 | 74.11 | 60.54 |
| w/o S | 62.94 | 66.03 | 51.87 | 48.93 | 68.85 | 59.72 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Cluster | Theme / dimension | Capability requirements |
|---|---|---|
| C1 | Preference / Value Fidelity ( F ) | A01 Preference alignment; A02 Viewpoint and value fidelity |
| C2 | Persona / Expression Fidelity ( F ) | A03 Persona maintenance; A04 Individual expression-style fidelity |
| C3 | Behavioral / Decision Fidelity ( F ) | A05 Individual response and choice fidelity; A06 User-intent fidelity; A10 Individual cognitive and decision-pattern fidelity |
| C4 | Goal / Outcome Realization ( O ) | A07 Individual goal realization |
| C5 | Human-like Expression / Behavior ( N ) | A08 Natural dialogue behavior; A09 Human-like behavior |
| C6 | Trajectory / Narrative Coherence ( T ) | A11 Behavior-sequence and cross-turn consistency; A12 Narrative progression and response consistency |
| Dataset / task | Task interpretation | Requirements | Clusters | Dimensions |
|---|---|---|---|---|
| Surveyed dataset/task entries | ||||
| alignx_v2 † | Preference-conditioned response selection. | A01 | C1 | F |
| convokit_casino-corpus | Negotiation with private item preferences and individual objectives. | A01, A07 | C1, C4 | F, O |
| convokit_chromium-corpus | Participants in software-review discussions. | A08, A16 | C5, C8 | N, S |
| convokit_emotional-support | Emotional-support conversations. | A15, A18 | C8 | S |
| convokit_mediasum-corpus | Dialogue simulation from interview transcripts. | A08 | C5 | N |
| Benchmark / entry | Cluster / dimension | Selected metrics and interpretation |
|---|---|---|
| LifeChoices | C1, C3 / F | Choice accuracy: persona-conditioned preference and choice fidelity. |
| AlignX | C1 / F | Direct-choice accuracy: personalized preference fidelity. |
| HumanLLM | C1 / F | Hit@5: personalized item selection; the story/scenario membership in the survey is outside this metric. |
| BehaviorChain | C3 / F; C6 / T | Node accuracy (F); chain-level CumScore (T). |
| CoSER | C2 / F; C5 / N; C6 / T | Character fidelity (F); anthropomorphism (N); storyline consistency and storyline quality (T). |
| HUMANUAL | C1–C3 / F | Response alignment and state alignment. The six domain entries supply the relevant individual-reference requirements. |
| Dataset / source | Before | After | Removed | Transformed | Share (%) |
|---|---|---|---|---|---|
| alignx_v2 | 14,649,170 | 14,086,414 | 562,756 | 1,097 | 74.0663 |
| socsci210 | 2,618,745 | 2,547,535 | 71,210 | 30,311 | 13.3949 |
| human_llm | 1,316,829 | 1,316,811 | 18 | 535 | 6.9238 |
| wildchat | 165,346 | 157,220 | 8,126 | 3,726 | 0.8267 |
| coser | 114,831 | 114,831 | 0 | 8 | 0.6038 |
| lmsys | 79,929 | 79,908 | 21 | 167 | 0.4202 |
| Dataset | AlignX | BehaviorChain | CoSER | FANToM | HumanLLM | HUMANUAL-Book | HUMANUAL-Chat | HUMANUAL-Opinion | HUMANUAL-Politics | LifeChoices | MirrorBench | Social-R1 | SOTOPIA | UserLM-S3 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RL training instances | 8,192 | 1,024 | 1,024 | 1,024 | 1,024 | 1,000 | 1,000 | 1,000 | 1,000 | 1,150 | 1,024 | 687 | 405 | 1,024 |
| Dim. | Benchmark | Metric | Description |
|---|---|---|---|
| F | LifeChoices | Acc. | Accuracy of persona-conditioned choices against the reference decisions. |
| F | AlignX | Choice acc. | Accuracy of direct choices against the target individual’s reference preferences. |
| F | HumanLLM | Hit@5 | Fraction of cases in which the reference item appears among the top five predicted items. |
| F | BehaviorChain | Node | Micro-averaged correctness of individual behavior predictions across chain nodes. |
| F | CoSER | Char. fid. | Scene-level assessment of consistency with the assigned character’s attributes and behavior. |
| F | HUMANUAL | Response | Alignment of generated responses with the target person’s reference responses. |
| Dataset / benchmark | Sampling unit | Sampled units | Cases |
|---|---|---|---|
| IndiEval | |||
| FANToM | Complete question set (one per conversation) | 100 | 1,468 |
| Social-R1 | Question | 100 | 100 |
| LifeChoices | Decision question (balanced sampling across books) | 500 | 500 |
| BehaviorChain | Complete persona behavior chain | 100 | 1,529 |
| AlignX | Core preference sample (five conditioning variants each) | 100 | 500 |
| Benchmark | Cases/run | Metric | Mean SD | Min. | Max. | Range |
|---|---|---|---|---|---|---|
| LifeChoices | 500 | Accuracy | 79.000 | 80.600 | 1.600 | |
| HUMANUAL | 600 | Response alignment | 11.716 | 13.325 | 1.609 | |
| State alignment | 35.676 | 37.023 | 1.347 | |||
| MirrorBench | 200 | GTEval | 53.975 | 54.925 | 0.950 | |
| PI | 60.833 | 64.333 | 3.500 | |||
| RNR | 73.250 | 77.500 | 4.250 |