MetaPersona: Task-Grounded Synthetic Populations from Empirical Social Science
Organizations: University of Southern California · University of California, Irvine · 2077AI · Arizona State University
Abstract
Personas used to seed LLM social simulations face a cold-start problem: existing methods lack a principled basis for deciding which attributes to include and how to assign their values. As a result, synthetic populations may misrepresent the demographic composition, latent attributes, and dependency structure that shape downstream behavior. We introduce MetaPersona-DB, a dataset of 11,000+ empirical human-subjects studies annotated with task-relevant variables, reported relationships, and aggregate-level population statistics. Building on this resource, we propose MetaPersona, a framework that retrieves task-relevant evidence, constructs literature-derived persona dependency graphs, and samples synthetic populations from empirical priors linking demographics, latent attributes, and outcomes. Across three downstream case studies, three baselines, and three frontier models, results vary by task and model: MetaPersona performs strongly on misinformation belief and AI-tool sentiment, while results on income redistribution are mixed. It also reduces persona-construction cost to under $0.5 per task using GPT-5.2. Finally, we present MetaPersona-Studio, a prototype interactive interface for empirically grounded persona generation.
Figures & tables
| Task | Discipline | Dataset | Predicted Outcome Variable |
|---|---|---|---|
| Misinformation belief | Communication, political science | MIST ( Maertens et al., 2024 ) | Ability to distinguish real from fake headlines, measured by numerical MIST-20 score |
| Income redistribution attitude | Economics, sociology | ISSP 2019 Social Inequality V (ZA7600, v3.0.0) | Agreement that government should reduce income differences, measured by ordinal scale |
| AI-tool sentiment & adoption | Behavioral science, HCI | Stack Overflow Survey ( Stack Overflow, 2025 ) | Sentiment toward AI tools and expected future use, measured by ordinal scales |
| Case Study 1: Misinformation (MIST) | Case Study 2: Redistribution (ISSP) | Case Study 3: AI Tool Sentiment (SO) | ||||||||||||
| Backbone | Method | RMSE | MAE | Bias | W-dist | KS | MAE | W-1 Acc | QWK | W-dist | MAE | W-1 Acc | QWK | W-dist |
| GPT-5.2 | Demo-Only | 0.310 | 0.257 | 0.248 | 0.248 | 0.619 | 0.812 | 0.817 | 0.086 | 0.583 | 0.248 | 0.743 | 0.059 | 0.808 |
| Backstory | 0.352 | 0.304 | 0.303 | 0.303 | 0.896 | 0.828 | 0.819 | 0.111 | 0.627 | 0.250 | 0.749 | 0.036 | 0.863 | |
| LLM-Heuristic | 0.322 | 0.268 | 0.255 | 0.257 | 0.666 | 0.842 | 0.808 | 0.128 | 0.570 | 0.252 | 0.735 | 0.073 | 0.871 | |
| MetaPersona | 0.270 | 0.216 | 0.118 | 0.118 | 0.346 | 0.832 | 0.814 | 0.087 | 0.642 | 0.217 | 0.743 | 0.320 | 0.616 | |
| Claude- Haiku 4.5 | Demo-Only | 0.238 | 0.192 | 0.151 | 0.153 | 0.475 | 0.884 | 0.802 | 0.170 | 0.416 | 0.293 | 0.667 | 0.052 | 0.931 |
| Task / Path | Method | ||
|---|---|---|---|
| MIST CRT discern. (expect ) | Ground truth | 0.313 | — |
| LLM-Heuristic | 0.238 | 0.075 | |
| MetaPersona | 0.293 | 0.020 | |
| ISSP education support (expect /mixed) | Ground truth | 0.043 | — |
| Demo-Only | 0.107 | 0.064 | |
| Backstory | 0.147 | 0.104 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | MIST | ISSP | SO |
|---|---|---|---|
| Number of personas | 1000 | 2915 | 2000 |
| LLM-Heuristic | $2.27 | $7.91 | $4.64 |
| Backstory | $8.52 | $27.68 | $16.84 |
| MetaPersona | $0.16 | $0.25 | $0.18 |
| Validation component | Measure | Result |
| Validation 1: Empirical paper labeling | ||
| Human agreement | Cohen’s | 0.94 |
| Human agreement | Raw agreement | 97% |
| LLM vs. adjudicated ground truth | Precision | 86% |
| LLM vs. adjudicated ground truth | Recall | 97% |
| LLM vs. adjudicated ground truth | F1 | 91% |
| Field | Definition | Examples / Allowed Values |
| Study-level descriptors | ||
| Domain or subfield | The specific empirical or disciplinary area of the study. | Political communication; health psychology; labor economics; organizational behavior. |
| Agent types | The social actors involved in the study. | Voters; consumers; patients; employees; managers; firms; institutions. |
| Context: country | The geographic setting of the study. | USA; Germany; China; cross-national sample. |
| Context: population | The sample, population, or empirical setting studied. | Registered voters; undergraduate students; low-income households; Fortune 500 employees. |
| Main finding | A one-sentence summary of the paper’s primary empirical result. | “Economic insecurity is associated with lower institutional trust.” |
| Task / Path | Model | Method | ||
|---|---|---|---|---|
| MIST CRT discern. (expect ) | Claude Haiku 4.5 | Ground truth | 0.313 | — |
| LLM-Heuristic | -0.172 | 0.485 | ||
| MetaPersona | -0.021 | 0.334 | ||
| DeepSeek v3.2 | Ground truth | 0.313 | — | |
| LLM-Heuristic | 0.423 | 0.110 | ||
| MetaPersona | 0.052 | 0.261 |