Artificial Societies Benchmark: A Validation Framework for Synthetic Research
Organizations: Artificial Societies University of Oxford, UK
Abstract
A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.
Figures & tables
| Test | Measure | Comparison | What it captures |
|---|---|---|---|
| Internal validity | |||
| IV-1 | Repeated questions | Repeat the same item | Excess consistency or variability |
| IV-2 | Linked questions | Check logical constraints | Departures from human violation patterns |
| IV-3 | Paraphrased questions | Ask equivalent wording | Changes in answers across wording |
| IV-4 | Perturbed surveys | Change survey administration | Human sensitivity to the instrument |
| Construct validity | |||
| Model | Weight availability | Temperature | Thinking |
|---|---|---|---|
| GPT-5.6 Sol | Proprietary | 1 | Off |
| Claude Opus 5 | Proprietary | 1 | Off |
| Gemini 3.8 Flash | Proprietary | Ignored | Low |
| Grok 4.3 | Proprietary | 1 | Off |
| Mistral Small 2603 | Open-weight | 1 | Off |
| DeepSeek Flash | Open-weight | 1 | Off |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Test | Measure | Human-referenced diagnostics |
|---|---|---|
| IV-1 | Repeated questions | Repeated-answer agreement, kappa, and numerical change |
| IV-2 | Linked questions | Constraint violations and their joint patterns |
| IV-3 | Paraphrased questions | Agreement and response-transition distance |
| IV-4 | Perturbed surveys | Response shifts under instrument perturbations |
| CV-1 | Item covariance | Reliability and within-construct item associations |
| CV-2 | Independent criteria | Associations with independently measured criteria |
| Condition | Information we supply to the model | Output | Sources |
|---|---|---|---|
| L1 aggregate | Survey context, question, and options for the national adult population | Percentages over options | GSS |
| L2 composition | L1 plus the unweighted demographic composition of the task roster in five declared age bands | Percentages over options | GSS |
| L3 demographic | One respondent’s recorded age, sex, ethnicity, education, urbanicity, and political affiliation, or the recorded fields a source provides | One answer | GSS and every profiled task |
| L3 plus history | L3 plus an extractive summary of the same person’s earlier closed-ended answers | One answer | Twin |
| L4 verbatim | L3 plus original free text, being three self-descriptions in Twin and pre-election likes, dislikes, and problem statements in ANES 2020 | One answer | Twin, ANES 2020 |
| L4 distilled | L3 plus a seven-field summary of the free text, which one shared preprocessing model extracts as quotations and we audit against the source, with and without history | One answer | Twin, 430 people |
| Model | Provider, route | Regime | Decoding controls |
|---|---|---|---|
| gpt-5.6-sol | OpenAI, batch | Thinking disabled | Temperature 1, reasoning effort none, caps 100/512 |
| claude-opus-5 | Anthropic, batch | Thinking disabled | Temperature 1, thinking off, effort high (provider default), caps 100/512 |
| gemini-3.8-flash | Google, online | Low thinking, separate | Backend ignores temperature, low thinking, cap 2,048 |
| grok-4.3 | xAI, batch | Thinking disabled | Temperature 1, reasoning off, caps 100/512 |
| mistral-small-2603 | Mistral, batch | Thinking disabled | Temperature 1, caps 100/512; both synchronous and Batch image limits refused one ESS continuation |
| deepseek-flash | DeepSeek, online | Thinking disabled | Temperature 1, thinking off, caps 100/512 |
| Test | IV-1 | IV-2 | IV-3 | IV-4 | CV-1 | CV-2 | CV-3 | EV-1 | EV-2 | EV-3 | EV-4 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| yielding an estimate | 5 | 5 | 5 | 12 | 8 | 5 | 4–5 | 17 | 14 | 9 | 5 |
| run | 5 | 5 | 5 | 12 | 9 | 5 | 5 | 17 | 15 | 9 | 6 |
| Model | Original text | Distilled text | History and distilled text |
|---|---|---|---|
| GPT-5.6 Sol | , | , | , |
| Claude Opus 5 | , | , | , |
| Gemini 3.8 Flash | , | , | , |
| Grok 4.3 | , | , | , |
| Mistral Small 2603 | , | , | , |
| DeepSeek Flash | , | , | , |
| Model | Original text | Distilled text | History and distilled text |
|---|---|---|---|
| GPT-5.6 Sol | |||
| Claude Opus 5 | |||
| Gemini 3.8 Flash | |||
| Grok 4.3 | |||
| Mistral Small 2603 | |||
| DeepSeek Flash |