cs.CLSep 24, 2026

Artificial Societies Benchmark: A Validation Framework for Synthetic Research

Authors: Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He

Organizations: Artificial Societies University of Oxford, UK

Abstract

A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research

    Sep 23, 2026Florian Kutzner, Celina Kacperski, Laura de Molière +4ValidationSurvey

  2. When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

    Jul 28, 2026Zihan Chen, Di Zhu, Lei Nico ZhengSurveyDemographics

  3. Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data

    Feb 10, 2026Jason Miklian, Kristian Hoelscher, John E. KatsosSurveySynthetic Data