Authors: Augusto Gonzalez-Bonorino, Kseniia Biriukova, Monica Capra
Organizations: Department of Economics, Arizona State University · EconLLM Lab · Department of Information Systems, Arizona State University · Department of Economics, Claremont Graduate University
Population prompts are widely used to generate synthetic survey responses, but they combine information supplied at inference with associations already encoded during pretraining. We introduce an alternative construction that maps declared aggregate preference anchors into group-indexed choice policies. For each population, the signs of six Global Preferences Survey (GPS) coordinates deterministically label a shared bank of paired synthetic responses, and Direct Preference Optimization fits a parameter-efficient adapter to those comparisons. We evaluate the adapters on candidate World Values Survey (WVS) items using prompts that omit country names and distinguish four questions: recovery of the imposed labels, transfer of the anchor signal to new text, coherence between the GPS anchors and human WVS responses, and agreement between adapter and human scores. The adapters recover the imposed pairwise labels. On a purposively selected sixteen-country development panel, adapter trust scores completely separate the two GPS-sign groups and have a rank correlation of (0.74) with continuous GPS trust scores. Human-GPS and adapter-human associations remain unresolved on the same panel, and results for the other preference dimensions are heterogeneous. These findings show that an anchored policy can retain a declared aggregate signal without thereby reproducing human response patterns. The contribution is therefore both an inspectable construction and an evaluation framework that separates anchor transfer from human criterion agreement.
Figures & tables
Approach
Information supplied
Fitting or evaluation target
CultureLLM [ 11 ]
WVS seeds and semantic augmentation
Culture-specific adaptation
Cao et al. [ 12 ]
Country survey distributions
Distributional fit and cross-survey evaluation
SubPOP [ 13 ]
Subpopulation survey responses
Distribution prediction and transfer
Jung and Kim [ 14 ]
Korean cultural policy and synthetic triplets
DPO for cultural response coherence
Krsteski et al. [ 15 ]
Limited human responses
Allocation to synthesis and rectification
Present construction
Aggregate preference signs and a shared contrast bank
Table 1: Selected methodological comparisons of information inputs and evaluation targets.
Figure 1: Construction and readout. Profile prose and score magnitudes document the anchors but do not feed the deterministic label rule. The shared bank and the pretrained model also influence the resulting policy.
Anchor
WVS items
Interpretation to be assessed
Trust
Q57–Q64, Q73
Generalized, interpersonal, and institutional trust/confidence
Patience
Q43; Q50
Less importance placed on work; financial satisfaction
Risk taking
Q106, Q107, Q109, Q178
Economic values and fare avoidance
Positive reciprocity
Q81
Confidence in charitable organizations
Negative reciprocity
Q176, Q177, Q179, Q195
Moral clarity and justifiability judgments
Altruism
Q99, Q101, Q103
Membership in environmental, charitable, and self-help groups
Table 2: Candidate indicators on the strict common-support panel. Each row defines a candidate item set; latent-scale validity remains to be assessed. Patience items are interpreted separately.
Members
trust
risk
patience
altruism
pos. recip.
neg. recip.
EGY, IDN
+
−
−
+
+
+
GRC, IND
−
−
−
−
−
+
MEX, RUS
−
−
−
−
−
−
ARG
−
+
−
+
+
−
BRA
−
−
−
+
+
−
CHN
+
−
+
+
+
+
Table 3: Thirteen distinct GPS sign profiles among the sixteen adapter countries. Coordinates are ordered trust, risk taking, patience, altruism, positive reciprocity, and negative reciprocity, matching Appendix A . The United States is non-negative on every coordinate; Mexico and Russia are negative on every coordinate. Identical sign vectors induce identical labels on the shared bank.
Figure 2: Sixteen-country panel. Panel A is association with the GPS anchors. Panel B is agreement with WVS country responses. Markers are Spearman correlations; bars are 95% country-bootstrap intervals from 2,000 resamples. The intervals are conditional within-panel stability summaries. They exclude training-seed variation, respondent resampling, and GPS measurement error. Patience is omitted because Q43 and Q50 are reported separately.
Anchor
ρ
Interval
Trust
-0.07
[-0.55, 0.45]
Risk taking
-0.11
[-0.72, 0.51]
Positive reciprocity
0.02
[-0.47, 0.48]
Negative reciprocity
0.34
[-0.17, 0.82]
Altruism
-0.23
[-0.64, 0.26]
Table 4: Country-named persona–GPS association on the strict panel, with 95% country-bootstrap intervals. This supplies the anchor comparator for Figure 2 .
Pairing
Q43
Q50
Human–GPS
-0.56 [-0.86, -0.04]
0.62 [0.11, 0.91]
Adapter–GPS
0.01 [-0.47, 0.44]
0.44 [-0.12, 0.79]
Adapter–human
0.29 [-0.19, 0.73]
0.25 [-0.32, 0.70]
Persona–GPS
-0.62 [-0.90, -0.10]
0.62 [0.16, 0.87]
Persona–human
0.59 [0.11, 0.82]
0.70 [0.23, 0.91]
Table 5: Patience candidate indicators, separately scored on sixteen countries. Entries are Spearman correlations; brackets give 95% country-bootstrap intervals.
Configuration
Trust–GPS ρ
Mean TVD
Base model + neutral prompt
—
0.462
Base model + country prompt
−0.07
0.369
Profile adapter + neutral prompt
0.74
0.487
Profile adapter + country prompt
0.67
0.401
Table 6: Configuration sensitivity on the strict panel. TVD averages 368 country–item cells. The Profile adapter + country prompt trust cell ( 0.67 ) was not recomputed from the two saved jobs. The base TVD uses banked runs. No base rank association with the anchor is interpreted, because run variation is confounded with the bank. Rank associations on this table use the same country bootstrap as Figure 2 , which excludes training-seed variation and GPS measurement error.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
WVS
GPS link
Abbreviated wording
g(v)
Range
Status
Q57
GPS trust
Most people can be trusted (1) vs need to be very careful (2)
1{v=1}
0–1
primary rectangle
Q58
GPS trust
Trust: your family (1 trust completely … 4 do not trust at all)
5−v
1–4
primary rectangle
Q59
GPS trust
Trust: your neighborhood (1 trust completely … 4 do not trust at all)
5−v
1–4
primary rectangle
Q60
GPS trust
Trust: people you know personally (1 trust completely … 4 do not trust at all)
5−v
1–4
primary rectangle
Q61
GPS trust
Trust: people you meet for the first time (1 trust completely … 4 do not trust at all)
5−v
1–4
primary rectangle
Q62
GPS trust
Trust: people of another religion (1 trust completely … 4 do not trust at all)
Table 8: Human–GPS country-rank associations on the sixteen-country item map, using the same recodes as the sixteen-country rectangle. Rows require at least fifty valid weighted responses per item. Complete-item rows retain only countries observed on every listed item. Intervals are 95% country bootstrap (2,000 draws) on the human panel and omit GPS measurement error.
A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.
Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis +1
Generative Synthetic Populations (GSP) -- the convergence of population synthesis, agent-based modelling, and LLM agents -- are attracting growing interest for urban simulation and institutional communication research. Before any GSP instrument is used on a real population, a more basic question must be answered: does it respond to stimuli of known valence in an ordered, replicable, group-structured way? We call this controllability. We ask not whether a synthetic population tracks humans, but whether it tracks itself: whether the latent structure we impose on it is recovered in its own responses. This internal-validity question is logically prior to any claim about external validity, just as characterising an instrument's response function must precede using it to test a theory. We report SIVE (Synthetic Instrument Validation Experiment): a fictional municipality (Montelago) with 120 synthetic personas of known latent structure, exposed to seven conditions spanning strongly positive to strongly negative institutional communications about a water network. Seven pre-registered criteria, evaluated across a temperature sweep, jointly assess fidelity, stability, noise floor, specificity, sensitivity, and ordering. All seven pass at every temperature. A central finding turns a calibration failure into a diagnostic success: a message designed as "weakly positive" was identified by the instrument as functionally negative, traced to unresolved problems, uncertainty, and institutional passivity in its text; a redesigned version restored the expected ordering and interacts with agents' latent trust in unanticipated ways. A noise sub-experiment shows the instrument's intrinsic noise is roughly half the cross-agent estimate and stable across temperatures. Individual trajectories reveal coherent micro-dynamics that summary statistics obscure. Full data are available via an interactive explorer.
Mirko Degli Esposti
Department of Physics and Astronomy, University of Bologna
Demographic synthetic survey panels are often validated by matching aggregate answers to published surveys. We test what that certificate establishes across six multiselect batteries from four survey organisations in three countries. The headline analysis is restricted to three instruments whose synthetic cohort and human target share the stated population frame; three other batteries remain sensitivity analyses. The response contract dominates measured fidelity. In the aligned instruments, committed sets leave 66 of 128 model-battery option slots empty in panels of up to 500 respondents, versus 0 of 128 under per-option probability elicitation. Across eight uncapped model-instrument comparisons, probabilities reduce option-marginal MAE by 4.53 to 7.30 points. The capped instrument reverses on two models until the vectors are projected onto its stated maximum. These are measurement effects: human targets are realised check-all responses, whereas the vectors are latent inclusion propensities. Published marginal agreement also fails to discriminate respondent simulation from direct population estimation. On nine aligned model-battery pairs, a no-persona population-prevalence query averages 6.27 MAE versus 12.39 for committed panels and wins all nine comparisons. Constraint-aware probability vectors average 5.34 and beat the query on four of nine, so the baseline challenges the validation criterion rather than proving direct estimation uniformly best. On three unpublished demographic cells, neither approach beats reciting the national distribution. Population-marginal agreement is therefore evidence about an elicitation contract and an estimand obtainable without simulated respondents, not evidence of individual simulation.