Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferences across the whole questionnaire, not merely accurate distributions for isolated items. Existing single-item methods cannot accurately reproduce how the same person answers a complete survey. We propose FullRespondent-LLM (FR-LLM), which fine-tunes two specialized LLMs: a marginal model for each item's response distribution and a respondent-level autoregressive model for dependencies across answers. Marginal-Constrained Joint Projection (MCJP) then projects the autoregressive joint distribution onto the set satisfying the item-level marginals learned by the first model. This yields complete questionnaires with realistic cross-item relationships while retaining strong item-level accuracy. On two real-world social survey datasets, FR-LLM more accurately reproduces multi-question response patterns, maintains competitive single-item accuracy, and generalizes better to unseen populations and questions. In a small commercial-survey dataset, we use simulated responses to make pricing and stocking decisions; FR-LLM achieves the highest realized profit.
Figures & tables
Dataset
Train n
M1 test n
M2 test n
M3 test n
ESS11
28,268
28,268
7,312
7,312
TALIS 2018
12,000
8,000
8,000
8,000
Table 1: Distinct respondents by split. M1: unseen constructs; M2: unseen countries; M3: both. M2 and M3 share respondents but use different items.
Mode
Method
Qwen3.5-9B
Ministral-3-8B
Cronbach’s α MAE
AVE MAE
Item JSD
Construct JSD
Cronbach’s α MAE
AVE MAE
Item JSD
Construct JSD
M1
Zero-shot
.507
.348
.075
.124
.442
.316
.106
.168
Single FT
.363
.272
.024
.037
.367
.274
.031
.048
Sequential FT
.084
.075
.044
.047
.114
.103
.089
.091
FR-LLM
.078
.071
.024
.021
.046
.040
.031
.032
M2
Zero-shot
.625
.405
.084
.175
.524
.370
.107
.201
Table 2: ESS11 results for both backbone models. Lower is better. Bold is best and underline is second best within each backbone and transfer mode; ties receive the same mark.
Mode
Method
Qwen3.5-9B
Ministral-3-8B
Cronbach’s α MAE
AVE MAE
Item JSD
Construct JSD
Cronbach’s α MAE
AVE MAE
Item JSD
Construct JSD
M1
Zero-shot
.809
.395
.061
.296
.808
.392
.141
.474
Single FT
.642
.362
.017
.144
.831
.396
.020
.168
Sequential FT
.094
.142
.015
.058
.114
.181
.029
.075
FR-LLM
.008
.007
.017
.043
.049
.058
.020
.052
M2
Zero-shot
.786
.387
.069
.309
.811
.388
.150
.486
Table 3: TALIS 2018 results for both backbone models. Lower is better. Bold is best and underline is second best within each backbone and transfer mode; ties receive the same mark.
Method
Train n
Test n
α MAE
AVE MAE
Item JSD
Construct JSD
Zero-shot
0
1,000
.494
.329
.073
.131
Single FT
1,000
1,000
.406
.290
.054
.042
FR-LLM
1,000
1,000
.066
.061
.054
.022
Single FT
28,268
28,268
.363
.272
.024
.037
FR-LLM
28,268
28,268
.078
.071
.024
.021
Table 4: ESS11 M1 low-data results (Qwen). Full-data rows reproduce Table 2 ; bold marks the best low-data value. Lower is better.
Background change
Zero-shot
Single FT
FR-LLM
Low to high education
.261
.153
.149
Age 18–29 to 60+
.399
.231
.216
Urban to rural
.123
.112
.111
Low/young to high/old
.298
.198
.177
Urban/male to rural/female
.229
.150
.168
Macro average
.262
.169
.164
Table 5: ESS11 background-shift MAE across five contrasts (response-scale points; lower is better).
Human
Excess error above human resampling
FR gap
Metric
error
Zero-shot
Single FT
Sequential FT
FR-LLM
closed
α MAE
.007
.500
.356
.077
.071
85.9%
AVE MAE
.007
.342
.265
.068
.065
81.1%
Item JSD
.007
.067
.017
.037
.017
74.4%
Construct JSD
< .001
.124
.037
.047
.021
82.9%
Table 6: ESS11 M1 gap to human resampling (Qwen). Excess error is model error minus the mean human–human error over 1,000 stratified resample pairs; lower is better. Gap closed is (EZS−EFR)/(EZS−Ehuman) ; higher is better.
Configuration
α MAE
AVE MAE
Item JSD
Construct JSD
Zero-shot
.507
.348
.075
.124
Single FT
.363
.272
.024
.037
Joint proposal (16 samples)
.076
.070
.027
.028
MCJP (1 sample)
.087
.080
.024
.022
MCJP (16 samples)
.078
.071
.024
.022
Table 7: ESS11 M1 component and candidate-sampling ablation (Qwen; lower is better). The joint proposal uses 16 samples without projection; MCJP uses learned Stage-1 margins. Rows with different candidate counts are frozen-model re-evaluations.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Data / backbone
Method
M1
M2
M3
ESS11 / Qwen
Zero-shot
.247
.366
.296
Single FT
.099
.142
.160
Sequential FT
.110
.165
.184
FR-LLM
.059
.081
.123
ESS11 / Ministral
Zero-shot
.291
.376
.340
Single FT
.119
.150
.182
Appendix
Table 8: Two-construct JSD. Lower is better. Bold is best and underline is second best within each dataset, backbone, and applicable mode.
Candidate questionnaires
Item JSD
Two-construct JSD
Maximum marginal error
Adjusted cells
Effective sample fraction
1
.0244
.0693
.1952
469
.330
2
.0244
.0621
.0748
201
.422
4
.0244
.0596
.0190
76
.506
8
.0244
.0596
.0112
18
.555
16
.0244
.0592
.0013
2
.578
Appendix
Table 9: ESS11 M1 candidate-support audit. Item JSD remains anchored while joint support and feasibility improve.
ρ
Item JSD
Construct JSD
Two-construct JSD
Support adjustment
Effective sample fraction
0.00
.0267
.0280
.0708
.0000
1.000
0.25
.0251
.0254
.0661
.0013
.962
0.50
.0242
.0240
.0636
.0027
.860
0.75
.0240
.0224
.0606
.0040
.722
1.00
.0244
.0217
.0596
.0053
.578
Appendix
Table 10: ESS11 M1 anchor path. The proposal and candidate support are fixed; only the target margin changes.
Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated. We ask a simple question: given a small pilot sample of human responses, can an LLM recover the statistical characteristics of a broader population? We decompose recovery along three axes: structural fidelity, marginal fidelity, and individual fidelity. Using a COVID-19 misinformation survey as a case study, we benchmark three families of approaches: prompting, rectification, and fine-tuning. The findings suggest that fine-tuning on small pilot samples offers a balanced approach for achieving multiple forms of fidelity, but the levels of such fidelity can vary across subsamples, potentially threatening pluralistic alignment.
Eun Cheol Choi, Youngrae Kim, Prabhu Pugalenthi +2
Large Language Models (LLMs) offer a scalable way to simulate survey respondents using demographic profiles and observed reference responses. However, the conventional approach of predicting one question per prompt repeatedly encodes the same context, limits each target to a narrow set of reference responses, and prevents later predictions from using information in earlier answers. Predicting multiple questions in one prompt can reduce these costs, share a broader pool of references, and let later predictions build on earlier ones. This requires forming coherent batches, selecting shared references, and ordering questions and references effectively. We propose Semantically Coherent Batching and Ordering (SCBO), a training-free framework that addresses these challenges. SCBO first uses an LLM to extract compact semantic representations from survey items and filter out template noise. It then groups related questions into batches and builds a shared reference bank using target-specific retrieval and centroid-based completion. Finally, it orders target questions from easy to hard and arranges references according to their semantic alignment with those questions. Experiments on four large-scale survey datasets and four LLMs show that SCBO substantially reduces token consumption and inference time while generally improving prediction accuracy over a non-batched baseline. Code is available at https://anonymous.4open.science/r/SCBO-41D8.
Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.
Chan-Tung Ku, Chan Hsu, Pei-Cing Huang +3
Department of Information Management National Sun Yat-sen University Kaohsiung, Taiwan · Institute of Political Science National Sun Yat-sen University Kaohsiung, Taiwan · Graduate Institute of Library and Information Science National Chung Hsing University Taichung, Taiwan