Large Language Models (LLMs) offer a scalable way to simulate survey respondents using demographic profiles and observed reference responses. However, the conventional approach of predicting one question per prompt repeatedly encodes the same context, limits each target to a narrow set of reference responses, and prevents later predictions from using information in earlier answers. Predicting multiple questions in one prompt can reduce these costs, share a broader pool of references, and let later predictions build on earlier ones. This requires forming coherent batches, selecting shared references, and ordering questions and references effectively. We propose Semantically Coherent Batching and Ordering (SCBO), a training-free framework that addresses these challenges. SCBO first uses an LLM to extract compact semantic representations from survey items and filter out template noise. It then groups related questions into batches and builds a shared reference bank using target-specific retrieval and centroid-based completion. Finally, it orders target questions from easy to hard and arranges references according to their semantic alignment with those questions. Experiments on four large-scale survey datasets and four LLMs show that SCBO substantially reduces token consumption and inference time while generally improving prediction accuracy over a non-batched baseline. Code is available at https://anonymous.4open.science/r/SCBO-41D8.
Figures & tables
Avg. tokens
Token budget (M)
Estimated cost (USD)
Dataset
Users
Questions
(User, Question)
Context
Question
Input
Output
GPT-6
Fable-5
Gemini-3.8
WVS
97,220
285
25,038,918
491.1
70.6
14,064
125
$11,126+
$58,552+
$11,017+
ANES
5,521
952
3,541,672
551.4
62.9
2,176
18
$1,714+
$9,026+
$1,698+
GSS
3,986
539
1,001,083
539.5
55.1
595
5
$470+
$2,473+
$465+
BSA
4,120
251
531,064
509.5
63.2
304
3
$240+
$1,265+
$238+
Table 1: Estimated token usage and API cost for exhaustive one-question-per-prompt prediction. Prices are based on the public rates listed at https://aiapi.world/pricing .
Figure 1: Overview of SCBO. As preprocessing, Question Instantiation extracts compact topic, intent, and entity from survey questions to filter template noise for the two modules: (1) Semantic Batching clusters semantically cohesive questions and builds shared reference banks through target-specific and centroid-based retrieval; and (2) Curriculum-based Ordering sorts target questions from easy to hard and orders references within each shared bank by semantic alignment.
LLM
Budget
Setting
WVS
GSS
ANES
BSA
ACC
MAE
F1
ACC
MAE
F1
ACC
MAE
F1
ACC
MAE
F1
Deepseek V4-Flash
1-budget
w/o
0.506
0.928
0.493
0.491
0.742
0.478
0.508
0.917
0.492
0.486
0.713
0.476
w
0.543
0.798
0.516
0.517
0.678
0.488
0.531
0.861
0.510
0.545
0.607
0.533
Δ (%)
7.3
14.0
4.6
5.3
8.6
2.1
4.5
6.1
3.7
12.1
14.9
12.0
2-budget
w/o
0.522
0.883
0.512
0.528
0.676
0.510
0.527
0.868
0.508
0.492
0.674
0.485
w
0.548
0.761
0.525
0.547
0.615
0.527
0.556
0.793
0.535
0.552
0.589
0.532
Table 2: Effectiveness across reference budgets, LLMs, and datasets. (w/o) denotes the non-batched baseline, (w) denotes SCBO, and Δ (%) reports relative change. Green indicates improvement, while gray indicates degradation.
LLM
Metric
Setting
WVS
GSS
ANES
BSA
1
2
3
1
2
3
1
2
3
1
2
3
DeepSeek V4-Flash
TPQ
w/o
577
661
747
589
659
727
621
698
774
541
613
681
w
194
280
380
170
253
401
181
257
344
187
247
329
Redu.(%)
66.3
57.7
49.1
71.1
61.6
44.8
70.8
63.3
55.5
65.4
59.7
51.7
SPQ
w/o
0.947
1.085
0.839
0.851
0.927
0.829
0.889
1.026
0.859
1.162
1.063
0.978
w
0.093
0.115
0.146
0.105
0.114
0.142
0.086
0.093
0.100
0.101
0.117
0.124
Table 3: Efficiency across reference budgets, LLMs, and datasets. Columns 1–3 denote the per-question reference budget; (w/o) denotes the non-batched baseline and (w) denotes SCBO. Green indicates improvement, while gray indicates degradation.
Figure 2: Accuracy and token cost across batch sizes on WVS. The figure compares 1-budget and 3-budget settings using DeepSeek-V4-Flash and Qwen3.7-Max.
Figure 3: Ablation study on WVS across different component variants and reference budgets.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Capacity-constrained semantic batching. KMeans++ identifies semantic centers from question embeddings. Each center is expanded into slots according to its capacity, and the Hungarian algorithm computes a one-to-one assignment that produces capacity-constrained batches.
Figure 5: Dataset scale and evaluation allocation. (a) Complete-release scale, where the horizontal axis denotes the number of respondents, the vertical axis denotes the number of survey questions, and marker area represents the number of observed respondent–question pairs. (b) Number of target questions assigned to each respondent under reference budgets. (c) Total target pairs evaluated in each method–model run. WVS uses Wave 7, while GSS, ANES, and BSA use their 2024 releases.
Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferences across the whole questionnaire, not merely accurate distributions for isolated items. Existing single-item methods cannot accurately reproduce how the same person answers a complete survey. We propose FullRespondent-LLM (FR-LLM), which fine-tunes two specialized LLMs: a marginal model for each item's response distribution and a respondent-level autoregressive model for dependencies across answers. Marginal-Constrained Joint Projection (MCJP) then projects the autoregressive joint distribution onto the set satisfying the item-level marginals learned by the first model. This yields complete questionnaires with realistic cross-item relationships while retaining strong item-level accuracy. On two real-world social survey datasets, FR-LLM more accurately reproduces multi-question response patterns, maintains competitive single-item accuracy, and generalizes better to unseen populations and questions. In a small commercial-survey dataset, we use simulated responses to make pricing and stocking decisions; FR-LLM achieves the highest realized profit.
Ji Huang, Mengfei Li, Shuai Shao
School of Computer Science, University of Science and Technology of China · School of Management, Fudan University
Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions. We study the design problem of allocating a fixed budget of human respondents across estimation tasks when cheap LLM predictions are available for every task. Our framework combines three components. First, building on Prediction-Powered Inference, we characterize a question-specific rectification difficulty that governs how quickly the estimator's variance decreases with human sample size. Second, we derive a closed-form optimal allocation rule that directs more human labels to tasks where the LLM is least reliable. Third, since rectification difficulty depends on unobserved human responses for new surveys, we propose a meta-learning approach, trained on historical data, that predicts it for entirely new tasks without pilot data. The framework extends to general M-estimation, covering regression coefficients and multinomial logit partworths for conjoint analysis. We validate the framework on two datasets spanning different domains, question types, and LLMs, showing that our approach captures 61-79% of the theoretically attainable efficiency gains, achieving 11.4% and 10.5% MSE reductions without requiring any pilot human data for the target survey.
Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.
Chan-Tung Ku, Chan Hsu, Pei-Cing Huang +3
Department of Information Management National Sun Yat-sen University Kaohsiung, Taiwan · Institute of Political Science National Sun Yat-sen University Kaohsiung, Taiwan · Graduate Institute of Library and Information Science National Chung Hsing University Taichung, Taiwan