Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonstrate the framework's utility in two ways. First, we reproduce a prior study of "machine bias" in LLM survey responses and apply the framework to its models and more recent ones, showing that poor representation reflects not only insufficient between-group variation but also variation assigned to the wrong groups. Second, we evaluate one proposed approach to improving models' population representativeness: cultural fine-tuning. We find that cultural fine-tuning can improve alignment with the survey center without improving the representation of within-population differences, a distinction that measures of aggregate agreement do not capture. We argue that representing a population requires models to reproduce several features of human attitudinal variation simultaneously. Our framework organizes these features and provides reusable code, data, and trained models for evaluating population fidelity across substantive domains and assessing proposed alignment methods.
Figures & tables
Figure 1: Stylized example with three subpopulations (A–C). Colored curves show each group’s share of respondents and sum to the pooled distribution (gray), which is identical in every panel ( Scenter=1 for all models); dashed lines mark group means. Calculations are in Appendix B.3 .
Figure 2: Population Fidelity for happiness under FA elicitation. (a) Components, PFS, and center alignment for the indicated cell sets. (b) Fine-tuned minus as-released changes on target-country cells. Markers show three-run FA means using each run’s paired support; Table 10 , Appendix E.1 , uses the stricter all-runs common-cell intersection. All-cell results are in Figure 13 , Appendix D.5 .
Figure 3: PFS and center alignment by demographic level for happiness under FA elicitation. Boxes summarize 19 model conditions. Results for all questions: Figure 10 .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Reproduction of Figure 2 from Boelaert et al. (2025) : cell-level nEMD densities for the six published model–elicitation series, shown with the demographic and random baselines and the original response-quality bands.
Figure 5: The Population Fidelity framework compares survey and model distributions cell by cell ( Sacc ), in spread ( Sadapt ), and in arrangement ( Sstruct ); center alignment ( Scenter ) is reported separately. Section 4 defines each. A reference population provides n demographic cells with survey distributions si . Model responses to the corresponding demographic profiles produce distributions mi . Accuracy, adaptability, and structure compare the two sets of distributions and combine into the PFS by geometric mean; center alignment is reported separately. The metrics can also be recomputed within demographic subsets.
Variable
Topic
Options
d_happy
Happiness
4
d_polpos
Political ideology
10
d_religiousp
Religious attendance
7
d_trust
Social trust
2
Appendix
Table 4: Attitudinal items used in the evaluation.
Figure 7: PFS and center alignment by demographic level for the two proprietary models under FA elicitation, shown separately by question. Within each question, the upper row shows PFS and the lower row center alignment; demographic columns follow Figure 3 . Blue circles denote archived GPT-3 outputs from the original study, and green triangles denote GPT-5.6 Terra as released.
Figure 8: Cell-level error and adaptability under FA elicitation. The horizontal axis reports mean nEMD to the WVS cells, with lower values indicating greater accuracy. The vertical axis reports the adaptability ratio A . Values below A=1 indicate compression of between-group variation, while values above one indicate amplification. Adaptability must be interpreted alongside structure because matching the amount of variation does not establish that it occurs between the corresponding social groups.
Figure 9: Change in PFS by country after cultural fine-tuning for happiness under FA elicitation. Each row represents a fine-tuned condition and each column a surveyed country. Green indicates improvement relative to the corresponding as-released model, and red indicates decline. For happiness under FA, target-country changes are not consistently more favorable than changes in the other surveyed countries.
Figure 10: PFS and center alignment by demographic level under FA elicitation, shown separately by question. Filled boxes show the distribution of PFS and hollow boxes the distribution of center alignment across 19 model conditions; Mixtral reference series are excluded. Demographic columns follow Figure 3 .
Figure 11: PFS and center alignment by demographic level for individual model conditions under FA elicitation, shown separately by question. Within each question, the upper row shows PFS and the lower row center alignment; demographic columns follow Figure 3 . Color and marker shape identify the model. Hollow markers denote as-released conditions, filled markers German fine-tuning, and half-filled markers Mexican fine-tuning.
Figure 12: Adaptability and structure under NTP elicitation, shown separately by question. The horizontal axis reports the adaptability ratio A , and the vertical axis the untransformed structure correlation ρ . Dashed lines mark A=1 and ρ=0 . Dotted connectors link each population-wide estimate to the corresponding estimate recomputed within the target-country subset: Germany for German fine-tuning and Mexico for Mexican fine-tuning. As-released conditions show both target-country estimates, whereas fine-tuned conditions show only the estimate for their target country. Conditions with A>1 and ρ≈0 exhibit amplified between-group variation without recovering the observed structure of group differences.
Figure 13: Population Fidelity for happiness under FA elicitation on all retained cells. (a) Component scores for all 21 conditions, ordered by PFS; diamonds show PFS and hollow squares center alignment. (b) Fine-tuned minus as-released changes computed on the cells retained by both conditions. Figure 2 shows the corresponding target-country results. Plotting conventions are as in Figure 2 .
Figure 14: Changes in PFS and center alignment after cultural fine-tuning under FA elicitation, for the cross-question summary, political ideology, religious attendance, and social trust. Each point compares a fine-tuned condition with its corresponding as-released model. Movement to the right indicates improved PFS, and upward movement indicates improved center alignment. The happiness result appears in Figure 2(b) .
Figure 15: Population Fidelity components under FA elicitation, for the cross-question summary, political ideology, religious attendance, and social trust. Accuracy is generally the strongest component, whereas structure is the weakest and most often limiting; happiness is shown in Figure 2(a) and movement intervals in Table 9 . Fine-tuned conditions are labelled German ( Fde ) and Mexican ( Fmx ).
Question
Sacc
Sadapt
Sstruct
PFS
Scenter
ρ (PFS)
Happiness
0.001
0.013
0.015
0.019
0.001
0.960
Political ideology
0.001
0.008
0.017
0.027
0.001
0.920
Religious attendance
0.002
0.008
0.007
0.029
0.002
0.922
Social trust
0.002
0.021
0.012
0.045
0.002
0.926
All questions
0.002
0.013
0.013
0.031
0.002
0.972
Appendix
Table 9: Run-to-run uncertainty across repeated FA runs. Entries are pooled 95% CI half-widths, hpool , across the condition–question pairs in each row. The final column reports the mean of the three pairwise Spearman rank correlations in PFS between repeated runs; for “All questions,” correlations are computed after pooling the condition–question observations across the four questions. Smaller CI half-widths and larger rank correlations indicate greater stability across repeated elicitation.
Figure 16: Population Fidelity and center alignment for five Qwen3-VL-2B-Thinking conditions: as released, Fde , Fmx , distribution-matched, and subgroup-matched. (a,b) Component profiles on all retained cells, German cells, and Mexican cells for social trust under NTP and happiness under FA, respectively. Diamonds show PFS and hollow squares center alignment. (c,d) Adapted minus as-released changes on all retained cells for the same two settings.
Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated. We ask a simple question: given a small pilot sample of human responses, can an LLM recover the statistical characteristics of a broader population? We decompose recovery along three axes: structural fidelity, marginal fidelity, and individual fidelity. Using a COVID-19 misinformation survey as a case study, we benchmark three families of approaches: prompting, rectification, and fine-tuning. The findings suggest that fine-tuning on small pilot samples offers a balanced approach for achieving multiple forms of fidelity, but the levels of such fidelity can vary across subsamples, potentially threatening pluralistic alignment.
Eun Cheol Choi, Youngrae Kim, Prabhu Pugalenthi +2
Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in values, beliefs, and cultural norms that distinguish one group from another. We introduce Population Aligned Language Models (PALMs), a suite of models each aligned to specific populations, covering five countries: USA, India, Brazil, France and Italy. PALMs are created by synthesizing rationales grounded in psychological and cultural constructs and using these as latent supervision during preference tuning for population-specific alignment. Evaluated across four dimensions: personality, values and beliefs, cultural norms, and morality, PALMs consistently outperform baselines, including culture-specialized models, achieving an average of 8.59% relative improvement over the best baseline across all five populations. Notably, construct-grounded rationales outperform both demographic prompting and survey-based fine-tuning, suggesting that grounding preference learning in psychology and culture provides a richer inductive signal than surface-level response distributions. We further demonstrate strong generalization to downstream applications with- out task-specific supervision: outperforming best baselines by 5.19% in personalized reward modeling, 6.34% in population simulation, and showing strong transfer to social reasoning tasks. Datasets and code are available at: https://github.com/limenlp/PALMs.
Priyanka Dey, Brihi Joshi, Preyashi Poddar +2
University of Southern California · Information Sciences Institute
Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produce reasonable approximations of cultural taste remains an open empirical question that becomes more urgent by the day, with market research companies already offering provisional 'synthetic' survey panels and the contamination of standard survey data from LLM-generated responses. In this study, we build on past work on silicon sampling by extending considerations of their ecological, relational, and positional fidelity in the doomain of cultural tastes. We use large-language models from OpenAI, Anthropic, and DeepSeek to produce 554,940 silicon surrogates of survey respondents from the Survey of Public Participation in the Arts (SPPA). We find these silicon surrogates' tastes to be highly stylized facsimiles of human tastes. First, silicon samples are super-omnivorous with a systematic postive-bias for liking. These individual-level bias of silicon samples are not well-explained by the WEIRD-bias often discussed in the literature. Second, the complex relationality in real taste structures is completely distorted among silicon samples. Third, very little of the known cultural alignment between tastes and social space are preserved. Silicon samples juvenilize age-taste associations, resurrect anachronistic class-taste associations, and caricaturize gender- and race-taste associations. Key words: AI, taste, consumption, culture, silicon sampling, meta-analysis.