cs.CLSep 28, 2026

Population Fidelity: Evaluating Population Representativeness in LLMs

Authors: Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva

Organizations: University of Toronto, Canada · Federal University of Technology – Paran´a, Brazil

Abstract

Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonstrate the framework's utility in two ways. First, we reproduce a prior study of "machine bias" in LLM survey responses and apply the framework to its models and more recent ones, showing that poor representation reflects not only insufficient between-group variation but also variation assigned to the wrong groups. Second, we evaluate one proposed approach to improving models' population representativeness: cultural fine-tuning. We find that cultural fine-tuning can improve alignment with the survey center without improving the representation of within-population differences, a distinction that measures of aggregate agreement do not capture. We argue that representing a population requires models to reproduce several features of human attitudinal variation simultaneously. Our framework organizes these features and provides reusable code, data, and trained models for evaluating population fidelity across substantive domains and assessing proposed alignment methods.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

    Jun 27, 2026Eun Cheol Choi, Youngrae Kim, Prabhu Pugalenthi +2SurveyBiases

  2. PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs

    Aug 2, 2026Priyanka Dey, Brihi Joshi, Preyashi Poddar +2Large Language Model AlignmentSocial Reasoning

  3. Tastes without distinction: silicon samples and the synthetic construction of tastes

    Jun 29, 2026Xiangyu Ma, Mengmi Zhang, Shannon Ang +1Large Language Model BiasSurvey