We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces stereotyped predictions (lacking natural diversity). It has previously been shown that LLMs can be bound to personas using naturalistic, freetext dialog avoiding stereotyping. Here we show that binding can also be achieved using short, individual samples of dialog from specific people. Demographics can be added later without negative effects by simply querying the model. We use the term Persona Mixture Models (PMMs) for well-calibrated human models, currently realized as pretrained base models. We show that PMMs produce more accurate predictions than instruction-tuned models and retain more of the lexical, semantic, and pragmatic diversity found in human dialog. We measure realism and diversity of LLMs simulating human interlocutors across a diverse set of corpora spanning open-domain text, human-AI chat, and task-oriented dialogue between human speakers. However, base pretrained models can produce out-of-domain dialog and may lose some of the human's internal state over long contexts. We propose and explore tandem models which combine a pre-trained model with an instruction-tuned supervisor. Tandem models achieve the best overall accuracy and diversity in our experiments.
Figures & tables
Corpus
Dialogue type
Sampled Contexts
Total
Reddit (ConvoKit) [ 67 ]
Open-domain; Human–Human
1,024
8,192
WildChat-1M [ 68 ]
Human-LLM Chat; Human–AI
1,024
8,192
LMSYS-Chat-1M [ 69 ]
Human-LLM Chat; Human–AI
1,024
8,192
MultiWOZ 2.1 [ 70 ]
Task-Oriented; Human–Human
807
6,456
DialOp [ 8 ]
Task-Oriented; Human–Human
306
2,448
Table 1 : Conversational corpora used in our case study. For all corpora, we use similar next-utterance prediction setups with LMs and sample 8 generations per context. For Multiwoz and DialOp, we take 128 and 117 dialogues, respectively, and generate utterances of all human-user turns; for all other datasets, we take the final user turn and generate model utterances in place of the human interlocutor.
Figure 1 : Per-token self-entropy and negative log-likelihood. For each model, we report model self-entropy (shown as solid bars) and negative log-likelihood of human utterances given dialogue context. See Appendix E for full results.
Figure 2 : MAUVE results. MAUVE is a sample-based score in [0,1] that quantifies how close two text distributions are. The score is high when the set of model samples and the set of human references cover the same regions of the distribution under a text embedding model; conversely, score is low when either distribution places mass where the other does not. A value near 1.0 indicates that the two samples are statistically indistinguishable, while lower values mean the model is sampling from a different region of utterance space than humans.
JSD ( ↓ )
V Overlap ( ↑ )
Model
N=1
2
3
N=1
2
3
Tandem
.088
.219
.280
0.08
.75
.84
M24B-B
.181
.277
.330
0.08
.75
.74
M24B-I
.195
.299
.346
0.08
.68
.67
Q2.5-14B-B
.175
.272
.325
0.08
.68
.76
Q2.5-14B-I
.119
.273
.321
0.08
.68
.75
Table 2: Full JSD and Vocabulary Overlap results ( K=100 ). B, I, and R respectively indicate Base, Instruct, and Reasoning model variants.
Figure 3 : Jensen–Shannon divergence (JSD) between human and model dialogue-act N -gram distributions on MultiWOZ-2.1. For each N∈{1,…,3} , we form the empirical distribution of speaker-marked dialogue-act N -grams from the aligned human reference ( PN ) and from the model under evaluation ( QN ), then plot JSD(PN,QN) . JSD is bounded in [0,1] ; lower is better , with 0 corresponding to identical distributions. Higher N probes longer pragmatic context—a model that matches marginal act frequencies (low JSD at N=1 ) can still fail to reproduce the ordering of acts across turns (high JSD at N≥3 ).
Figure 4 : Utterance-level diversity diagnostics on Reddit. For each model we report three complementary measures of how varied the model-generated continuations are within a given dialogue context, averaged across contexts. Self-BLEU (left) computes BLEU between every pair of continuations from the same context: lower is better , since a low value means the continuations are lexically distinct from one another. Self-BERT-F1 (center) is the analogous quantity in contextual-embedding space, capturing semantic rather than surface variation: again, lower indicates greater diversity . Vendi score (right) reports the effective number of semantically distinct samples, computed from the eigenspectrum of the embedding similarity kernel [ 78 ] : higher is better , with a value near 1 corresponding to total mode collapse.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
13.2424
19.2998
0.6743
0.9602
0.0122
0.0776
0.4796
6.3915
0.9650
Llama-3.1-8B-Instruct
23.1513
6.9341
0.5984
0.9339
0.0545
0.1304
0.5337
4.7601
0.6946
Mistral-24B-Base
10.7243
17.2362
0.6562
0.9581
0.0152
0.0845
0.4884
6.2010
0.9658
Mistral-24B-Instruct
13.4477
10.4384
0.6147
0.9263
0.0790
0.1473
0.5560
5.0822
0.8999
Qwen2.5-14B
12.8851
16.4860
0.6770
0.9593
0.0109
0.0799
0.4831
6.4020
0.9666
Qwen2.5-14B-Instruct
32.1536
2.7724
0.5726
0.8520
0.2017
0.2176
0.6285
4.1150
0.7728
Appendix
Table 3 : Full Reddit point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
4.3322
8.1718
0.6483
0.8977
0.0785
0.1122
0.5141
6.1210
0.8608
Llama-3.1-8B-Instruct
5.2963
9.5169
0.6068
0.9044
0.0865
0.1510
0.5840
4.7968
0.7813
Mistral-24B-Base
3.5866
8.9617
0.6321
0.8875
0.0902
0.1170
0.5166
5.9821
0.9659
Mistral-24B-Instruct
4.0625
7.0132
0.6758
0.9275
0.0598
0.1315
0.5812
5.5360
0.6233
Qwen2.5-14B
3.8005
7.4436
0.6463
0.8871
0.0922
0.1114
0.5102
6.0735
0.9691
Qwen2.5-14B-Instruct
5.8571
2.4996
0.5625
0.7978
0.2501
0.2386
0.6641
4.0393
0.5007
Appendix
Table 4 : Full WildChat-1M point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
4.0134
7.1166
0.6723
0.8993
0.0660
0.1117
0.5177
6.3393
0.8206
Llama-3.1-8B-Instruct
5.8418
3.5852
0.5955
0.8857
0.1070
0.1662
0.5911
4.8261
0.7639
Mistral-24B-Base
3.3318
7.3487
0.6714
0.8959
0.0746
0.1162
0.5185
6.2141
0.9503
Mistral-24B-Instruct
4.3267
4.9720
0.6673
0.9010
0.0918
0.1625
0.5977
5.3387
0.7026
Qwen2.5-14B
3.4565
6.1836
0.6676
0.8830
0.0837
0.1194
0.5196
6.1912
0.9499
Qwen2.5-14B-Instruct
12.5903
2.1503
0.5200
0.7306
0.3556
0.3049
0.6864
3.6589
0.6036
Appendix
Table 5 : Full LMSYS-Chat-1M point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
5.5176
7.9276
0.7019
0.9066
0.0809
0.1862
0.5428
5.1383
0.6172
Llama-3.1-8B-Instruct
10.7274
2.8377
0.5258
0.8028
0.2477
0.2610
0.6700
3.6958
0.3526
Mistral-24B-Base
4.8951
5.8217
0.6910
0.8948
0.0956
0.1845
0.5804
5.0722
0.6743
Mistral-24B-Instruct
5.0876
3.6643
0.5731
0.8003
0.2424
0.2851
0.6932
3.8587
0.6687
Qwen2.5-14B
5.1235
6.3018
0.6654
0.8824
0.1153
0.2005
0.6367
4.9447
0.7132
Qwen2.5-14B-Instruct
20.0783
1.7114
0.4251
0.6084
0.5541
0.4712
0.6949
2.5738
0.7499
Appendix
Table 6 : Full MultiWOZ continuation point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
12.4540
16.5786
0.7327
0.9570
0.0125
0.0829
0.5614
6.5064
0.4638
Llama-3.1-8B-Instruct
23.6024
4.4736
0.5588
0.8957
0.0868
0.1711
0.6723
4.1654
0.3901
Mistral-24B-Base
10.9471
12.0142
0.7194
0.9522
0.0180
0.0821
0.5996
6.4174
0.5942
Mistral-24B-Instruct
12.3744
9.9723
0.5982
0.9307
0.0434
0.1367
0.6078
4.8471
0.4900
Qwen2.5-14B
12.3213
8.0706
0.7019
0.9454
0.0245
0.0911
0.5743
6.2673
0.5640
Qwen2.5-14B-Instruct
57.3444
2.4584
0.5177
0.7909
0.2729
0.2582
0.6922
3.4711
0.3958
Appendix
Table 7 : Full DialOp point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
Figure 5 : Individual (N=1) dialogue act distributions for each model compared against the human distribution.
JSD ( ↓ )
Overlap K=10 ( ↑ )
Overlap K=100 ( ↑ )
Overlap K=200 ( ↑ )
System
N=1
2
3
N=1
2
3
N=1
2
3
N=1
2
3
Llama-3.1-8B-Base
0.198
0.293
0.346
0.700
0.600
0.600
0.080
0.700
0.770
0.040
0.600
0.645
Llama-3.1-8B-Inst
0.265
0.366
0.407
0.400
0.700
0.700
0.070
0.640
0.640
0.035
0.495
0.570
Mistral-24B-Base
0.181
0.277
0.330
0.700
0.800
0.500
0.080
0.750
0.740
0.040
0.620
0.660
Mistral-24B-Inst
0.195
0.299
0.346
0.700
0.700
0.800
0.080
0.680
0.670
0.040
0.575
0.635
Qwen2.5-7B-Base
0.168
0.271
0.325
0.700
0.800
0.600
0.080
0.760
0.780
0.040
0.635
0.660
Appendix
Table 8 : Comprehensive results for N-gram Dialogue Act sequences ( N=1 to 3 ). Results include Jensen-Shannon Divergence (JSD ↓ ) and Vocabulary Overlap ( ↑ ) at multiple K thresholds.
Figure 6 : Transition probabilities for N=2 (i.e., probability of user intent given previous assistant intent).
Figure 7 : Comparison of distributional divergence and vocabulary overlap between model outputs and human-human (H-H) bootstrapped baselines. H-H results are generated via 100 × random corpus splits. We observe that for N≤3 , H-H performance consistently exceeds model-human similarity, establishing N=3 as our threshold complexity limit.
Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks. In this work, we formalize these assumptions and systematically evaluate them across multiple architectures, scales, and simulation settings. We identify a fundamental limitation we term persona manifold collapse, where increasingly expressive persona specifications lead to systematic contraction of representational and behavioral diversity. Across models, increasing persona complexity consistently reduces inter-persona separation in latent space and weakens behavioral differentiation in downstream simulation tasks. These effects persist across multiple analyses as richer personas fail to preserve human subgroup disagreement, performance varies across attribute combinations of similar size, and adding descriptive detail often degrades rather than improves simulation fidelity. Surprisingly, simple Age-Gender personas consistently outperform richly specified Ideal Customer Profiles (ICPs) across industries, achieving substantially higher downstream prediction accuracy. We find that collapse is not uniform across attributes. Certain combinations remain behaviorally stable and preserve stronger alignment with human responses, forming localized regions we term alignment bridges. Together, our results provide empirical and conceptual foundations for understanding the limits of persona-conditioned simulation, highlighting the need for representation-aware persona construction rather than increasing persona expressivity alone.
Aanisha Bhattacharyya, Yaman Kumar Singla, Rajiv Ratn Shah +2
Adobe Media and Data Science Research (MDSR) · IIIT-Delhi · SUNY at Buffalo
Large Language Model (LLM) agents are increasingly deployed in settings where they interact with diverse users, including those who are unclear, impatient, or reluctant to share information. However, collecting real interaction data at scale remains expensive. The field has turned to LLM-based \emph{user simulators} as stand-ins, but these simulators inherit the behavior of their underlying models: cooperative and homogeneous. As a result, agents that appear strong in simulation often fail in real human interactions. To narrow this gap, we introduce Persona Policies (PPol), a plug-and-play control layer that induces realistic behavioral variation in user simulators while preserving original task goals. Rather than hand-crafting personas, we employ an evolutionary coding agent to discover persona generation programs optimized for human-likeness and behavioral coverage over real user conversations. The evolved program generates diverse, human-like personas for any task in the domain. Across 4 benchmarks--including τ2-bench Retail and Airline, ColBench, and WildChat--evolved PPol yield 28-72% absolute gains in fitness score over the baseline simulator. In blinded evaluations, annotators judged PPol users as 'human' 80.4% of the time, nearly 2x more than the baseline simulators. Training agents with PPol also improves real-world performance: our user study with live human-agent interactions showed that fine-tuning with our method boosted task success by +23% over default baselines. PPol thus offers a novel approach to strengthen simulator-based evaluation and training without changing underlying tasks.
Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term \emph{Persona Collapse}: agents each assigned a distinct profile nonetheless converge into a narrow behavioral mode, producing a homogeneous simulated population. To quantify persona collapse, we propose a framework that measures how much of the persona space a population occupies (Coverage), how evenly agents spread across it (Uniformity), and how rich the resulting behavioral patterns are (Complexity). Evaluating ten LLMs on personality simulation (BFI-44), moral reasoning, and self-introduction, we observe persona collapse along two axes: (1) Dimensions: a model can appear diverse on one axis yet structurally degenerate on another, and (2) Domains: the same model may collapse the most in personality yet be the most diverse in moral reasoning. Furthermore, item-level diagnostics reveal that behavioral variation tracks coarse demographic stereotypes rather than the fine-grained individual differences specified in each persona. Counter-intuitively, \textbf{the models achieving the highest per-persona fidelity consistently produce the most stereotyped populations}. We release our toolkit and data to support population-level evaluation of LLMs.