We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces stereotyped predictions (lacking natural diversity). It has previously been shown that LLMs can be bound to personas using naturalistic, freetext dialog avoiding stereotyping. Here we show that binding can also be achieved using short, individual samples of dialog from specific people. Demographics can be added later without negative effects by simply querying the model. We use the term Persona Mixture Models (PMMs) for well-calibrated human models, currently realized as pretrained base models. We show that PMMs produce more accurate predictions than instruction-tuned models and retain more of the lexical, semantic, and pragmatic diversity found in human dialog. We measure realism and diversity of LLMs simulating human interlocutors across a diverse set of corpora spanning open-domain text, human-AI chat, and task-oriented dialogue between human speakers. However, base pretrained models can produce out-of-domain dialog and may lose some of the human's internal state over long contexts. We propose and explore tandem models which combine a pre-trained model with an instruction-tuned supervisor. Tandem models achieve the best overall accuracy and diversity in our experiments.
Figures & tables
Corpus
Dialogue type
Sampled Contexts
Total
Reddit (ConvoKit) [ 67 ]
Open-domain; Human–Human
1,024
8,192
WildChat-1M [ 68 ]
Human-LLM Chat; Human–AI
1,024
8,192
LMSYS-Chat-1M [ 69 ]
Human-LLM Chat; Human–AI
1,024
8,192
MultiWOZ 2.1 [ 70 ]
Task-Oriented; Human–Human
807
6,456
DialOp [ 8 ]
Task-Oriented; Human–Human
306
2,448
Table 1 : Conversational corpora used in our case study. For all corpora, we use similar next-utterance prediction setups with LMs and sample 8 generations per context. For Multiwoz and DialOp, we take 128 and 117 dialogues, respectively, and generate utterances of all human-user turns; for all other datasets, we take the final user turn and generate model utterances in place of the human interlocutor.
Figure 1 : Per-token self-entropy and negative log-likelihood. For each model, we report model self-entropy (shown as solid bars) and negative log-likelihood of human utterances given dialogue context. See Appendix E for full results.
Figure 2 : MAUVE results. MAUVE is a sample-based score in [0,1] that quantifies how close two text distributions are. The score is high when the set of model samples and the set of human references cover the same regions of the distribution under a text embedding model; conversely, score is low when either distribution places mass where the other does not. A value near 1.0 indicates that the two samples are statistically indistinguishable, while lower values mean the model is sampling from a different region of utterance space than humans.
JSD ( ↓ )
V Overlap ( ↑ )
Model
N=1
2
3
N=1
2
3
Tandem
.088
.219
.280
0.08
.75
.84
M24B-B
.181
.277
.330
0.08
.75
.74
M24B-I
.195
.299
.346
0.08
.68
.67
Q2.5-14B-B
.175
.272
.325
0.08
.68
.76
Q2.5-14B-I
.119
.273
.321
0.08
.68
.75
Table 2: Full JSD and Vocabulary Overlap results ( K=100 ). B, I, and R respectively indicate Base, Instruct, and Reasoning model variants.
Figure 3 : Jensen–Shannon divergence (JSD) between human and model dialogue-act N -gram distributions on MultiWOZ-2.1. For each N∈{1,…,3} , we form the empirical distribution of speaker-marked dialogue-act N -grams from the aligned human reference ( PN ) and from the model under evaluation ( QN ), then plot JSD(PN,QN) . JSD is bounded in [0,1] ; lower is better , with 0 corresponding to identical distributions. Higher N probes longer pragmatic context—a model that matches marginal act frequencies (low JSD at N=1 ) can still fail to reproduce the ordering of acts across turns (high JSD at N≥3 ).
Figure 4 : Utterance-level diversity diagnostics on Reddit. For each model we report three complementary measures of how varied the model-generated continuations are within a given dialogue context, averaged across contexts. Self-BLEU (left) computes BLEU between every pair of continuations from the same context: lower is better , since a low value means the continuations are lexically distinct from one another. Self-BERT-F1 (center) is the analogous quantity in contextual-embedding space, capturing semantic rather than surface variation: again, lower indicates greater diversity . Vendi score (right) reports the effective number of semantically distinct samples, computed from the eigenspectrum of the embedding similarity kernel [ 78 ] : higher is better , with a value near 1 corresponding to total mode collapse.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
13.2424
19.2998
0.6743
0.9602
0.0122
0.0776
0.4796
6.3915
0.9650
Llama-3.1-8B-Instruct
23.1513
6.9341
0.5984
0.9339
0.0545
0.1304
0.5337
4.7601
0.6946
Mistral-24B-Base
10.7243
17.2362
0.6562
0.9581
0.0152
0.0845
0.4884
6.2010
0.9658
Mistral-24B-Instruct
13.4477
10.4384
0.6147
0.9263
0.0790
0.1473
0.5560
5.0822
0.8999
Qwen2.5-14B
12.8851
16.4860
0.6770
0.9593
0.0109
0.0799
0.4831
6.4020
0.9666
Qwen2.5-14B-Instruct
32.1536
2.7724
0.5726
0.8520
0.2017
0.2176
0.6285
4.1150
0.7728
Appendix
Table 3 : Full Reddit point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
4.3322
8.1718
0.6483
0.8977
0.0785
0.1122
0.5141
6.1210
0.8608
Llama-3.1-8B-Instruct
5.2963
9.5169
0.6068
0.9044
0.0865
0.1510
0.5840
4.7968
0.7813
Mistral-24B-Base
3.5866
8.9617
0.6321
0.8875
0.0902
0.1170
0.5166
5.9821
0.9659
Mistral-24B-Instruct
4.0625
7.0132
0.6758
0.9275
0.0598
0.1315
0.5812
5.5360
0.6233
Qwen2.5-14B
3.8005
7.4436
0.6463
0.8871
0.0922
0.1114
0.5102
6.0735
0.9691
Qwen2.5-14B-Instruct
5.8571
2.4996
0.5625
0.7978
0.2501
0.2386
0.6641
4.0393
0.5007
Appendix
Table 4 : Full WildChat-1M point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
4.0134
7.1166
0.6723
0.8993
0.0660
0.1117
0.5177
6.3393
0.8206
Llama-3.1-8B-Instruct
5.8418
3.5852
0.5955
0.8857
0.1070
0.1662
0.5911
4.8261
0.7639
Mistral-24B-Base
3.3318
7.3487
0.6714
0.8959
0.0746
0.1162
0.5185
6.2141
0.9503
Mistral-24B-Instruct
4.3267
4.9720
0.6673
0.9010
0.0918
0.1625
0.5977
5.3387
0.7026
Qwen2.5-14B
3.4565
6.1836
0.6676
0.8830
0.0837
0.1194
0.5196
6.1912
0.9499
Qwen2.5-14B-Instruct
12.5903
2.1503
0.5200
0.7306
0.3556
0.3049
0.6864
3.6589
0.6036
Appendix
Table 5 : Full LMSYS-Chat-1M point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
5.5176
7.9276
0.7019
0.9066
0.0809
0.1862
0.5428
5.1383
0.6172
Llama-3.1-8B-Instruct
10.7274
2.8377
0.5258
0.8028
0.2477
0.2610
0.6700
3.6958
0.3526
Mistral-24B-Base
4.8951
5.8217
0.6910
0.8948
0.0956
0.1845
0.5804
5.0722
0.6743
Mistral-24B-Instruct
5.0876
3.6643
0.5731
0.8003
0.2424
0.2851
0.6932
3.8587
0.6687
Qwen2.5-14B
5.1235
6.3018
0.6654
0.8824
0.1153
0.2005
0.6367
4.9447
0.7132
Qwen2.5-14B-Instruct
20.0783
1.7114
0.4251
0.6084
0.5541
0.4712
0.6949
2.5738
0.7499
Appendix
Table 6 : Full MultiWOZ continuation point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
System
PPL-H
PPL-G rand
Dist-1
Dist-2
Self-BLEU
Self-ROUGE-L
Self-BERT-F1
Vendi
MAUVE
Llama-3.1-8B
12.4540
16.5786
0.7327
0.9570
0.0125
0.0829
0.5614
6.5064
0.4638
Llama-3.1-8B-Instruct
23.6024
4.4736
0.5588
0.8957
0.0868
0.1711
0.6723
4.1654
0.3901
Mistral-24B-Base
10.9471
12.0142
0.7194
0.9522
0.0180
0.0821
0.5996
6.4174
0.5942
Mistral-24B-Instruct
12.3744
9.9723
0.5982
0.9307
0.0434
0.1367
0.6078
4.8471
0.4900
Qwen2.5-14B
12.3213
8.0706
0.7019
0.9454
0.0245
0.0911
0.5743
6.2673
0.5640
Qwen2.5-14B-Instruct
57.3444
2.4584
0.5177
0.7909
0.2729
0.2582
0.6922
3.4711
0.3958
Appendix
Table 7 : Full DialOp point-estimate results. PPL-H scores aligned human continuations; PPL-G rand scores one randomly sampled generated continuation per context.
Figure 5 : Individual (N=1) dialogue act distributions for each model compared against the human distribution.
JSD ( ↓ )
Overlap K=10 ( ↑ )
Overlap K=100 ( ↑ )
Overlap K=200 ( ↑ )
System
N=1
2
3
N=1
2
3
N=1
2
3
N=1
2
3
Llama-3.1-8B-Base
0.198
0.293
0.346
0.700
0.600
0.600
0.080
0.700
0.770
0.040
0.600
0.645
Llama-3.1-8B-Inst
0.265
0.366
0.407
0.400
0.700
0.700
0.070
0.640
0.640
0.035
0.495
0.570
Mistral-24B-Base
0.181
0.277
0.330
0.700
0.800
0.500
0.080
0.750
0.740
0.040
0.620
0.660
Mistral-24B-Inst
0.195
0.299
0.346
0.700
0.700
0.800
0.080
0.680
0.670
0.040
0.575
0.635
Qwen2.5-7B-Base
0.168
0.271
0.325
0.700
0.800
0.600
0.080
0.760
0.780
0.040
0.635
0.660
Appendix
Table 8 : Comprehensive results for N-gram Dialogue Act sequences ( N=1 to 3 ). Results include Jensen-Shannon Divergence (JSD ↓ ) and Vocabulary Overlap ( ↑ ) at multiple K thresholds.
Figure 6 : Transition probabilities for N=2 (i.e., probability of user intent given previous assistant intent).
Figure 7 : Comparison of distributional divergence and vocabulary overlap between model outputs and human-human (H-H) bootstrapped baselines. H-H results are generated via 100 × random corpus splits. We observe that for N≤3 , H-H performance consistently exceeds model-human similarity, establishing N=3 as our threshold complexity limit.