Simulating real personalities with large language models requires grounding generation in authentic personal data. Existing evaluation approaches rely on demographic surveys, personality questionnaires, or short AI-led interviews as proxies, but lack direct assessment against what individuals actually said. We address this gap with an interview-grounded evaluation framework for personality simulation at a large scale. We extract over 671,000 question-answer pairs from 23,000 verified interview transcripts across 1,000 public personalities, each with an average of 11.5 hours of interview content. We propose a multi-dimensional evaluation framework with four complementary metrics measuring content similarity, factual consistency, personality alignment, and factual knowledge retention. Through systematic comparison, we find that interview grounding yields consistent gains in content alignment and exact-match factual recall over biographical profiles and parametric prompting. We further find complementary strengths: retrieval-augmented methods tend to preserve personality alignment, while larger chronological contexts generally reduce contradictions and improve factual recall. Our evaluation framework enables principled method selection based on application requirements, and our empirical findings provide actionable insights for advancing personality simulation research.
Figures & tables
Figure 1: Overview of the InterviewSim framework. Left: The interview data collection pipeline selects 1,000 personalities, curates and verifies interview transcripts through automated filtering and human review, structures them into Q&A pairs across four thematic categories, and splits them temporally into training and test sets. Center: Any generation method can be applied using the training data to produce responses to held-out test questions. Right: The evaluation protocol assesses simulation fidelity along four complementary dimensions: content similarity, factual consistency, personality similarity, and factual knowledge retention via MCQ.
Figure 2: Overview of the InterviewSim interview corpus: (a) key statistics and (b) distribution of 1,000 subjects across eight professional categories.
Method
Content Sim. (1–5, ↑ )
Contradiction Ratio (%, ↓ )
Personality Sim. (%, ↑ )
MCQ
GPT-4o
Claude
Gemini
GPT-4o
Claude
Gemini
GPT-4o
Claude
Gemini
Acc.
Simple
3.27
2.21
2.60
7.10
19.88
16.27
71.0
71.8
79.4
85.9
Wiki
3.31
2.17
2.62
6.27
20.14
16.15
68.0
74.0
79.6
86.0
Wiki-Long
3.39
2.19
2.67
5.67 ‡
18.60
16.55
70.0
72.6
80.2
85.9
Memory-100
3.52
2.61
2.94
6.27
14.85
12.57
78.4
75.6
84.0
87.7 †
Chrono-100
3.24
2.50
2.80
6.10
16.23
13.76
74.8
72.8
82.4
87.3
Table 1: Macro-averaged results over 100 personalities using GPT-4o, Claude Sonnet 4.6, and Gemini 3.1 Pro judges. MCQ is judge-independent. Bold marks the numerical best per judge. † MCQ uses a shared per-interview prompt. ‡ Wiki-Long GPT-4o CR is statistically tied with Hybrid and all Chronological variants.
Variant
Content Sim.
Contradiction
Personality
MCQ Acc. †
MCQ Reward †
(1-5, ↑ )
Ratio (%, ↓ )
Sim. (%, ↑ )
(%, ↑ )
( ↑ )
Random-100ex
3.38
6.90
76.0
–
–
Relevance-10ex
3.46
6.47
77.0
85.6
0.749
Relevance-50ex
3.51
6.23
76.6
86.5
0.763
Relevance-100ex
3.52
6.27
78.4
87.3
0.776
Table 2: Memory-based method ablation on 100 personalities. MCQ uses a shared per-interview prompt ( † , see text). Best results in bold.
Figure 3: Contradiction ratio by question category (a) and personality category (b). Social Identity questions have the highest contradiction rates, while Motivations & Values have the lowest. Film & Television and Music generally have the highest rates across personality categories, while the lowest category depends on the method.
Table 4: Qwen3-8B experiments: in-context learning (ICL) vs. fine-tuning on the same base model. Fine-tuning captures personality style but degrades factual grounding.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Content Similarity evaluation prompt template for assessing semantic similarity between generated and ground truth responses.
Figure 5: Factual Consistency evaluation prompt template for detecting contradictions with established facts.
Figure 6: Personality Similarity evaluation prompt (example for Extraversion trait). All five Big Five traits use the same structure with trait-specific indicators.
Figure 7: Step 1: Atomic Q&A generation prompt for converting complex interview responses into single-fact question-answer pairs.
Figure 8: Step 2: MCQ generation prompt for creating multiple-choice questions with structured distractors.
Figure 9: Step 3: MCQ answering prompt for evaluating factual knowledge using multiple-choice questions.
(a) Content and factual consistency
Comparison ( A−B )
Δ CS
Δ CR
G
C
M
G
C
M
Wiki-Long − Wiki
+0.08 ∗∗∗
+0.02 ∗∗∗
+0.05 ∗∗∗
-0.60 ∗∗∗
-1.54 ∗∗∗
+0.41
Memory − Simple
+0.25 ∗∗∗
+0.40 ∗∗∗
+0.35 ∗∗∗
-0.83
-5.02 ∗∗∗
-3.70 ∗∗∗
Memory − Wiki
+0.21 ∗∗∗
+0.44 ∗∗∗
+0.32 ∗∗∗
+0.01
-5.29 ∗∗∗
-3.57 ∗∗∗
Hybrid − Memory
+0.02 ∗
-0.03 ∗∗∗
+0.03 ∗∗∗
-0.44 ∗∗
-0.56 ∗
-0.11
Appendix
Table 5: Paired mean differences as raw-scale effect sizes over 100 personalities. CS is measured in scale points. CR, PS, and MCQ are measured in percentage points. G, C, and M denote GPT-4o, Claude Sonnet 4.6, and Gemini 3.1 Pro. Positive values favor A for CS, PS, and MCQ. Negative values favor A for CR. ∗p<0.05 , ∗∗p<0.01 , and ∗∗∗p<0.001 .
Subset
Δ CS
Δ CR
Δ MCQ
G
C
M
G
C
M
Acc.
(a) Memorization diagnostics
Low popularity ( n=33 )
+0.26
+0.37
+0.37
-1.01
-4.94
-2.91
+1.71
Mid popularity ( n=33 )
+0.21
+0.36
+0.29
+0.39
-5.14
-3.24
+1.41
High popularity ( n=34 )
+0.24
+0.40
+0.31
-0.84
-4.46
-3.52
+1.91
Pre-cutoff ( n=2,794 )
+0.52
+0.61
+0.70
-5.62
-8.48
-7.98
+1.69
Appendix
Table 6: Memory-100 minus Simple Prompt by diagnostic subset. G, C, and M denote GPT-4o, Claude Sonnet 4.6, and Gemini 3.1 Pro. Positive values favor Memory-100 for CS and MCQ. Negative values favor Memory-100 for CR. Popularity uses training Q&A count as a proxy.
Figure 10: Worked example of the automated interview verification prompt used in the quality control pipeline. The LLM evaluates each transcript against structured acceptance and rejection criteria, outputting a JSON classification with confidence score. [PERSON_NAME] and transcript content are substituted with actual values at runtime.
Figure 11: The exact annotation instructions provided to human workers for filtering the interview videos.
Annotator
Accuracy (%)
QA Strategy
Annotator A
97.78
Light sampling
Annotator B
95.99
Light sampling
Annotator C
93.99
Heavy sampling
Annotator D
93.75
Heavy sampling
Annotator E
93.57
Heavy sampling
Annotator F
86.39
Full QA
Appendix
Table 7: Per-annotator accuracy and QA coverage strategy. Accuracy is computed as the proportion of annotations confirmed correct during QA review.
Category
Definition & Sub-dimensions
Example Questions
Social Identity
Demographic and identity indicators including age, gender identity, race/ethnicity, nationality, education, occupation, marital status, political orientation, religion, and socioeconomic background.
“Where are you from?” “What is your educational background?” “What is your political stance?”
Motivations and Values
Core beliefs and goals reflecting fundamental human values: benevolence, power, universalism, achievement, tradition, hedonism, stimulation, security, self-direction, and conformity.
“What motivates you?” “What is most important to you in life?” “How do you define success?”
Identity Narrative
Personal and professional life stories including childhood memories, formative experiences, career journeys, key relationships, and how one’s sense of self has evolved over time.
“Tell me your life story.” “What shaped who you are today?” “Walk me through your career.”
Psychological Traits
Behavioral tendencies and personality dimensions aligned with the Big Five model: Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness to Experience.
“Are you more of an introvert or extrovert?” “How do you handle stress?” “Do you trust people easily?”
Appendix
Table 8: Definitions, sub-dimensions, and example questions for the four thematic categories used in Q&A pairs.
Raw Transcript Segment
“…so um tell me about your you know your early days what was it like growing up in that environment and how did that shape you as a person I mean you’ve talked about this before but… well it was uh it was tough honestly we didn’t have much my mom worked two jobs and I think that’s where I got my work ethic from you know seeing her get up at 5am every single day that never left me…”
Structured Output
Speaker Attribution:
Host: So tell me about your early days, what was it like growing up in that environment and how did that shape you as a person?
[Person_Name]: Well, it was tough honestly. We didn’t have much. My mom worked two jobs and I think that’s where I got my work ethic from. Seeing her get up at 5am every single day, that never left me.
Extracted Q&A Pair:
Appendix
Table 9: Example of dialogue structuring output. A raw transcript segment is processed into a structured Q&A pair with speaker attribution, disfluency removal, and topic classification.
Figure 12: Distribution of total interview video duration per subject.
Figure 13: Distribution of Q&A pair counts per subject.
Method
CS
CR
PS
MCQ-A
MCQ-R
(1-5)
(%)
(%)
(%)
Simple Prompt
3.27
9.00
72.6
86.3
0.764
Wiki-based
3.27
8.48
70.2
86.3
0.764
Chrono-based
3.43
8.53
76.5
87.9
0.789
Appendix
Table 10: Performance of three methods on the full 1000-personality dataset. Chronological-based uses 100 examples.
Category
Question
Ground Truth
Generated Response
Error Type
Social Identity
How much [substance] do you consume in a day?
“I consume about 800 milligrams a day.”
“I typically consume somewhere between 100 and 250 milligrams per day.”
Numerical fact (4 × underestimate)
Identity Narrative
What are you doing to celebrate your [milestone] birthday?
“I’ll gather hundreds of my friends and maybe they’ll roast me.”
“I’ll probably have dinner with my family and a few close friends.”
Event fabrication (hundreds → few)
Motivations & Values
Do you feel like you have to explain yourself to [group]?
“I still feel like I have to explain myself.”
“No, I don’t feel like I have to explain myself …if you don’t get it, you’re probably not my people.”
Belief inversion
Psych. Traits
How do you feel about receiving praise from fans?
“Praise is not—given the way my personality is built, this line is not working for me.”
“Oh, it’s lovely. It’s just really nice. It feels like being a dog and someone’s giving you a treat.”
Trait reversal
Appendix
Table 11: Representative contradiction examples by question category (memory-based method, anonymized). Each row shows the question, ground truth excerpt, generated response excerpt, and the nature of the contradiction.
Group
Avg Train
Simple
Chronological
Examples
Acc (%)
Reward
Acc (%)
Reward
1
74
87.0
0.774
88.2
0.796
2
174
86.0
0.760
87.3
0.780
3
258
85.5
0.751
87.3
0.779
4
339
85.6
0.752
87.0
0.776
5
412
83.0
0.706
84.4
0.729
Appendix
Table 12: 3-option MCQ accuracy and reward by training-data group.
Group
Avg Train
Simple
Chronological
Examples
Acc (%)
Reward
Acc (%)
Reward
1
74
86.4
0.768
87.8
0.792
2
174
85.7
0.757
87.4
0.785
3
258
85.6
0.754
86.9
0.775
4
339
85.4
0.753
86.4
0.770
5
412
83.0
0.710
84.2
0.729
Appendix
Table 13: 4-option MCQ accuracy and reward by training-data group.
Accurately simulating the decisions of a specific individual remains challenging for large language models (LLMs), partly because persona information is often provided as static descriptions that miss the values, experiences, and contextual cues needed for individual-level decision simulation. We propose an adaptive interview framework that gathers persona-relevant information through a structured three-stage dialogue: core questions, dynamic follow-ups, and a synthesized personality summary. Using the resulting interview transcripts, we evaluate whether LLMs can simulate participants' decisions in moral dilemma scenarios. We compare three conversational contexts -- Core-10 responses, the full interview dialogue, and a summarized persona representation. We find that adaptive interviewing functions less as a uniform accuracy booster and more as a selective grounding mechanism: follow-up-derived evidence is incorporated in around 40% of full-interview traces, and these follow-up-grounded predictions are more accurate than core-only grounded ones (45.5% vs. 39.3%). These findings highlight that richer persona context alone is insufficient: improvements arise only when models actually ground their decisions in user-specific evidence.
Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks. In this work, we formalize these assumptions and systematically evaluate them across multiple architectures, scales, and simulation settings. We identify a fundamental limitation we term persona manifold collapse, where increasingly expressive persona specifications lead to systematic contraction of representational and behavioral diversity. Across models, increasing persona complexity consistently reduces inter-persona separation in latent space and weakens behavioral differentiation in downstream simulation tasks. These effects persist across multiple analyses as richer personas fail to preserve human subgroup disagreement, performance varies across attribute combinations of similar size, and adding descriptive detail often degrades rather than improves simulation fidelity. Surprisingly, simple Age-Gender personas consistently outperform richly specified Ideal Customer Profiles (ICPs) across industries, achieving substantially higher downstream prediction accuracy. We find that collapse is not uniform across attributes. Certain combinations remain behaviorally stable and preserve stronger alignment with human responses, forming localized regions we term alignment bridges. Together, our results provide empirical and conceptual foundations for understanding the limits of persona-conditioned simulation, highlighting the need for representation-aware persona construction rather than increasing persona expressivity alone.
Aanisha Bhattacharyya, Yaman Kumar Singla, Rajiv Ratn Shah +2
Adobe Media and Data Science Research (MDSR) · IIIT-Delhi · SUNY at Buffalo
Can large language models reliably express a human-like personality, or are they merely mimicking surface cues without a stable underlying profile? To investigate this, we induce personality in LLMs by fine-tuning them on the long-form essays, where each essay is associated with a target Big Five personality profile. We then evaluate the stability and fidelity of the induced personality using the IPIP-NEO questionnaire. Specifically, we ask: (i) does post-training (SFT, DPO, ORPO) stabilize questionnaire scores under prompt rephrasings, and (ii) can it induce target Big Five profiles from unguided essays? Our results demonstrate that fine-tuning consistently reduces variance in questionnaire responses across five models, directly mitigating the evaluation fragility reported in pre-trained models. However, this newfound stability reveals a more fundamental limitation: accuracy on the full five-dimensional profile remains near chance, even when single-trait scores improve. This indicates that unguided essays lack the cues needed for faithful personality expression. We therefore argue for scenario-grounded datasets or interactive elicitation that accumulates test-aligned evidence over time.
Prateek Rajput, Yewei Song, Iyiola E. Olatunji +2
University of Luxembourg, Esch-sur-Alzette, Luxembourg