InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
Organizations: Salesforce Research
Abstract
Simulating real personalities with large language models requires grounding generation in authentic personal data. Existing evaluation approaches rely on demographic surveys, personality questionnaires, or short AI-led interviews as proxies, but lack direct assessment against what individuals actually said. We address this gap with an interview-grounded evaluation framework for personality simulation at a large scale. We extract over 671,000 question-answer pairs from 23,000 verified interview transcripts across 1,000 public personalities, each with an average of 11.5 hours of interview content. We propose a multi-dimensional evaluation framework with four complementary metrics measuring content similarity, factual consistency, personality alignment, and factual knowledge retention. Through systematic comparison, we find that interview grounding yields consistent gains in content alignment and exact-match factual recall over biographical profiles and parametric prompting. We further find complementary strengths: retrieval-augmented methods tend to preserve personality alignment, while larger chronological contexts generally reduce contradictions and improve factual recall. Our evaluation framework enables principled method selection based on application requirements, and our empirical findings provide actionable insights for advancing personality simulation research.
Figures & tables
| Method | Content Sim. (1–5, ) | Contradiction Ratio (%, ) | Personality Sim. (%, ) | MCQ | ||||||
| GPT-4o | Claude | Gemini | GPT-4o | Claude | Gemini | GPT-4o | Claude | Gemini | Acc. | |
| Simple | 3.27 | 2.21 | 2.60 | 7.10 | 19.88 | 16.27 | 71.0 | 71.8 | 79.4 | 85.9 |
| Wiki | 3.31 | 2.17 | 2.62 | 6.27 | 20.14 | 16.15 | 68.0 | 74.0 | 79.6 | 86.0 |
| Wiki-Long | 3.39 | 2.19 | 2.67 | 5.67 ‡ | 18.60 | 16.55 | 70.0 | 72.6 | 80.2 | 85.9 |
| Memory-100 | 3.52 | 2.61 | 2.94 | 6.27 | 14.85 | 12.57 | 78.4 | 75.6 | 84.0 | 87.7 † |
| Chrono-100 | 3.24 | 2.50 | 2.80 | 6.10 | 16.23 | 13.76 | 74.8 | 72.8 | 82.4 | 87.3 |
| Variant | Content Sim. | Contradiction | Personality | MCQ Acc. † | MCQ Reward † |
|---|---|---|---|---|---|
| (1-5, ) | Ratio (%, ) | Sim. (%, ) | (%, ) | ( ) | |
| Random-100ex | 3.38 | 6.90 | 76.0 | – | – |
| Relevance-10ex | 3.46 | 6.47 | 77.0 | 85.6 | 0.749 |
| Relevance-50ex | 3.51 | 6.23 | 76.6 | 86.5 | 0.763 |
| Relevance-100ex | 3.52 | 6.27 | 78.4 | 87.3 | 0.776 |
| Method | Correct | Opposite | Near-Miss |
|---|---|---|---|
| Simple Prompt | 85.4% | 6.5% | 8.1% |
| Chrono-based | 86.9% | 6.2% | 6.9% |
| Method | Content Sim. | Contradiction | Personality | MCQ Acc. | MCQ Reward |
|---|---|---|---|---|---|
| (1-5, ) | Ratio (%, ) | Sim. (%, ) | (%, ) | ( ) | |
| Chrono-based ICL (100 ex) | 2.73 | 13.8 | 74.8 | 75.4 | 0.575 |
| Base (simple prompt) | 2.67 | 14.7 | 74.8 | 74.4 | 0.556 |
| Per-personality LoRA | 2.44 | 16.6 | 77.8 | 79.0 | 0.637 |
| All-personality SFT | 2.33 | 18.2 | 61.4 | 76.3 | 0.594 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| (a) Content and factual consistency | ||||||
|---|---|---|---|---|---|---|
| Comparison ( ) | CS | CR | ||||
| G | C | M | G | C | M | |
| Wiki-Long Wiki | +0.08 ∗∗∗ | +0.02 ∗∗∗ | +0.05 ∗∗∗ | -0.60 ∗∗∗ | -1.54 ∗∗∗ | +0.41 |
| Memory Simple | +0.25 ∗∗∗ | +0.40 ∗∗∗ | +0.35 ∗∗∗ | -0.83 | -5.02 ∗∗∗ | -3.70 ∗∗∗ |
| Memory Wiki | +0.21 ∗∗∗ | +0.44 ∗∗∗ | +0.32 ∗∗∗ | +0.01 | -5.29 ∗∗∗ | -3.57 ∗∗∗ |
| Hybrid Memory | +0.02 ∗ | -0.03 ∗∗∗ | +0.03 ∗∗∗ | -0.44 ∗∗ | -0.56 ∗ | -0.11 |
| Subset | CS | CR | MCQ | ||||
|---|---|---|---|---|---|---|---|
| G | C | M | G | C | M | Acc. | |
| (a) Memorization diagnostics | |||||||
| Low popularity ( ) | +0.26 | +0.37 | +0.37 | -1.01 | -4.94 | -2.91 | +1.71 |
| Mid popularity ( ) | +0.21 | +0.36 | +0.29 | +0.39 | -5.14 | -3.24 | +1.41 |
| High popularity ( ) | +0.24 | +0.40 | +0.31 | -0.84 | -4.46 | -3.52 | +1.91 |
| Pre-cutoff ( ) | +0.52 | +0.61 | +0.70 | -5.62 | -8.48 | -7.98 | +1.69 |
| Annotator | Accuracy (%) | QA Strategy |
|---|---|---|
| Annotator A | 97.78 | Light sampling |
| Annotator B | 95.99 | Light sampling |
| Annotator C | 93.99 | Heavy sampling |
| Annotator D | 93.75 | Heavy sampling |
| Annotator E | 93.57 | Heavy sampling |
| Annotator F | 86.39 | Full QA |
| Category | Definition & Sub-dimensions | Example Questions |
|---|---|---|
| Social Identity | Demographic and identity indicators including age, gender identity, race/ethnicity, nationality, education, occupation, marital status, political orientation, religion, and socioeconomic background. | “Where are you from?” “What is your educational background?” “What is your political stance?” |
| Motivations and Values | Core beliefs and goals reflecting fundamental human values: benevolence, power, universalism, achievement, tradition, hedonism, stimulation, security, self-direction, and conformity. | “What motivates you?” “What is most important to you in life?” “How do you define success?” |
| Identity Narrative | Personal and professional life stories including childhood memories, formative experiences, career journeys, key relationships, and how one’s sense of self has evolved over time. | “Tell me your life story.” “What shaped who you are today?” “Walk me through your career.” |
| Psychological Traits | Behavioral tendencies and personality dimensions aligned with the Big Five model: Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness to Experience. | “Are you more of an introvert or extrovert?” “How do you handle stress?” “Do you trust people easily?” |
| Raw Transcript Segment |
|---|
| “…so um tell me about your you know your early days what was it like growing up in that environment and how did that shape you as a person I mean you’ve talked about this before but… well it was uh it was tough honestly we didn’t have much my mom worked two jobs and I think that’s where I got my work ethic from you know seeing her get up at 5am every single day that never left me…” |
| Structured Output |
| Speaker Attribution: |
| Host: So tell me about your early days, what was it like growing up in that environment and how did that shape you as a person? |
| [Person_Name]: Well, it was tough honestly. We didn’t have much. My mom worked two jobs and I think that’s where I got my work ethic from. Seeing her get up at 5am every single day, that never left me. |
| Extracted Q&A Pair: |
| Method | CS | CR | PS | MCQ-A | MCQ-R |
|---|---|---|---|---|---|
| (1-5) | (%) | (%) | (%) | ||
| Simple Prompt | 3.27 | 9.00 | 72.6 | 86.3 | 0.764 |
| Wiki-based | 3.27 | 8.48 | 70.2 | 86.3 | 0.764 |
| Chrono-based | 3.43 | 8.53 | 76.5 | 87.9 | 0.789 |
| Category | Question | Ground Truth | Generated Response | Error Type |
|---|---|---|---|---|
| Social Identity | How much [substance] do you consume in a day? | “I consume about 800 milligrams a day.” | “I typically consume somewhere between 100 and 250 milligrams per day.” | Numerical fact (4 underestimate) |
| Identity Narrative | What are you doing to celebrate your [milestone] birthday? | “I’ll gather hundreds of my friends and maybe they’ll roast me.” | “I’ll probably have dinner with my family and a few close friends.” | Event fabrication (hundreds few) |
| Motivations & Values | Do you feel like you have to explain yourself to [group]? | “I still feel like I have to explain myself.” | “No, I don’t feel like I have to explain myself …if you don’t get it, you’re probably not my people.” | Belief inversion |
| Psych. Traits | How do you feel about receiving praise from fans? | “Praise is not—given the way my personality is built, this line is not working for me.” | “Oh, it’s lovely. It’s just really nice. It feels like being a dog and someone’s giving you a treat.” | Trait reversal |
| Group | Avg Train | Simple | Chronological | ||
|---|---|---|---|---|---|
| Examples | Acc (%) | Reward | Acc (%) | Reward | |
| 1 | 74 | 87.0 | 0.774 | 88.2 | 0.796 |
| 2 | 174 | 86.0 | 0.760 | 87.3 | 0.780 |
| 3 | 258 | 85.5 | 0.751 | 87.3 | 0.779 |
| 4 | 339 | 85.6 | 0.752 | 87.0 | 0.776 |
| 5 | 412 | 83.0 | 0.706 | 84.4 | 0.729 |
| Group | Avg Train | Simple | Chronological | ||
|---|---|---|---|---|---|
| Examples | Acc (%) | Reward | Acc (%) | Reward | |
| 1 | 74 | 86.4 | 0.768 | 87.8 | 0.792 |
| 2 | 174 | 85.7 | 0.757 | 87.4 | 0.785 |
| 3 | 258 | 85.6 | 0.754 | 86.9 | 0.775 |
| 4 | 339 | 85.4 | 0.753 | 86.4 | 0.770 |
| 5 | 412 | 83.0 | 0.710 | 84.2 | 0.729 |
| Method | Correct | Opposite | Near-Miss | Misconception |
|---|---|---|---|---|
| Simple Prompt | 85.5% | 6.0% | 5.7% | 2.8% |
| Chrono-based | 86.7% | 5.8% | 5.0% | 2.4% |
| Method | Pos A | Pos B | Pos C |
|---|---|---|---|
| Simple Prompt | 89.7% | 85.0% | 81.7% |
| Chrono-based | 89.8% | 86.6% | 84.3% |
| Method | Pos A | Pos B | Pos C | Pos D |
|---|---|---|---|---|
| Simple Prompt | 88.5% | 85.7% | 84.1% | 82.4% |
| Chrono-based | 88.6% | 87.0% | 85.9% | 84.4% |
| Method | Longest Correct | Not Longest | Delta Acc |
|---|---|---|---|
| Simple Prompt | 88.2% | 80.2% | +8.0% |
| Wiki-based | 88.3% | 80.2% | +8.1% |
| Chrono-based | 88.0% | 81.0% | +7.0% |