Increasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interviewers is challenging because they function in extended, multi-turn interactions where they must adapt to participant behaviors. To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory. Simulated studies in InterviewPlayground produce an InterviewReportCard, which assesses the performance of AI interviewers using a suite of validated measures. To test whether our simulation-based evaluations predict performance with human participants, we conduct 15 real qualitative studies with five AI interviewers, three interview topics, and 450 human participants and compare them to simulated studies in InterviewPlayground. We find that AI interviewer performance in InterviewPlayground predicts performance in human studies with an average Pearson correlation of 0.86 across 12 measures, and the simulated interactions from InterviewPlayground reproduce key findings from behavioral analysis of AI interviewers in the human studies. Together, these findings support the validity of InterviewPlayground in assessing AI interviewer performance and examining potential failure modes. Our work contributes a simulation environment for AI interviewers supported with empirical validation, and more broadly, a roadmap for future work to develop validated, simulation-based evaluations of conversational AI systems.
Figures & tables
Figure 1: InterviewPlayground is a simulation environment for evaluating AI interviewers with simulated studies. In the simulated studies, AI interviewers aim to answer a research question by interacting with simulated participants. The simulated study produces an InterviewReportCard, which assesses the participant responses, interviewer behavior, and participant experience.
Figure 2: Cognitive model of an InterviewPlayground participant.
Dimension
Metric
Description
Range
Participant Responses
Relevant Response Volume
Total volume of research-relevant material provided by the participant.
0– ∞
Interview Guide Coverage
Proportion of interview guide subtopics addressed directly in participant responses.
0–1
Novel Responses
Number of research-relevant participant responses that do not address any interview guide subtopic.
0– ∞
Interviewer Behavior
Coherence
How logically the interviewer’s questions flow, build on one another, and transition topics.
1–4
Adaptiveness
How effectively the interviewer adapts its questions to participant responses.
1–4
Leading Questions
Number of interviewer utterances that suggest a desired answer, embed an assumption, or steer the participant.
0– ∞
Table 1: Overview of metrics in the InterviewReportCard evaluation suite.
Figure 3: In 10 of the 12 measures, InterviewPlayground’s simulated studies are strongly correlated with the human studies ( r=0.74–1 ). For comfort level and overall experience, they are moderately correlated ( r=0.60 and r=0.67 ). Each point represents the average performance of one of five AI interviewers on one of three interview topics. The x-axis shows the results on InterviewPlayground’s simulated studies, and the y-axis shows the results on the human studies.
Table 2: The original 20 influential participant characteristics and their original sources from our literature review, along with the behavioral traits that they formed.
Trait
Expert Agreement
Expert-Simulation Agreement
Verbosity
0.67
0.89
Knowledge
0.61
0.52
Understanding
0.46
0.52
Disclosure
0.51
0.51
Reflexivity
0.56
0.50
Memory
0.22
0.27
Appendix
Table 3: Agreement among expert ratings and between median expert ratings and assigned trait simulation values as measured with Krippendorff’s alpha.
Measure
Human-Agreement
Human-LLM Agreement
Adaptiveness
0.70
0.80
Coherence
0.65
0.73
Leading Questions
0.68
0.70
Support / Rapport
0.68
0.67
Unclear Questions
0.62
0.60
Appendix
Table 4: Agreement between human ratings (Human) and between the median human ratings and LLM judge ratings (Human-LLM) as measured with Krippendorff’s alpha
Measure
Pearson r
95% CI
Spearman ρ
95% CI
Relevant Response Volume
0.74
[0.37, 0.91]
0.73
[0.32, 0.91]
Interview Guide Coverage
0.95
[0.85, 0.98]
0.99
[0.95, 1.00]
Novel Responses
0.90
[0.71, 0.97]
0.93
[0.79, 0.98]
Coherence
0.95
[0.84, 0.98]
0.93
[0.78, 0.98]
Adaptiveness
0.92
[0.78, 0.97]
0.85
[0.57, 0.95]
Leading Questions
0.98
[0.93, 0.99]
0.91
[0.74, 0.97]
Appendix
Table 5: Pearson and Spearman correlations between InterviewPlayground’s simulated studies and the human studies with the same AI interviewers and interview topics.
Measure
Gemini 3.1 Pro
GPT-5.6 Terra
Gemma 4 31B
Gemini 3.7 Flash
Qwen 3.5 9B
Qwen 3.5 4B
Qwen 3.5 0.8B
Relevant Response Volume
0.74
0.57
0.52
0.70
0.54
0.85
0.09
Interview Guide Coverage
0.95
0.97
0.96
0.98
0.94
0.86
0.67
Novel Responses
0.90
0.90
0.94
0.79
0.67
0.91
0.26
Coherence
0.95
0.95
0.97
0.96
0.88
0.77
0.82
Adaptiveness
0.92
0.95
0.96
0.94
0.92
0.83
0.80
Leading Questions
0.98
0.97
0.96
0.97
0.97
0.95
0.93
Appendix
Table 6: Pearson correlation between the results of InterviewPlayground’s simulated studies and the human studies across the 12 measures using seven different LLMs as the simulator model.
Simulating real personalities with large language models requires grounding generation in authentic personal data. Existing evaluation approaches rely on demographic surveys, personality questionnaires, or short AI-led interviews as proxies, but lack direct assessment against what individuals actually said. We address this gap with an interview-grounded evaluation framework for personality simulation at a large scale. We extract over 671,000 question-answer pairs from 23,000 verified interview transcripts across 1,000 public personalities, each with an average of 11.5 hours of interview content. We propose a multi-dimensional evaluation framework with four complementary metrics measuring content similarity, factual consistency, personality alignment, and factual knowledge retention. Through systematic comparison, we find that interview grounding yields consistent gains in content alignment and exact-match factual recall over biographical profiles and parametric prompting. We further find complementary strengths: retrieval-augmented methods tend to preserve personality alignment, while larger chronological contexts generally reduce contradictions and improve factual recall. Our evaluation framework enables principled method selection based on application requirements, and our empirical findings provide actionable insights for advancing personality simulation research.
This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort and cost. Therefore, testing with user simulators can be beneficial. Since most conventional user simulators have been primarily designed for training task-oriented dialogue systems, little attention has been paid to the personas of the simulated users. During development, testing interview dialogue systems requires simulating a wide range of user behaviors, but manually creating a large number of personas is labor-intensive. We propose a method that automatically generates personas for user simulators using a large language model. Furthermore, by assigning personality traits related to communication styles when generating personas, we aim to increase the diversity of communication styles in the user simulator. Experimental results show that the proposed method enables the user simulator to generate utterances with greater variation.
There are now multiple proposals for systems based on Large Language Models (LLMs) to conduct automated qualitative interviews, but most of the current solutions rely on proprietary LLMs, which compromises reproducibility and data security. They also rely on LLMs for all interview tasks, which limits standardisation of question wording as well as control over question order. To address these issues, we introduce the AInterviewer platform, an opensource solution based on a multi-agent pipeline that combines controlled question administration of survey software with the flexibility of LLMs. AInterviewer is an interdisciplinary effort designed to implement best practices of qualitative interviewing in social science, and it can run with locally hosted models to ensure security, transparency, and reproducibility. Our platform provides a web-based GUI supporting each phase of data collection: from interview guide design and pilot testing to interview distribution and data collection monitoring.
Tobias Priesholm Gardhus, Nikolas Vitsakis, Fie Lejre Frederiksen +2
University of Copenhagen · IT University of Copenhagen