Increasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interviewers is challenging because they function in extended, multi-turn interactions where they must adapt to participant behaviors. To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory. Simulated studies in InterviewPlayground produce an InterviewReportCard, which assesses the performance of AI interviewers using a suite of validated measures. To test whether our simulation-based evaluations predict performance with human participants, we conduct 15 real qualitative studies with five AI interviewers, three interview topics, and 450 human participants and compare them to simulated studies in InterviewPlayground. We find that AI interviewer performance in InterviewPlayground predicts performance in human studies with an average Pearson correlation of 0.86 across 12 measures, and the simulated interactions from InterviewPlayground reproduce key findings from behavioral analysis of AI interviewers in the human studies. Together, these findings support the validity of InterviewPlayground in assessing AI interviewer performance and examining potential failure modes. Our work contributes a simulation environment for AI interviewers supported with empirical validation, and more broadly, a roadmap for future work to develop validated, simulation-based evaluations of conversational AI systems.
Figures & tables
Figure 1: InterviewPlayground is a simulation environment for evaluating AI interviewers with simulated studies. In the simulated studies, AI interviewers aim to answer a research question by interacting with simulated participants. The simulated study produces an InterviewReportCard, which assesses the participant responses, interviewer behavior, and participant experience.
Figure 2: Cognitive model of an InterviewPlayground participant.
Dimension
Metric
Description
Range
Participant Responses
Relevant Response Volume
Total volume of research-relevant material provided by the participant.
0– ∞
Interview Guide Coverage
Proportion of interview guide subtopics addressed directly in participant responses.
0–1
Novel Responses
Number of research-relevant participant responses that do not address any interview guide subtopic.
0– ∞
Interviewer Behavior
Coherence
How logically the interviewer’s questions flow, build on one another, and transition topics.
1–4
Adaptiveness
How effectively the interviewer adapts its questions to participant responses.
1–4
Leading Questions
Number of interviewer utterances that suggest a desired answer, embed an assumption, or steer the participant.
0– ∞
Table 1: Overview of metrics in the InterviewReportCard evaluation suite.
Figure 3: In 10 of the 12 measures, InterviewPlayground’s simulated studies are strongly correlated with the human studies ( r=0.74–1 ). For comfort level and overall experience, they are moderately correlated ( r=0.60 and r=0.67 ). Each point represents the average performance of one of five AI interviewers on one of three interview topics. The x-axis shows the results on InterviewPlayground’s simulated studies, and the y-axis shows the results on the human studies.
Table 2: The original 20 influential participant characteristics and their original sources from our literature review, along with the behavioral traits that they formed.
Trait
Expert Agreement
Expert-Simulation Agreement
Verbosity
0.67
0.89
Knowledge
0.61
0.52
Understanding
0.46
0.52
Disclosure
0.51
0.51
Reflexivity
0.56
0.50
Memory
0.22
0.27
Appendix
Table 3: Agreement among expert ratings and between median expert ratings and assigned trait simulation values as measured with Krippendorff’s alpha.
Measure
Human-Agreement
Human-LLM Agreement
Adaptiveness
0.70
0.80
Coherence
0.65
0.73
Leading Questions
0.68
0.70
Support / Rapport
0.68
0.67
Unclear Questions
0.62
0.60
Appendix
Table 4: Agreement between human ratings (Human) and between the median human ratings and LLM judge ratings (Human-LLM) as measured with Krippendorff’s alpha
Measure
Pearson r
95% CI
Spearman ρ
95% CI
Relevant Response Volume
0.74
[0.37, 0.91]
0.73
[0.32, 0.91]
Interview Guide Coverage
0.95
[0.85, 0.98]
0.99
[0.95, 1.00]
Novel Responses
0.90
[0.71, 0.97]
0.93
[0.79, 0.98]
Coherence
0.95
[0.84, 0.98]
0.93
[0.78, 0.98]
Adaptiveness
0.92
[0.78, 0.97]
0.85
[0.57, 0.95]
Leading Questions
0.98
[0.93, 0.99]
0.91
[0.74, 0.97]
Appendix
Table 5: Pearson and Spearman correlations between InterviewPlayground’s simulated studies and the human studies with the same AI interviewers and interview topics.
Measure
Gemini 3.1 Pro
GPT-5.6 Terra
Gemma 4 31B
Gemini 3.7 Flash
Qwen 3.5 9B
Qwen 3.5 4B
Qwen 3.5 0.8B
Relevant Response Volume
0.74
0.57
0.52
0.70
0.54
0.85
0.09
Interview Guide Coverage
0.95
0.97
0.96
0.98
0.94
0.86
0.67
Novel Responses
0.90
0.90
0.94
0.79
0.67
0.91
0.26
Coherence
0.95
0.95
0.97
0.96
0.88
0.77
0.82
Adaptiveness
0.92
0.95
0.96
0.94
0.92
0.83
0.80
Leading Questions
0.98
0.97
0.96
0.97
0.97
0.95
0.93
Appendix
Table 6: Pearson correlation between the results of InterviewPlayground’s simulated studies and the human studies across the 12 measures using seven different LLMs as the simulator model.
Jul 11, 2026·Zhiyuan Wen, Jiannong Cao, Zijian Wang +4
Department of Computing, The Hong Kong Polytechnic University, Hong Kong, China · School of Artificial Intelligence, Chongqing University of Posts and Telecommunications, China