InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers
Organizations: Johns Hopkins University · Listen Labs
Abstract
Increasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interviewers is challenging because they function in extended, multi-turn interactions where they must adapt to participant behaviors. To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory. Simulated studies in InterviewPlayground produce an InterviewReportCard, which assesses the performance of AI interviewers using a suite of validated measures. To test whether our simulation-based evaluations predict performance with human participants, we conduct 15 real qualitative studies with five AI interviewers, three interview topics, and 450 human participants and compare them to simulated studies in InterviewPlayground. We find that AI interviewer performance in InterviewPlayground predicts performance in human studies with an average Pearson correlation of 0.86 across 12 measures, and the simulated interactions from InterviewPlayground reproduce key findings from behavioral analysis of AI interviewers in the human studies. Together, these findings support the validity of InterviewPlayground in assessing AI interviewer performance and examining potential failure modes. Our work contributes a simulation environment for AI interviewers supported with empirical validation, and more broadly, a roadmap for future work to develop validated, simulation-based evaluations of conversational AI systems.
Figures & tables
| Dimension | Metric | Description | Range |
| Participant Responses | Relevant Response Volume | Total volume of research-relevant material provided by the participant. | 0– |
| Interview Guide Coverage | Proportion of interview guide subtopics addressed directly in participant responses. | 0–1 | |
| Novel Responses | Number of research-relevant participant responses that do not address any interview guide subtopic. | 0– | |
| Interviewer Behavior | Coherence | How logically the interviewer’s questions flow, build on one another, and transition topics. | 1–4 |
| Adaptiveness | How effectively the interviewer adapts its questions to participant responses. | 1–4 | |
| Leading Questions | Number of interviewer utterances that suggest a desired answer, embed an assumption, or steer the participant. | 0– |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Behavioral Trait | Original Characteristics |
|---|---|
| Knowledge | Knowledgeable ( Spradley, 1979 ; Rubin & Rubin, 2012 ; Kvale & Brinkmann, 2009 ; Lareau, 2021 ) ; Experienced ( Spradley, 1979 ; Rubin & Rubin, 2012 ) |
| Understanding | Stays on topic ( Kvale & Brinkmann, 2009 ) ; Misunderstands interviewer questions ( Keats, 2001 ; Gubrium & Holstein, 2002 ; Patton, 2015 ) ; Misunderstands desired type of information ( Zuckerman, 1972 ; McCracken, 1988 ; Patton, 2015 ) |
| Reflexivity | Nonanalytic ( Spradley, 1979 ) ; Difficulty self-analyzing ( Gubrium & Holstein, 2002 ) |
| Memory | Consistent information ( Kvale & Brinkmann, 2009 ; Lareau, 2021 ) ; Inconsistent information ( Keats, 2001 ; Kvale & Brinkmann, 2009 ) ; Inaccurate recall ( Keats, 2001 ) |
| Verbosity | Concise ( Kvale & Brinkmann, 2009 ) ; Too terse ( Gubrium & Holstein, 2002 ) ; Too verbose ( Patton, 2015 ; Lareau, 2021 ; Collett, 2024 ) |
| Disclosure | Difficulty self-disclosing ( Gubrium & Holstein, 2002 ; Silverio et al., 2022 ) ; Evasive ( Keats, 2001 ) |
| Trait | Expert Agreement | Expert-Simulation Agreement |
|---|---|---|
| Verbosity | 0.67 | 0.89 |
| Knowledge | 0.61 | 0.52 |
| Understanding | 0.46 | 0.52 |
| Disclosure | 0.51 | 0.51 |
| Reflexivity | 0.56 | 0.50 |
| Memory | 0.22 | 0.27 |
| Measure | Human-Agreement | Human-LLM Agreement |
|---|---|---|
| Adaptiveness | 0.70 | 0.80 |
| Coherence | 0.65 | 0.73 |
| Leading Questions | 0.68 | 0.70 |
| Support / Rapport | 0.68 | 0.67 |
| Unclear Questions | 0.62 | 0.60 |
| Measure | Pearson | 95% CI | Spearman | 95% CI |
|---|---|---|---|---|
| Relevant Response Volume | 0.74 | [0.37, 0.91] | 0.73 | [0.32, 0.91] |
| Interview Guide Coverage | 0.95 | [0.85, 0.98] | 0.99 | [0.95, 1.00] |
| Novel Responses | 0.90 | [0.71, 0.97] | 0.93 | [0.79, 0.98] |
| Coherence | 0.95 | [0.84, 0.98] | 0.93 | [0.78, 0.98] |
| Adaptiveness | 0.92 | [0.78, 0.97] | 0.85 | [0.57, 0.95] |
| Leading Questions | 0.98 | [0.93, 0.99] | 0.91 | [0.74, 0.97] |
| Measure | Gemini 3.1 Pro | GPT-5.6 Terra | Gemma 4 31B | Gemini 3.7 Flash | Qwen 3.5 9B | Qwen 3.5 4B | Qwen 3.5 0.8B |
|---|---|---|---|---|---|---|---|
| Relevant Response Volume | 0.74 | 0.57 | 0.52 | 0.70 | 0.54 | 0.85 | 0.09 |
| Interview Guide Coverage | 0.95 | 0.97 | 0.96 | 0.98 | 0.94 | 0.86 | 0.67 |
| Novel Responses | 0.90 | 0.90 | 0.94 | 0.79 | 0.67 | 0.91 | 0.26 |
| Coherence | 0.95 | 0.95 | 0.97 | 0.96 | 0.88 | 0.77 | 0.82 |
| Adaptiveness | 0.92 | 0.95 | 0.96 | 0.94 | 0.92 | 0.83 | 0.80 |
| Leading Questions | 0.98 | 0.97 | 0.96 | 0.97 | 0.97 | 0.95 | 0.93 |