Evidence-Traceable Dynamic Interviewer Architecture for Expertise-Adaptive Qualitative Interviews Using Local LLMs
Organizations: Department of Computing, University of Turku, Turku, Finland
Abstract
Automated interviewers and conversational agents are increasingly used in research, recruitment, customer service, and education. However, many existing systems rely on fixed question sequences and provide limited context-based personalization without considering participants' knowledge, which can lead to repetitive or irrelevant follow-up questions. Therefore, there is a need for an adaptive interviewing system that can adjust question depth while maintaining conversational continuity and semantic progression. To address this, an Evidence-Traceable Dynamic Interviewer Architecture is presented using a locally hosted Large Language Model (LLM), with the interview continuously adapted throughout the entire conversation based on the participant's responses and evolving context. The interviewer profiles participants' expertise in real time to generate knowledge-appropriate questions, well-articulated responses, and smooth transition messages that support conversational continuity. A five-module prompt-driven architecture and persistent interview-state record support these functions. The interviewer was evaluated with 246 participants. Expertise Profiling module (M3) showed 78.9% exact agreement with independently reported participant expertise, with a weighted Cohen's K of 0.80. Generate Iterative Questions module (M4) showed a strong expertise-complexity association (p=.79, p<.001), and participants reported high relevance (mean 4.41), engagement (mean 4.32), and satisfaction (mean 4.38), providing evidence that the architecture's adaptive components operated consistently with their intended functions while participants reported a positive interview experience.
Figures & tables
| System-Prompt Component | Main Function |
|---|---|
| Role and Interview Identity | Defines the AI interviewer role and responsibilities |
| General Information | Specifies objectives, research areas, priorities, and question allocation |
| Purpose and Context | Defines language, question type, style, tone, and conversational flow |
| Ethical Guidelines | Controls privacy, sensitive questions, neutrality, and inclusive interaction |
| Knowledge Handling Rules | Controls reliable information use, uncertainty, and hallucination avoidance |
| Behavioral Consistency | Maintains the defined rules across interview turns |
| M1 User-Prompt Component | Main Function |
|---|---|
| LLM Role Specification | Defines the model as a system-prompt generator |
| Task Scope | Limits generation to the supplied interview configuration |
| Clarity of Instruction | Specifies which configuration elements must be included |
| Language Emphasis | Maintains the selected interview language and terminology |
| Output Format Enforcement | Requires a structured, machine-readable output |
| Action Limitation | Prevents M1 from conducting the interview or generating questions |
| M2 Component | Function |
|---|---|
| Research Area and Priority | Selects an appropriate high-priority area for beginning the interview |
| Question Complexity | Starts with a low-complexity, broadly understandable question |
| Contextual Framing | Aligns the question with the research objectives and interview context |
| Question Requirements | Produces one clear, open-ended, non-leading question |
| Behavioral Constraints | Prevents unnecessary explanation, repetition, or multiple questions |
| Output Structure | Returns the initial question in the required structured format |
| Expertise Factor | What M3 Examines |
|---|---|
| F1. Domain Terminology and Specificity | Appropriate use of domain terminology, concepts, technologies, processes, or specific examples |
| F2. Conceptual Correctness and Relevance | Whether objectively verifiable domain statements are coherent, relevant, and technically appropriate |
| F3. Reasoning Depth | Ability to explain relationships, causes, consequences, dependencies, or trade-offs |
| F4. Application and Evaluation | Ability to apply knowledge to situations, compare alternatives, justify decisions, or propose solutions |
| M4 Output | Function |
|---|---|
| Response Message | Briefly acknowledges or reflects the participant’s previous response |
| Transition Message | Creates a natural conversational connection between the previous response and the next question |
| Next Question | Generates a context-aware, expertise-aligned, open-ended follow-up question |
| M5 Component | Function |
|---|---|
| Candidate Question Input | Receives the proposed next question generated by M4 |
| Question-History Comparison | Compares the candidate question with previously asked questions |
| Semantic Overlap Assessment | Identifies conceptual similarity beyond exact wording |
| Research Progression Check | Determines whether the question contributes new information or advances the discussion |
| Validation Decision | Produces an Accept or Regenerate decision |
| Structured Output | Returns the validation result and internal justification |
| Research Area | Priority | Total Question Allocation |
|---|---|---|
| Awareness and knowledge of LLMs among employees | High | 4 |
| Application of LLMs in the organization | Medium | 3 |
| Skill levels and training in using LLMs | High | 3 |
| Data privacy and security in LLM use | Medium | 4 |
| Organizational guidelines for LLM use and adoption | Low | 2 |
| Evaluation Dimension | Purpose |
|---|---|
| Question Relevance and Coherence | Evaluates whether questions were appropriate, logically connected, and relevant to the ongoing discussion |
| Cognitive and Emotional Engagement | Evaluates whether the interaction maintained participants’ attention and involvement |
| Overall User Satisfaction | Evaluates participants’ overall perception of the AI-powered interview experience |
| Verification Component | Comparison / Evidence | Main Purpose |
|---|---|---|
| Expertise Profiling | M3 expertise vs. participant-reported expertise | Assess agreement in expertise classification |
| Question Adaptation | M3 expertise vs. subsequent M4 question complexity | Assess expertise-driven adaptation |
| Question Uniqueness | M5 decisions and question-history comparisons | Examine semantic-overlap filtering and regeneration |
| Question Relevance and Coherence | Participant post-interview rating | Assess conversational quality |
| Engagement | Participant post-interview rating | Assess participant involvement |
| Overall Satisfaction | Participant post-interview rating | Assess overall interview experience |
| Characteristic | (%) | Characteristic | (%) |
|---|---|---|---|
| Gender | Age | ||
| Male | 130 (52.8) | 18–24 years | 31 (12.6) |
| Female | 111 (45.1) | 25–34 years | 78 (31.7) |
| Prefer not to say | 5 (2.0) | 35–44 years | 70 (28.5) |
| 45–54 years | 43 (17.5) | ||
| 55 years or above | 24 (9.8) | ||
| Self-Reported Expertise | M3 Novice | M3 Basic | M3 Advanced | M3 Expert |
|---|---|---|---|---|
| Novice | 44 | 7 | 1 | 0 |
| Basic Knowledge | 8 | 54 | 10 | 1 |
| Advanced Knowledge | 1 | 9 | 58 | 8 |
| Expert | 0 | 1 | 6 | 38 |
| M3 Expertise Level | Mean Complexity | SD |
|---|---|---|
| Novice | 1.42 | 0.55 |
| Basic Knowledge | 2.08 | 0.60 |
| Advanced Knowledge | 2.91 | 0.63 |
| Expert | 3.56 | 0.51 |
| Expertise Level | Representative Question |
|---|---|
| Novice | “Have you ever thought about privacy when using an LLM at work? What kinds of information would you avoid sharing with it?” |
| Basic Knowledge | “What privacy or security problems do you think could arise when employees enter work-related information into publicly available LLMs?” |
| Advanced Knowledge | “How would you assess the risk of sensitive organizational information being exposed through employees’ use of public LLM services?” |
| Expert | “Suppose an organization wants to permit employees to use public LLMs while protecting confidential information. What technical and governance controls would you recommend, and how would you evaluate whether those controls are effective?” |
| Evaluation Dimension | Mean | SD |
|---|---|---|
| Question Relevance and Coherence | 4.41 | 0.58 |
| Cognitive and Emotional Engagement | 4.32 | 0.64 |
| Overall User Satisfaction | 4.38 | 0.61 |