In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study these behaviors at scale. We present an de-identified corpus of 33 cognitive assessment conversations with 8,250 utterances annotated for three speaker roles and 56 dialogue acts. We use this corpus to benchmark large language models on fine-grained dialogue-act classification and next-patient-utterance generation. We also test whether out-of-domain instruction data and explanation-augmented training transfer to this clinical setting. Instruction tuning produces the strongest patient-utterance reference matching and improves classification accuracy. Reasoning-aware fine-tuning produces the strongest classification results among the LLaMA-3.1-8B variants. However, even the best models struggle to separate closely related dialogue acts, showing that broad conversational intent is easier to recognize than fine-grained communicative function. The corpus and benchmark make interaction structure measurable in cognitive assessments and support follow-up work on conversational markers, clinician education, and carefully validated simulated patients. This work does not make diagnostic claims. Instead, it provides the data and evaluation framework needed to study these applications.
Figures & tables
Figure 1: Overview of the data construction, model adaptation, and evaluation tasks.
Dataset
Task
Records
Alpaca
Instruction Following
25,000
DailyDialog
General Conversation
4,500
ECC
Emotion and Cause Reasoning
3,836
Behavior-SD
Behavioral Dialogue
2,500
WikiDialogue
Dialogue Act
2,500
CoQA
Question Answering
2,000
Table 3: Composition of the multi-task instruction corpus.
Model
Prompt
Acc.
mF1
wF1
Dummy Baselines
Majority Classifier
—
.136
.004
.033
Random Classifier (Stratified)
—
.081
.019
.081
Random Classifier (Uniform)
—
.018
.010
.027
Baseline Model
LLaMA3.1-8B
Zero-shot
.234
.117
.239
Table 4: Performance comparison of prompting strategies, instruction tuning, and reasoning-aware fine-tuning for dialogue act classification.
Model
Acc.
mF1
wF1
Dummy (Majority)
.139
.004
.034
Dummy (Uniform)
.018
.010
.026
Linear SVM
.546
.157
.511
RoBERTa-Large
.681
.266
.668
BioMedBERT
.654
.164
.612
LLaMA3.1-8B + Reasoning (Base)
.418
.113
.342
Table 5: Performance comparison of prompting vs. supervised models.
Model
BLEU
ROUGE-1
ROUGE-2
ROUGE-L
BERTScore-F1
Gemma-3-12B
.021
.140
.033
.139
.906
Qwen3-30B
.025
.168
.038
.161
.898
Qwen3-30B (Instruction)
.041
.184
.037
.181
.901
Mistral-3-24B
.018
.157
.023
.153
.874
LLaMA-3.1-8B
.023
.139
.026
.136
.900
LLaMA-3.1-8B (Instruction)
.043
.178
.036
.175
.906
Table 6: Patient utterance generation performance ( k=5 ). Best results are shown in bold and second-best results are underlined.