Substance Use Disorder (SUD) counseling requires patient responses that reflect underlying cognitive states such as beliefs, coping strategies, and readiness for change. Although large language models (LLMs) can generate fluent text, they often fail to produce cognitively coherent and clinically realistic patient behavior, especially under ethical and data-scarce clinical settings. Moreover, deploying frontier-scale LLMs in healthcare applications presents practical challenges including high computational cost, latency, privacy concerns, and limited deployability in resource-constrained environments, motivating the need for cognitively aligned small language models (SLMs). We propose a cognitively grounded framework for SUD patient dialogue generation that explicitly models and aligns latent cognitive components with patient histories and counselor questions. Our pipeline consists of two stages: cognitive component detection and cognitive component-aligned dialogue generation. To enable effective learning with smaller models, we combine knowledge distillation from high-capacity teacher models, preference optimization from human-annotations, and attention-guided reward shaping. Extensive evaluations using automatic scores like BERTScore, ROUGE, METEOR and BLEU, and LLM-as-judge hit-metrics against both human and teacher-model references show that cognitively informed fine-tuning substantially improves cognitive realization and alignment over a generic instruction-tuned baselines and mental health domain specific SLMs, with particularly strong gains for open-ended cognitive components.
Figures & tables
Fig. 1: Framework for Cognitive Components-Aligned SUD Patient Dialogue Generation in LLMs
Generations
Reference Labels
Human
GPT-5
BERT-F1
R-1
R-2
R-L
METEOR
BLEU
BERT-F1
R-1
R-2
R-L
METEOR
BLEU
Llama3-8B-instruct
0.805
0.0289
0.0074
0.0258
2.31
0.10
0.813
0.0666
0.0191
0.0515
6.81
0.35
MentaLLaMA7B
0.812
0.0259
0.0047
0.0221
1.90
0.09
0.819
0.0617
0.0150
0.0463
5.49
0.30
MentaLLaMA13B
0.815
0.0243
0.0037
0.0210
1.61
0.09
0.819
0.0576
0.0136
0.0450
4.84
0.35
MentaLLaMA13B-Fewshot
0.828
0.0678
0.0069
0.0647
3.46
0.78
0.830
0.1018
0.0185
0.0923
6.94
1.33
TABLE I: Cognitive Component Detection (CCD) performance across BERTScore, ROUGE, METEOR, and BLEU.
TABLE III: Per–cognitive component difference in reference hit ratio between GPT-5 and CCD when evaluated against human annotations ( ΔR=RGPT−RCCD ). Positive values indicate higher GPT coverage, while negative values indicate higher CCD coverage.
TABLE V: Cognitive Component Aligned Dialogue Generation (CADG) performance across various reference labels and baseline models. PM / FM / NM (%) represents Partial Match, Full Match, and No Match ratios respectively.
Model
ROUGE (R-1 / R-2 / R-L)
BLEU
METEOR
Llama8b-instruct
0.012/0.003 /0.010
0.03
1.55
Mentallama-7B
0.024 /0.004 /0.019
0.07
2.29
Mentallama-13B
0.025 / 0.005 / 0.020
0.08
2.39
Mentallama-13B+FS
0.021 / 0.003 / 0.017
0.06
2.02
CADG
0.018 / 0.004 / 0.016
0.04
1.67
TABLE VI: Automatic evaluation scores with ROUGE,BLEU and METEOR for human reference cognitive alignment analysis. Showing our CADG got low rouge score than other baselines except Llama 8B instruct and still achieves high LLM judge alignment score indicating model did not copy paste the CC but it paraphrased for alignment.
Model
Cognitive Component
Counselor Question
Length
ROUGE (R1/R2/RL)
BLEU
METEOR
ROUGE (R1/R2/RL)
BLEU
METEOR
#words <40 /#words > 40
CADG
0.056 / 0.034 / 0.050
1.05
6.75
0.093/0.030/0.068
0.93
14.63
0.0% / 100.0 %
CConlyAttendRwd
0.532 / 0.485 / 0.533
20.49
28.76
0.028/0.003/0.025
0.25
1.36
98.4%/ 1.6%
CQonlyAttendRwd
0.023/0.001/0.020
0.15
1.15
0.647/0.554/0.611
33.45
48.40
94.7% / 5.3%
lengthOnlyRwd
0.047/ 0.024/ 0.038
0.52
5.86
0.081/ 0.022/ 0.058
0.54
13.63
0.0% / 100.0%
TABLE VII: Analysis on GRPO reward ablation study on Cognitive Component attention Counselor Question attention and Length reward (reports average response length)
Fig. 2: Training reward progress of preference optimization (DPO) for chosen and rejected in cognitive aligned dialogue generator.
Fig. 3: Progress of attention- and length-based rewards during GRPO training. The attention-based reward shows how the model progressively increases its attention toward both the cognitive component and counselor question segments, indicating improved focus and alignment over training steps. The length-based reward shows how the model learns to produce responses within the desired length range over training steps.
Fig. 4: Harmonic mean of length and attention reward for different weight distributions.
In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study these behaviors at scale. We present an de-identified corpus of 33 cognitive assessment conversations with 8,250 utterances annotated for three speaker roles and 56 dialogue acts. We use this corpus to benchmark large language models on fine-grained dialogue-act classification and next-patient-utterance generation. We also test whether out-of-domain instruction data and explanation-augmented training transfer to this clinical setting. Instruction tuning produces the strongest patient-utterance reference matching and improves classification accuracy. Reasoning-aware fine-tuning produces the strongest classification results among the LLaMA-3.1-8B variants. However, even the best models struggle to separate closely related dialogue acts, showing that broad conversational intent is easier to recognize than fine-grained communicative function. The corpus and benchmark make interaction structure measurable in cognitive assessments and support follow-up work on conversational markers, clinician education, and carefully validated simulated patients. This work does not make diagnostic claims. Instead, it provides the data and evaluation framework needed to study these applications.
Vishalakshi Arumugam, Dan Schumacher, Veronica Rammouz +3
College of AI, Cyber and Computing University of Texas at San Antonio, USA
Rising demand for mental health support has increased interest in using Large Language Models (LLMs) for counseling, but adapting them to this safety-critical domain is hindered by limited real-world data due to privacy constraints. Synthetic datasets provide a promising alternative, but existing approaches often rely on unstructured or semi-structured text inputs and overlook structural dependencies between a client's cognitive, emotional, and behavioral states, leading to psychologically inconsistent and less realistic interactions. We introduce Graph2Counsel, a framework for generating synthetic counseling sessions grounded in Client Psychological Graphs (CPGs) that encode relationships among clients' thoughts, emotions, and behaviors. Graph2Counsel uses a structured prompting pipeline guided by counselor strategies and CPG, and explores prompting strategies including CoT and Multi-Agent Feedback. It produces 760 sessions from 76 CPGs across diverse client profiles. In expert evaluation, our dataset outperforms prior datasets on specificity, counselor competence, authenticity, conversational flow, and safety (Krippendorff's α = 0.70). Fine-tuning an open-source model on this dataset improves performance on several metrics on CounselingBench and CounselBench, while matching baselines on others. We further introduce Therapy-Eval, a multi-turn evaluation framework, and demonstrate the effectiveness of our fine-tuned model in realistic therapeutic conversations. We make our code, data and fine-tuned model public.
Aishik Mandal, Hiba Arnaout, Clarissa W. Ong +5
UKP Lab, Department of Computer Science and Hessian Center for AI (hessian.AI), Technische Universität Darmstadt · Zuse School · National Research Center for Applied Cybersecurity ATHENE +4
Mental health challenges are increasing worldwide, straining emotional support services and leading to counselor overload. This can result in delayed responses during critical situations, such as suicidal ideation, where timely intervention is essential. While large language models (LLMs) have shown strong generative capabilities, their application in low-resource languages, especially in sensitive domains like mental health, remains underexplored. Furthermore, existing LLM-based agents often struggle to replicate the supportive language and intervention strategies used by professionals due to a lack of training on large-scale, real-world datasets. To address this, we propose CARE (Counselor-Aligned Response Engine), a GenAI framework that assists counselors by generating real-time, psychologically aligned response recommendations. CARE fine-tunes open-source LLMs separately for Hebrew and Arabic using curated subsets of real-world crisis conversations. The training data consists of sessions rated as highly effective by professional counselors, enabling the models to capture interaction patterns associated with successful de-escalation. By training on complete conversation histories, CARE maintains the evolving emotional context and dynamic structure of counselor-help-seeker dialogue. In experimental settings, CARE demonstrates stronger semantic and strategic alignment with gold-standard counselor responses compared to non-specialized LLMs. These findings suggest that domain-specific fine-tuning on expert-validated data can significantly support counselor workflows and improve care quality in low-resource language contexts.
Hagai Astrin, Ayal Swaid, Avi Segal +1
Ben-Gurion University · School of Informatics, University of Edinburgh Edinburgh, UK