Substance Use Disorder (SUD) counseling requires patient responses that reflect underlying cognitive states such as beliefs, coping strategies, and readiness for change. Although large language models (LLMs) can generate fluent text, they often fail to produce cognitively coherent and clinically realistic patient behavior, especially under ethical and data-scarce clinical settings. Moreover, deploying frontier-scale LLMs in healthcare applications presents practical challenges including high computational cost, latency, privacy concerns, and limited deployability in resource-constrained environments, motivating the need for cognitively aligned small language models (SLMs). We propose a cognitively grounded framework for SUD patient dialogue generation that explicitly models and aligns latent cognitive components with patient histories and counselor questions. Our pipeline consists of two stages: cognitive component detection and cognitive component-aligned dialogue generation. To enable effective learning with smaller models, we combine knowledge distillation from high-capacity teacher models, preference optimization from human-annotations, and attention-guided reward shaping. Extensive evaluations using automatic scores like BERTScore, ROUGE, METEOR and BLEU, and LLM-as-judge hit-metrics against both human and teacher-model references show that cognitively informed fine-tuning substantially improves cognitive realization and alignment over a generic instruction-tuned baselines and mental health domain specific SLMs, with particularly strong gains for open-ended cognitive components.
Figures & tables
Fig. 1: Framework for Cognitive Components-Aligned SUD Patient Dialogue Generation in LLMs
Generations
Reference Labels
Human
GPT-5
BERT-F1
R-1
R-2
R-L
METEOR
BLEU
BERT-F1
R-1
R-2
R-L
METEOR
BLEU
Llama3-8B-instruct
0.805
0.0289
0.0074
0.0258
2.31
0.10
0.813
0.0666
0.0191
0.0515
6.81
0.35
MentaLLaMA7B
0.812
0.0259
0.0047
0.0221
1.90
0.09
0.819
0.0617
0.0150
0.0463
5.49
0.30
MentaLLaMA13B
0.815
0.0243
0.0037
0.0210
1.61
0.09
0.819
0.0576
0.0136
0.0450
4.84
0.35
MentaLLaMA13B-Fewshot
0.828
0.0678
0.0069
0.0647
3.46
0.78
0.830
0.1018
0.0185
0.0923
6.94
1.33
TABLE I: Cognitive Component Detection (CCD) performance across BERTScore, ROUGE, METEOR, and BLEU.
TABLE III: Per–cognitive component difference in reference hit ratio between GPT-5 and CCD when evaluated against human annotations ( ΔR=RGPT−RCCD ). Positive values indicate higher GPT coverage, while negative values indicate higher CCD coverage.
TABLE V: Cognitive Component Aligned Dialogue Generation (CADG) performance across various reference labels and baseline models. PM / FM / NM (%) represents Partial Match, Full Match, and No Match ratios respectively.
Model
ROUGE (R-1 / R-2 / R-L)
BLEU
METEOR
Llama8b-instruct
0.012/0.003 /0.010
0.03
1.55
Mentallama-7B
0.024 /0.004 /0.019
0.07
2.29
Mentallama-13B
0.025 / 0.005 / 0.020
0.08
2.39
Mentallama-13B+FS
0.021 / 0.003 / 0.017
0.06
2.02
CADG
0.018 / 0.004 / 0.016
0.04
1.67
TABLE VI: Automatic evaluation scores with ROUGE,BLEU and METEOR for human reference cognitive alignment analysis. Showing our CADG got low rouge score than other baselines except Llama 8B instruct and still achieves high LLM judge alignment score indicating model did not copy paste the CC but it paraphrased for alignment.
Model
Cognitive Component
Counselor Question
Length
ROUGE (R1/R2/RL)
BLEU
METEOR
ROUGE (R1/R2/RL)
BLEU
METEOR
#words <40 /#words > 40
CADG
0.056 / 0.034 / 0.050
1.05
6.75
0.093/0.030/0.068
0.93
14.63
0.0% / 100.0 %
CConlyAttendRwd
0.532 / 0.485 / 0.533
20.49
28.76
0.028/0.003/0.025
0.25
1.36
98.4%/ 1.6%
CQonlyAttendRwd
0.023/0.001/0.020
0.15
1.15
0.647/0.554/0.611
33.45
48.40
94.7% / 5.3%
lengthOnlyRwd
0.047/ 0.024/ 0.038
0.52
5.86
0.081/ 0.022/ 0.058
0.54
13.63
0.0% / 100.0%
TABLE VII: Analysis on GRPO reward ablation study on Cognitive Component attention Counselor Question attention and Length reward (reports average response length)
Fig. 2: Training reward progress of preference optimization (DPO) for chosen and rejected in cognitive aligned dialogue generator.
Fig. 3: Progress of attention- and length-based rewards during GRPO training. The attention-based reward shows how the model progressively increases its attention toward both the cognitive component and counselor question segments, indicating improved focus and alignment over training steps. The length-based reward shows how the model learns to produce responses within the desired length range over training steps.
Fig. 4: Harmonic mean of length and attention reward for different weight distributions.
UKP Lab, Department of Computer Science and Hessian Center for AI (hessian.AI), Technische Universität Darmstadt · Zuse School · National Research Center for Applied Cybersecurity ATHENE +4