cs.CLFeb 6, 2026

An evidence-guided reinforcement learning method to improve psychiatric reasoning in small language models

Authors: Xinxin LinGuangxin DaiYi ZhongXiang LiXue XiaoJian LiuYixin ZhangLingming Hu+20 more

Organizations: The Chinese University of Hong Kong, Shatin, N.T., Hong Kong SAR, China · Shandong University, Jinan, China · Peking University Sixth Hospital, Beijing, China · Inspur Cloud Information Technology Co., Ltd., Jinan, China · Fudan University, Shanghai, China · Shandong Provincial Hospital Affiliated to Shandong First Medical University, Jinan, China · Jilin University, Changchun, China · Hong Kong ICI Cloud Service Limited, Hong Kong SAR, China · Shandong First Medical University, Jinan, China

Abstract

Privacy and computational constraints limit the use of large language models in psychiatry, while adapting small language models (SLMs) often requires substantial data and expert annotation. We developed ClinMPO, an evidence-guided reinforcement-learning framework guided by the psychiatrist-defined Clinical Psychiatry Thinking Strategy (CPTS). ClinMPO uses ClinRM, a reward model trained on 18,569 question--answer pairs from 4,474 psychiatry articles. We evaluated four Qwen3 sizes on 1,737 model-screened questions. ClinMPO outperformed Base, supervised fine-tuning and standard group relative policy optimization across scales. From responses by 300 senior pre-licensure medical students, we established the human baseline, a medical-student reference. The 4B model approached this baseline, whereas the 8B model surpassed it and ranked first among 31 models and post-training variants. ClinMPO improved performance across two complementary schemes covering ICD-11 diagnostic categories and psychiatric practice competencies. Blinded assessment by three clinicians showed improved rationale quality across CPTS criteria. These findings highlight how existing clinical evidence and specialist knowledge can be incorporated into the development of medical AI systems through evidence-guided learning.

Explore similar work

CardsList