cs.LGJul 17, 2026

TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue

Authors: Shuzhong LaiJunhong LaiChenxi LiQing ZhouHaifeng LiGang PanLin YaoYueming Wang

Organizations: 1Nanhu Brain-Computer Interface Institute · 2MOE Frontiers Science Center for Brain and Brain-Machine Integration, Zhejiang University · College of Computer Science and Technology, Zhejiang University · 4Children’s Hospital Zhejiang University School of Medicine · 5State Key Laboratory of Brain-Machine Intelligence · Department of Neurobiology, Affiliated Mental Health Center and Hangzhou Seventh People’s Hospital, Zhejiang University School of Medicine

Abstract

The sycophancy of large language models can increase the safety risk in intervention dialogue for autistic children. Supervised fine-tuning can somewhat reduce sycophancy, but relying solely on positive examples is often insufficient to identify and correct failure patterns. We observe that sycophancy behaviors can often be localized to a limited span within the model response. In this regime, sequence-level preference optimization can over-update preference-irrelevant tokens and degrade intervention ability. To address this, we propose the \textbf{M}inimal \textbf{E}dit \textbf{D}ata \textbf{A}ugmentation (MEDA) strategy to construct controlled, stable, minimal edit preference pairs and \textbf{T}oken-level \textbf{D}ifference \textbf{D}irect \textbf{P}reference \textbf{O}ptimization (TD-DPO), which upweights difference tokens between chosen and rejected responses while downweighting shared tokens to suppress background drift. Extensive experiments across multiple backbones and evaluators show that TD-DPO achieves a better trade-off between sycophancy mitigation and intervention ability retention in our offline settings, highlighting its potential as a practical alignment approach for autism intervention.

Explore similar work

CardsList