cs.CLSep 28, 2026

Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

Authors: Xu Wang, Difan Zou, Xuansheng Wu

Organizations: The University of Hong Kong · Shenzhen Loop Area Institute · Shanghai Artificial Intelligence Laboratory

Abstract

Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate this possibility using compensatory feature injection (CFI), a training technique designed to limit the acquisition of a target concept by supplying its associated activation during learning. Across three Qwen3.5 base models, we use sparse autoencoders (SAEs) to identify the top-ranked sycophancy feature from paired sycophantic and independent responses, then validate its behavioral influence through inference steering. We subsequently inject the selected feature during supervised fine-tuning on sycophantic targets. Positive injection reduces learned sycophancy after removal (by 62.0% relative to ordinary fine-tuning in 35B-A3B), whereas modest negative injection increases it. Unexpectedly, these reductions in sycophancy do not consistently improve direct refusal of harmful requests, motivating a narrower evaluation of the same harmful intents under user pressure. In this setting, ordinary fine-tuning on sycophantic responses substantially weakens refusal, while selected checkpoints trained with positive injection recover part of the loss, including approximately 95% in 35B-A3B. These findings show that persistent sycophancy reduction does not guarantee stronger direct refusal, while identifying recovery under user pressure as a distinct, conditional benefit of training intervention.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models

    Jan 21, 2026Shahar Ben-Natan, Oren TsurSycophancyLarge Language Model Bias

  2. Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

    Sep 3, 2026Hoang Cuong Nguyen, Mark Dras, Usman NaseemRefusalsLarge Language Model Alignment