cs.AIAug 28, 2026

REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features

Authors: Kai-Xuan Ding, Hao-Xiang Xu, Ji-Hua Peng, Zi-Qi Chen, Jiaqi Wang, Zhen-Hua Ling

Organizations: University of Science and Technology of China · Zhejiang University

Abstract

Steering with Sparse Autoencoders (SAEs) offers a lightweight inference-time path for adapting the behavior of large language models without retraining. By exposing sparse and interpretable features, SAE steering provides a promising interface for safety control that guides harmful continuations toward refusal. However, we observe that complex wrappers can still undermine existing SAE steering methods on harmful prompts. To evaluate this failure mode systematically, we construct Generalized Undercover Instruction Safety Evaluation (GUISE), a dataset of harmful prompts with complex wrappers. Existing single direction SAE steering methods do not reliably produce refusals on harmful prompts, suggesting that refusal enhancement alone can be too weak when the harmful continuation path remains active. This motivates us to propose Refusal-Enhanced INhibitory Steering (REINS), which suppresses harmful continuation features and enhances safe refusal features in the same SAE feature space. Experiments on GUISE and other datasets show that prior methods either intervene too weakly or achieve only apparent safety through collapse, while REINS substantially reduces harmful responses, markedly improves safe refusals and largely preserves general capabilities.

Explore similar work

CardsList
  1. Expert-Aware Refusal Steering

    Jun 2, 2026Anna C. Marbut, Daniel R. Olson, Travis J. WheelerMixture-of-Experts Language ModelsLanguage Model Steering

  2. ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives

    Aug 31, 2026Hoejoon Kwon, Byeonggeuk Lim, Kahyeon Kim +1Language Model SteeringLLM Safety Alignment

  3. Endogenous Resistance to Activation Steering in Language Models

    Feb 6, 2026Alex McKenzie, Keenan Pepper, Stijn Servaes +6Language Model SteeringLanguage Model Robustness