cs.CVOct 1, 2026

RASteer: Retain-Aware Activation Steering for Concept Erasure in Diffusion Models

Authors: Yongliang Wu, Haori Lu, Yulun Wu, Jinqi Luo, Xingyu Zhu, Yaoyao Liu

Organizations: University of Illinois Urbana-Champaign · University of Pennsylvania · National University of Singapore

Abstract

Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-image diffusion model while preserving its ability to generate other content. Existing activation steering methods build an erasure direction mainly from the target concept and adjust model activations along it at inference time. However, target and retained concepts often overlap in the model's representation space, so this direction also contains shared components that retained concepts rely on. Steering directly along this direction can therefore suppress retained concepts and harm the generation of non-target content. To address this issue, we propose Retain-aware Activation Steering (RASteer), a training-free method. RASteer first builds a retain subspace from the concepts to preserve. Retain-Orthogonal Steering (ROS) then removes components aligned with this subspace from the erasure direction, making steering more specific to the target. Since fully removing the shared components can weaken erasure, we further introduce Overlap-Adaptive Calibration (OAC). At each layer and denoising step, OAC uses the overlap between the erasure direction and the retain subspace to control how much of each shared component is removed, balancing target erasure and concept preservation. Experiments on unsafe-content, instance, and artistic-style erasure across multiple backbones and benchmarks show that RASteer matches or outperforms the activation steering and weight editing baselines we evaluate, achieving a better balance between erasure and preservation.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Erasing Without Collateral Damage: Precise Concept Removal in Diffusion Models

    Jul 6, 2026Parth Upman, Nishita Jain, Shreyank N GowdaConcept ErasureText-To-Image Diffusion Models

  2. Continual Concept Erasure in Diffusion Models by Suppressing Cross-Edit Interference

    Oct 1, 2026Yongliang Wu, Haori Lu, Jinqi Luo +3Concept ErasureText-To-Image Diffusion Models

  3. PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders

    Aug 11, 2026Man Jiang, Ouxiang Li, Weibao Xue +4Concept ErasureText-To-Image Diffusion Models