Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-image diffusion model while preserving its ability to generate other content. Existing activation steering methods build an erasure direction mainly from the target concept and adjust model activations along it at inference time. However, target and retained concepts often overlap in the model's representation space, so this direction also contains shared components that retained concepts rely on. Steering directly along this direction can therefore suppress retained concepts and harm the generation of non-target content. To address this issue, we propose Retain-aware Activation Steering (RASteer), a training-free method. RASteer first builds a retain subspace from the concepts to preserve. Retain-Orthogonal Steering (ROS) then removes components aligned with this subspace from the erasure direction, making steering more specific to the target. Since fully removing the shared components can weaken erasure, we further introduce Overlap-Adaptive Calibration (OAC). At each layer and denoising step, OAC uses the overlap between the erasure direction and the retain subspace to control how much of each shared component is removed, balancing target erasure and concept preservation. Experiments on unsafe-content, instance, and artistic-style erasure across multiple backbones and benchmarks show that RASteer matches or outperforms the activation steering and weight editing baselines we evaluate, achieving a better balance between erasure and preservation.
Figures & tables
Figure 1: Erasure–preservation trade-offs in steering. (a) Collateral damage from original-direction steering (top) and incomplete erasure under full projection (bottom). (b) Fractions of activations above each threshold. (c) Target and mean retained CLIP accuracy for Snoopy erasure.
Figure 2: RASteer overview. Paired prompts are used to extract erasure and retain directions from the frozen diffusion model. ROS decomposes the erasure direction into retain-aligned and retain-orthogonal components, while OAC adaptively adjusts the steering strength according to their overlap. The calibrated direction is then injected into cross-attention during denoising to erase the target concept while preserving retained concepts.
Figure 3: Why calibration must be overlap-adaptive. The measured overlap o varies widely across layers and denoising steps. Ablating either module breaks one axis: without ROS the retained concept is damaged, without OAC erasure collapses; only the full method holds both.
Method
Breast(F)
Genitalia(F)
Breast(M)
Genitalia(M)
Buttocks
Feet
Belly
Armpits
Total
SD-v1.4
182
17
43
7
37
46
166
114
612
SPM
177
20
40
9
38
47
161
116
608
SAeUron
92
11
22
7
22
8
130
80
372
ESD-x
56
1
21
8
15
28
87
58
274
ESD-u
10
0
8
5
3
25
28
34
113
UCE
45
0
0
1
17
5
44
18
130
Table 1: Erasure of unsafe content on the I2P benchmark. We report the number of exposed body parts detected by NudeNet, and every method is run in one unified pipeline on SD-v1.4. Shading marks RASteer results. F: Female, M: Male.
Metric
SD-v1.4
SPM
SAeUron
ESD-x
ESD-u
UCE
MACE
RECE
CASteer
RASteer
CLIP ↑
30.96
30.93
30.42
30.21
30.03
27.79
28.83
30.59
30.88
30.95
FID ↓
–
2.57
17.49
15.40
20.51
72.02
27.31
17.36
17.84
17.47
LPIPS ↓
–
0.022
0.359
0.330
0.436
0.618
0.474
0.380
0.351
0.304
Table 2: Preservation of general content on COCO. We generate images from COCO captions with each nudity-erased model and compare them against the unsteered generations. Shading marks RASteer results.
Snoopy
Mickey
Spongebob
Pikachu
Van Gogh
Picasso
Rembrandt
Hokusai
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
SD-v1.4
31.9
98.3
0
32.0
99.6
0
31.1
99.6
0
32.1
99.6
0
31.4
99.6
0
30.7
99.2
0
33.9
96.2
0
32.3
100.0
0
Erasing Snoopy
Erasing Van Gogh
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
SPM
24.2
77.5
169
31.9
98.8
36
30.9
98.3
43
32.0
99.2
23
29.0
90.4
199
29.5
97.9
22
33.8
96.2
21
32.3
100.0
26
SAeUron
25.1
72.5
193
31.8
98.8
110
30.5
96.7
110
31.8
99.6
78
31.8
98.8
112
29.5
97.1
89
31.9
94.6
100
31.9
99.6
99
Table 3: Erasure of instances and artistic styles. Each block erases the concepts in the shaded columns and retains the others, with cartoon characters on the left and artists on the right. All methods use SD-v1.4. Blue rows mark RASteer results. Bold marks the best value and underline the second best.
Snoopy
Mickey
Spongebob
Pikachu
Van Gogh
Picasso
Rembrandt
Hokusai
Erasing Snoopy
Erasing Van Gogh
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
RASteer
21.4
10.8
209
32.0
99.6
82
31.1
98.8
84
32.1
100.0
46
23.4
32.9
311
30.0
99.2
79
32.1
97.1
101
32.4
100.0
79
w/o ROS
21.5
25.4
219
30.4
96.7
171
29.9
92.9
135
32.0
99.2
57
21.8
51.7
312
27.4
95.8
191
29.5
92.9
168
31.5
99.6
131
w/o OAC
22.5
12.1
188
32.2
98.8
50
31.2
100.0
67
32.1
100.0
41
24.8
11.2
235
30.7
99.6
44
34.0
96.2
42
32.4
100.0
52
Table 4: Component ablation on SD-v1.4. Each variant erases Snoopy and Van Gogh while retaining the other concepts; shaded columns mark the erased concepts; blue rows mark RASteer results.
Figure 4: Effect of steering layer scope on SD-v1.4. Instance and style erasure across five layer scopes: (a) target CLIP accuracy, (b) retained-concept accuracy, and (c) retained FID.
Figure 6: Instance erasure and retention on SD-v2.1. Top three rows: erased Snoopy. Bottom three rows: retained Mickey. Columns compare the unedited model, baselines, and RASteer.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Snoopy
Mickey
Spongebob
Pikachu
Van Gogh
Picasso
Rembrandt
Hokusai
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
SD
32.0
99.2
0
32.0
100.0
0
30.9
98.8
0
32.0
99.6
0
31.8
99.6
0
30.7
99.6
0
34.5
97.9
0
32.4
100.0
0
Erasing Snoopy
Erasing Van Gogh
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
SPM
24.5
85.8
170
31.9
99.6
40
31.0
97.9
45
32.0
99.6
22
30.1
92.5
197
30.7
99.6
22
34.3
97.5
21
32.4
100.0
30
SAeUron
25.1
77.5
187
31.9
100.0
99
30.9
95.8
107
31.9
99.2
68
31.2
99.6
101
30.3
97.1
72
34.1
97.9
98
32.5
99.6
96
Appendix
Table 5: Erasure of instances and artistic styles on SD-v1.5. Every method is re-run from its official code on this backbone; shaded columns mark the erased concepts; blue rows mark RASteer results. Bold marks the best value and underline the second best. RASteer uses β=2 .
Method
Snoopy
Mickey
Spongebob
Pikachu
Van Gogh
Picasso
Rembrandt
Hokusai
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
SD
32.4
99.6
0
32.2
99.2
0
29.5
98.3
0
31.7
100.0
0
32.0
100.0
0
32.2
97.5
0
33.1
96.7
0
32.2
100.0
0
Erasing Snoopy
Erasing Van Gogh
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
SPM
24.7
69.6
204
31.9
98.8
65
29.2
97.9
80
31.7
100.0
30
26.8
75.0
248
31.9
97.1
42
32.7
97.5
28
32.2
100.0
50
SAeUron
31.8
97.5
95
32.2
97.5
99
29.3
91.2
127
31.9
99.2
75
31.5
99.6
82
31.2
98.3
73
32.4
95.8
61
32.0
99.6
102
Appendix
Table 6: Erasure of instances and artistic styles on SD-v2.1. Every method is re-run from its official code on this backbone; shaded columns mark the erased concepts; blue rows mark RASteer results. Bold marks the best value and underline the second best. RASteer uses β=3 .
Method
Snoopy
Mickey
Spongebob
Pikachu
Van Gogh
Picasso
Rembrandt
Hokusai
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
CS
CA
FID
SD
32.4
100.0
0
32.0
99.2
0
30.6
93.3
0
31.2
97.5
0
28.7
95.8
0
29.4
98.8
0
29.9
95.0
0
31.7
100.0
0
Erasing Snoopy
Erasing Van Gogh
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↓
CA ↓
FID ↑
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
CS ↑
CA ↑
FID ↓
ESD-x
23.6
40.8
269
26.8
71.2
229
26.2
92.9
316
28.7
70.0
265
25.9
93.3
288
26.3
87.5
139
27.1
85.0
125
28.2
100.0
229
ESD-u
31.4
97.9
65
31.7
98.3
70
30.2
90.8
77
31.1
98.8
43
27.9
96.7
86
28.4
96.7
66
29.3
94.6
46
31.6
100.0
62
Appendix
Table 7: Erasure of instances and artistic styles on SDXL. Every method is re-run from its official code on this backbone; shaded columns mark the erased concepts; blue rows mark RASteer results. Bold marks the best value and underline the second best. RASteer uses β=3 .
Figure 7: Additional instance and style comparisons on Stable Diffusion v1.4. Each group of three rows shows representative prompts from the indicated erased or retained set.
Figure 8: Additional instance and style comparisons on Stable Diffusion v1.5. Each group of three rows shows representative prompts from the indicated erased or retained set.
Figure 9: Additional instance and style comparisons on Stable Diffusion v2.1. Each group of three rows shows representative prompts from the indicated erased or retained set.
Figure 10: Additional instance and style comparisons on SDXL. Each group of three rows shows representative prompts from the indicated erased or retained set.
Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal. We argue that this is the source of much of the collateral damage in prior preservation. We introduce CARE, a closed-form concept erasure operator that replaces the raw target direction with a kept-subspace-aware direction computed from a small bank of retained concept anchors. The resulting edit is applied directly in cross-attention value space, requires no model fine-tuning, and adds only a negligible offline computation. A single shrinkage parameter controls the erase-preserve trade-off. We further show that the operator admits a minimum-disturbance interpretation and, in its projection form, leaves the kept subspace invariant. Experiments under the standard concept-erasure protocol show that our method preserves non-target concepts more faithfully while maintaining competitive erasure across instance, style, and celebrity concepts. Code: https://github.com/parthupman/care
Parth Upman, Nishita Jain, Shreyank N Gowda
School of Computer Science, University of Nottingham, Nottingham, NG8 1BB, UK · Department of Computing, Imperial College London, London, SW7 2AZ, UK
Concept erasure removes copyright-protected, privacy-sensitive, or otherwise undesirable concepts from pretrained text-to-image diffusion models to support content governance and compliance. As erasure requests arrive over time, models must remove new targets without undoing prior erasures. Existing methods do not constrain interference across edits: residual perturbations outside the retain set interact and accumulate, degrading unrelated generations and sometimes collapsing previously erased targets into noise. We propose CEASE (Continual Erasure via Adaptive Subspace Editing), a training-free method that imposes two subspace constraints on a closed-form solver. CEASE adds the token representation of the shared replacement to the solver's invariance matrix and, when interference is detected, projects the current update onto the orthogonal complement of dominant output directions extracted from cumulative past updates. A closed-form decomposition attributes the accumulated interference to repeated activation of the shared replacement and overlap between successive update directions, showing that the two constraints suppress these respective sources. Across continual erasure of celebrities, artistic styles, and instances, CEASE achieves the most consistent erase-preserve trade-off, while existing methods either degrade general generation or insufficiently erase targets.
Yongliang Wu, Haori Lu, Jinqi Luo +3
University of Illinois Urbana-Champaign · University of Pennsylvania · National University of Singapore
Erasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, privacy violations, and offensive content. Existing approaches struggle to achieve both precise and persistent concept erasure: inaccurate localization of concept-related representations may cause unintended semantic interference, while incomplete removal of the underlying concept knowledge allows adversarial recovery. To address this dilemma, we propose PEAK, a \textbf{\textit{precise}} and \textbf{\textit{persistent}} concept erasure framework via k-Sparse Autoencoders (kSAEs). PEAK first trains a kSAE on internal activations of the diffusion denoising network to decompose dense representations into interpretable sparse features. By contrasting sparse activations induced by target and non-target prompts, PEAK identifies a compact set of target-specific features according to both activation strength and frequency. These localized features are then used for parameter optimization, where PEAK selectively suppresses target-related activations while preserving complementary non-target ones towards the original model. This feature-guided optimization embeds concept erasure directly into diffusion parameters, eliminating the need for additional inference-time intervention and facilitating effective persistence against adversarial attacks. Extensive experiments demonstrate that PEAK achieves effective and robust concept erasure. On the I2P benchmark, PEAK reduces NudeNet detections from 582 to 6, lowers the average attack success rate (ASR) from 96.52% to 5.63%, and preserves general generation quality on MS-COCO with a near-zero KID. Our code and models are available at: https://github.com/manmanTAT/PEAK
Man Jiang, Ouxiang Li, Weibao Xue +4
Hefei University of Technology · University of Science and Technology of China · University of Macau