Concept erasure removes copyright-protected, privacy-sensitive, or otherwise undesirable concepts from pretrained text-to-image diffusion models to support content governance and compliance. As erasure requests arrive over time, models must remove new targets without undoing prior erasures. Existing methods do not constrain interference across edits: residual perturbations outside the retain set interact and accumulate, degrading unrelated generations and sometimes collapsing previously erased targets into noise. We propose CEASE (Continual Erasure via Adaptive Subspace Editing), a training-free method that imposes two subspace constraints on a closed-form solver. CEASE adds the token representation of the shared replacement to the solver's invariance matrix and, when interference is detected, projects the current update onto the orthogonal complement of dominant output directions extracted from cumulative past updates. A closed-form decomposition attributes the accumulated interference to repeated activation of the shared replacement and overlap between successive update directions, showing that the two constraints suppress these respective sources. Across continual erasure of celebrities, artistic styles, and instances, CEASE achieves the most consistent erase-preserve trade-off, while existing methods either degrade general generation or insufficiently erase targets.
Figures & tables
Figure 1: Utility loss under continual celebrity erasure. (a) COCO FID on 1,000 unrelated prompts as the number of erased concepts grows from 10 to 100; lower is better. (b) COCO CLIP score on the same prompts after 100 edits; the dashed line marks the unedited model. (c) Generations after 100 edits for an unrelated COCO scene (top) and the erased target “Tom Cruise” (bottom).
Figure 2: Overview of CEASE for continual concept erasure. AIC preserves the shared anchor during the closed-form solve. HOC extracts dominant output directions from cumulative historical edits via SVD and, when needed, projects the new update onto their orthogonal complement before applying it to the current model.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Acc ↓
CS ↑
FID ↓
LPIPS ↓
Acc ↑
CS ↑
FID ↓
LPIPS ↓
SD v1.4
30.95
0.0
0.000
88.0
31.66
0.0
0.000
70.6
26.52
0.0
0.000
ESD
21.64
61.6
0.545
23.0
21.31
59.3
0.544
12.9
21.96
61.5
0.540
UCE
13.55
272.8
0.768
0.0
14.10
242.7
0.768
0.0
13.15
241.4
0.773
ConAbl
15.79
74.1
0.540
0.0
15.85
73.0
0.537
0.0
24.26
50.7
0.476
RECE
16.51
568.3
0.974
0.0
16.27
573.1
0.972
0.0
16.28
616.5
1.002
Table 1: Celebrity erasure after continually erasing 100 celebrities on SD v1.4. Bold and underline mark the best and second-best edited results.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Aes ↑
CS ↑
FID ↓
LPIPS ↓
Aes ↑
CS ↑
FID ↓
LPIPS ↓
SD v1.4
28.31
0.0
0.000
6.01
28.20
0.0
0.000
5.99
26.52
0.0
0.000
ESD
19.98
85.2
0.617
5.36
20.12
84.2
0.618
5.39
23.50
56.7
0.517
UCE
17.28
250.0
0.772
4.27
15.83
210.5
0.768
3.93
13.97
228.0
0.751
ConAbl
20.27
127.1
0.577
5.52
20.16
125.3
0.581
5.51
22.98
57.7
0.516
RECE
12.75
388.7
0.916
3.13
12.41
365.3
0.909
3.20
13.79
371.6
0.904
Table 2: Style erasure after continually erasing 100 artistic styles on SD v1.4. Bold and underline mark the best and second-best edited results.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Acc ↓
CS ↑
FID ↓
LPIPS ↓
Acc ↑
CS ↑
FID ↓
LPIPS ↓
SD v1.4
28.02
0.0
0.000
80.2
28.46
0.0
0.000
78.6
26.52
0.0
0.000
ESD
20.08
65.1
0.621
28.9
20.22
60.1
0.604
27.7
21.82
64.2
0.555
UCE
15.65
319.8
0.790
1.2
16.13
271.7
0.771
0.2
14.20
276.3
0.750
ConAbl
15.20
179.7
0.653
1.8
15.13
176.7
0.651
2.5
19.13
89.9
0.575
RECE
16.22
438.6
0.891
1.0
15.90
433.5
0.897
0.0
13.46
427.6
0.881
Table 3: Instance erasure after continually erasing 100 instances on SD v1.4. Bold and underline mark the best and second-best edited results.
Figure 3: Qualitative comparison after the 100-step continual erasure, on the erased target Hillary Clinton, the retained concept Jason Statham, and an unrelated MS-COCO prompt. CEASE redirects the target to a coherent anchor substitute while preserving the retained concept and the COCO scene.
Module
Erase set
Retain set
MS-COCO
AIC
HOC
CS ↓
FID ↑
LPIPS ↑
Acc ↓
CS ↑
FID ↓
LPIPS ↓
Acc ↑
CS ↑
FID ↓
LPIPS ↓
×
×
18.38
137.2
0.635
0.3
30.28
20.8
0.383
54.4
22.14
107.2
0.566
×
✓
20.84
57.8
0.528
1.5
31.44
13.5
0.298
63.7
25.54
59.7
0.457
✓
×
21.41
43.1
0.485
3.5
31.54
12.9
0.284
65.2
26.23
48.6
0.399
✓
✓
21.54
39.9
0.478
3.5
31.59
12.3
0.271
65.9
26.24
46.2
0.388
Table 4: Ablation of AIC and HOC on the celebrity task. Combining both modules improves retained-concept preservation and general generation quality. Bold marks the best results.
Figure 4: Cross-attention layer-scope ablation on the 100-step celebrity sequence. We compare four editing scopes using erase and retain CLIP scores and recognition accuracy, together with COCO CLIP score and FID. Arrows indicate the preferred direction.
Figure 5: Cross-edit interference over 100 celebrity erasures, comparing SPEED, RECE, UCE, and CEASE. The three panels track noise on unrelated COCO prompts, the shared replacement “person”, and the erased target “Tom Cruise” as successive edits accumulate.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Acc ↓
CS ↑
FID ↓
LPIPS ↓
Acc ↑
CS ↑
FID ↓
LPIPS ↓
SD v1.5
31.13
0.0
0.000
88.8
31.92
0.0
0.000
70.6
26.43
0.0
0.000
ESD
28.74
57.4
0.413
79.2
28.86
56.5
0.415
58.2
24.82
46.1
0.446
UCE
13.08
363.3
0.818
0.0
13.63
293.9
0.795
0.0
13.10
294.4
0.788
ConAbl
20.04
98.5
0.454
2.0
21.35
92.2
0.447
1.8
25.72
37.3
0.291
RECE
16.51
564.4
0.977
0.0
16.27
572.0
0.975
0.0
16.28
610.4
1.002
Appendix
Table 5: Celebrity erasure after continually erasing 100 celebrities on SD v1.5. Bold and underline mark the best and second-best edited results.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Aes ↑
CS ↑
FID ↓
LPIPS ↓
Aes ↑
CS ↑
FID ↓
LPIPS ↓
SD v1.5
28.34
0.0
0.000
6.01
28.24
0.0
0.000
6.00
26.43
0.0
0.000
ESD
24.50
50.9
0.485
5.76
24.53
49.4
0.485
5.77
25.71
42.9
0.388
UCE
17.09
227.8
0.860
4.59
15.41
182.4
0.795
4.24
13.70
210.2
0.776
ConAbl
23.70
66.4
0.468
5.67
24.01
61.3
0.468
5.66
25.37
40.6
0.355
RECE
14.60
663.7
0.996
4.60
14.05
668.9
0.997
4.60
16.28
610.4
1.002
Appendix
Table 6: Style erasure after continually erasing 100 artistic styles on SD v1.5. Bold and underline mark the best and second-best edited results.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Acc ↓
CS ↑
FID ↓
LPIPS ↓
Acc ↑
CS ↑
FID ↓
LPIPS ↓
SD v1.5
28.05
0.0
0.000
79.9
28.52
0.0
0.000
78.9
26.43
0.0
0.000
ESD
24.28
38.9
0.535
58.6
24.60
36.0
0.517
58.0
24.50
46.6
0.440
UCE
14.97
301.0
0.809
1.6
15.25
247.8
0.791
0.2
14.03
264.1
0.758
ConAbl
19.76
74.4
0.567
15.8
20.36
71.2
0.562
18.8
24.89
46.8
0.411
RECE
17.23
606.3
1.006
0.0
17.12
594.9
1.006
1.0
16.28
610.4
1.002
Appendix
Table 7: Instance erasure after continually erasing 100 instances on SD v1.5. Bold and underline mark the best and second-best edited results.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Acc ↓
CS ↑
FID ↓
LPIPS ↓
Acc ↑
CS ↑
FID ↓
LPIPS ↓
SD v2.1
30.60
0.0
0.000
88.8
31.03
0.0
0.000
75.6
26.32
0.0
0.000
ESD
25.85
103.5
0.520
62.0
25.48
109.2
0.523
41.8
23.43
52.8
0.481
UCE
27.82
47.1
0.398
47.0
30.74
42.2
0.362
68.2
26.11
48.5
0.442
ConAbl
15.63
214.3
0.603
0.0
15.68
206.1
0.604
0.0
24.60
47.9
0.451
RECE
16.51
549.3
0.973
0.0
16.27
562.3
0.971
0.0
16.28
614.3
1.002
Appendix
Table 8: Celebrity erasure after continually erasing 100 celebrities on SD v2.1. Bold and underline mark the best and second-best edited results.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Aes ↑
CS ↑
FID ↓
LPIPS ↓
Aes ↑
CS ↑
FID ↓
LPIPS ↓
SD v2.1
27.26
0.0
0.000
6.01
27.37
0.0
0.000
5.99
26.32
0.0
0.000
ESD
25.71
41.6
0.400
5.95
25.69
41.8
0.399
5.94
25.45
41.4
0.366
UCE
25.89
28.7
0.316
5.99
26.08
28.0
0.303
5.97
25.91
40.3
0.339
ConAbl
25.25
43.2
0.424
5.84
25.49
40.9
0.415
5.83
25.51
41.0
0.350
RECE
12.45
442.0
0.866
3.48
12.14
435.4
0.862
3.43
14.47
403.6
0.878
Appendix
Table 9: Style erasure after continually erasing 100 artistic styles on SD v2.1. Bold and underline mark the best and second-best edited results.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
Acc ↓
CS ↑
FID ↓
LPIPS ↓
Acc ↑
CS ↑
FID ↓
LPIPS ↓
SD v2.1
26.39
0.0
0.000
74.3
26.43
0.0
0.000
70.2
26.32
0.0
0.000
ESD
25.64
21.6
0.395
69.2
25.51
19.8
0.386
64.4
25.72
40.1
0.347
UCE
17.50
125.5
0.666
5.8
19.15
76.9
0.617
19.8
18.43
135.4
0.612
ConAbl
20.67
60.4
0.567
27.2
20.72
60.5
0.567
27.0
25.17
45.4
0.396
RECE
16.66
351.8
0.883
1.0
16.44
327.3
0.874
0.0
13.35
419.1
0.880
Appendix
Table 10: Instance erasure after continually erasing 100 instances on SD v2.1. Bold and underline mark the best and second-best edited results.
Erase set
Retain set
MS-COCO
CS ↓
FID ↑
LPIPS ↑
CS ↑
FID ↓
LPIPS ↓
CS ↑
FID ↓
LPIPS ↓
Ours
21.35 ± 0.06
81.5 ± 0.8
0.543 ± 0.002
31.57 ± 0.08
48.9 ± 1.0
0.428 ± 0.003
26.22 ± 0.05
46.3 ± 0.7
0.387 ± 0.009
Appendix
Table 11: Erasure-order robustness on the celebrity task: mean ± std over five random orderings of the 100 targets, using five templates and one image-sampling seed per template. The editing seed is fixed at 0, and CEASE includes the same retain augmentation as in the main table.
Figure 6: Single-request wall time on one H100 GPU. CEASE stays in the seconds-level regime of the closed-form methods, one to two orders of magnitude below the training-based baselines.
Figure 7: Additional celebrity examples after the full 100-step sequence on Stable Diffusion v1.4.
Figure 8: Additional style examples after the full 100-step sequence on Stable Diffusion v1.4.
Figure 9: Additional instance examples after the full 100-step sequence on Stable Diffusion v1.4.
Figure 10: Celebrity examples after the full 100-step sequence on Stable Diffusion v1.5.
Figure 11: Celebrity examples after the full 100-step sequence on Stable Diffusion v2.1.
Figure 12: Completed candidate comparisons on the common value-projection diagnostic, including methods with lower noise than CEASE. The one-epoch RECE configuration and UCE sequence are defined above.
Figure 13: Scaling on the celebrity benchmark at N∈{10,25,50,75,100} . (a) COCO CLIP score on 1,000 prompts; the gray line is the unedited SD reference. (b) Target-averaged value-projection diagnostic on a logarithmic axis. Markers denote measured checkpoints; COCO FID at the same checkpoints is shown in Fig. 1 (a).
Figure 14: COCO and erased-concept comparisons for SPEED, RECE, UCE, and CEASE after 100 celebrity erasures. Every row uses the same prompt and random seed with FP16 sampling, and erased-target sampling retains the original batch positions. RECE uses one adversarial epoch per request.
Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal. We argue that this is the source of much of the collateral damage in prior preservation. We introduce CARE, a closed-form concept erasure operator that replaces the raw target direction with a kept-subspace-aware direction computed from a small bank of retained concept anchors. The resulting edit is applied directly in cross-attention value space, requires no model fine-tuning, and adds only a negligible offline computation. A single shrinkage parameter controls the erase-preserve trade-off. We further show that the operator admits a minimum-disturbance interpretation and, in its projection form, leaves the kept subspace invariant. Experiments under the standard concept-erasure protocol show that our method preserves non-target concepts more faithfully while maintaining competitive erasure across instance, style, and celebrity concepts. Code: https://github.com/parthupman/care
Parth Upman, Nishita Jain, Shreyank N Gowda
School of Computer Science, University of Nottingham, Nottingham, NG8 1BB, UK · Department of Computing, Imperial College London, London, SW7 2AZ, UK
Erasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, privacy violations, and offensive content. Existing approaches struggle to achieve both precise and persistent concept erasure: inaccurate localization of concept-related representations may cause unintended semantic interference, while incomplete removal of the underlying concept knowledge allows adversarial recovery. To address this dilemma, we propose PEAK, a \textbf{\textit{precise}} and \textbf{\textit{persistent}} concept erasure framework via k-Sparse Autoencoders (kSAEs). PEAK first trains a kSAE on internal activations of the diffusion denoising network to decompose dense representations into interpretable sparse features. By contrasting sparse activations induced by target and non-target prompts, PEAK identifies a compact set of target-specific features according to both activation strength and frequency. These localized features are then used for parameter optimization, where PEAK selectively suppresses target-related activations while preserving complementary non-target ones towards the original model. This feature-guided optimization embeds concept erasure directly into diffusion parameters, eliminating the need for additional inference-time intervention and facilitating effective persistence against adversarial attacks. Extensive experiments demonstrate that PEAK achieves effective and robust concept erasure. On the I2P benchmark, PEAK reduces NudeNet detections from 582 to 6, lowers the average attack success rate (ASR) from 96.52% to 5.63%, and preserves general generation quality on MS-COCO with a near-zero KID. Our code and models are available at: https://github.com/manmanTAT/PEAK
Man Jiang, Ouxiang Li, Weibao Xue +4
Hefei University of Technology · University of Science and Technology of China · University of Macau
Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while preserving model utility on benign concepts. Current CETs face a trade-off between erasure robustness and utility: stronger edits erase the target more reliably but degrade utility on non-target concepts, and vice versa. This stems from how existing methods define what to erase and what to preserve. Many CETs rely on static concept banks specified manually, generated by LLMs, or selected by CLIP image-text similarity. Such banks do not model how prompts steer the model during denoising, leaving it vulnerable to triggers that reintroduce the target while suppressing nearby benign concepts. We present Preservation-aware Adaptive Ranked Subspace Expansion (PARSE), a training-free framework for robust concept erasure in latent diffusion models. Given a target, PARSE queries the diffusion model with classifier-free guidance to dynamically discover target-inducing erase concepts and nearby retain concepts in the model vocabulary. It then edits the cross-attention value space with a preservation-aware projection that removes target directions while leaving retain directions intact. For triggers beyond this vocabulary-indexed space, PARSE iteratively searches for re-emergence triggers by textual inversion and adaptively expands the erased subspace only when a new trigger direction does not conflict with retain semantics. We also introduce the Balanced Erasure Utility Score (BEUS), which combines robustness (ASR under multiple attacks) and utility preservation (FID) via bounded monotone transforms and harmonic mean aggregation. Experiments on NSFW, artistic style, and object erasure, with a large-scale robustness-utility analysis over many CET baselines, show that PARSE erases multiple concepts robustly without sacrificing post-edit utility.