Removing copyrighted, unsafe, or user-specified concepts from a deployed text-to-image diffusion model is now a practical requirement. Weight-editing methods can suppress fixed targets, but they require per-target retraining and modify the model checkpoint. Training-free methods, on the other hand, are deployment-friendly, but they suffer from text-side routing failures on compositional prompts. In such prompts, the erase target may be invoked through a related class rather than its lexical name, and its modifiers may migrate onto preserved objects. This paper proposes Hierarchically Grounded Semantic Surgery (HGSS), a training-free framework for compositional concept erasure. The framework lifts both the routing signal and the edit operator used by text-side erasure. First, hierarchical span grounding resolves erase-target spans through lexical, taxonomic, and semantic evidence, while guarding against broad-hypernym and compound-head false positives. Second, dynamic attribute binding refines the text conditioning during early denoising via a counterfactual reference and a preserve-aware cross-attention objective, keeping surviving attribute-noun bindings intact. HGSS selectively removes the erase target without updating model weights or adding learned parameters. On SEE, HGSS cuts hierarchical evasion from 29.54 to 10.02 and roughly halves pairwise attribute leakage, achieving the best Neighbor E and AttrP scores among the reported erasure methods. On UnlearnCanvas, HGSS slightly improves the six-metric average over the matched Semantic Surgery baseline, reaching state-of-the-art.
Figures & tables
Figure 1: Two motivating failures in compositional concept erasure. Top: under erase target “mammal”, prior methods can retain a hierarchical descendant. Bottom: under erase target “car”, ESD can retain the target, while SS removes it but transfers the target attribute to the preserved tree. HGSS removes the target concept while preserving non-target content.
Figure 2: HGSS pipeline. HGSS surrounds the SS-inherited weighted coarse edit (Eq. ( 1 )) with two training-free modules. Hierarchical Span Grounding produces the weighted edit plan from lexical, taxonomic, and semantic evidence, gated by broad-hypernym and compound-head guards (Sec. 3.3 ). Dynamic Attribute Binding stabilises the preserve-side bindings exposed by the resolver-expanded erase set, refining the live conditioning during the first 0.6 fraction of DDIM iterations via an asymmetric preserve-aware cross-attention objective (Sec. 3.4 ). The U-Net and the CLIP text encoder remain frozen throughout.
Method
Neighbor
Neighbor
Evas.
Attr E
Attr P
E↓
P↑
Avg ↓
/CLIP ↓
/CLIP ↓
Unedited SD v1.4 [ Rombach et al.(2022)Rombach, Blattmann, Lorenz, Esser, and Ommer ]
92.70
92.17
91.65
92.21
35.01
Weight-edit / fine-tuning baselines
UCE [ Gandikota et al.(2024)Gandikota, Orgad, Belinkov, Materzyńska, and Bau ]
30.00
66.85
84.25
31.56
52.14
RECE [ Gong et al.(2024)Gong, Chen, Wei, Chen, and Jiang ]
23.08
57.62
83.49
24.43
57.26
MACE [ Lu et al.(2024)Lu, Wang, Li, Liu, and Kong ]
28.68
55.47
81.34
29.33
58.87
Table 1: SEE aggregate results in the original paper direction. Neighbor E / P metrics from SEE Table 2; Evasion Avg. from SEE Table 3; Attr. Leakage columns from SEE Table 4 (pairwise CLIP without substring matching). Lower is better for E /Evas./Leak., higher is better for P . Best bold , second-best underlined , excluding the unedited SD v1.4 reference row. SS and HGSS rows use matched runs under the same evaluation protocol.
Method
Vehicle
Outdoor
Animal
Access.
Sports
Kitchen
Food
Furn.
Elec.
Appl.
Indoor
Avg
Unedited SD
95.65
91.04
92.56
91.72
88.29
94.52
94.31
96.97
86.02
91.04
85.99
91.65
UCE [ Gandikota et al.(2024)Gandikota, Orgad, Belinkov, Materzyńska, and Bau ]
94.19
89.54
94.81
81.60
83.43
63.12
89.83
97.02
81.05
90.55
61.56
84.25
RECE [ Gong et al.(2024)Gong, Chen, Wei, Chen, and Jiang ]
95.39
93.28
91.86
75.40
81.04
62.25
93.83
96.15
78.63
88.36
62.17
83.49
MACE [ Lu et al.(2024)Lu, Wang, Li, Liu, and Kong ]
91.55
88.93
89.68
77.88
82.02
58.87
88.01
93.64
78.91
87.83
57.38
81.34
SPM [ Lyu et al.(2024)Lyu, Yang, Hong, Chen, Jin, He, Xue, Han, and Ding ]
94.82
91.18
92.71
79.13
84.51
65.70
87.93
89.43
79.92
90.97
59.92
83.29
ESD [ Gandikota et al.(2023)Gandikota, Materzyńska, Fiotto-Kaufman, and Bau ]
94.00
91.50
94.20
81.30
84.80
66.10
90.10
90.60
81.20
90.90
62.20
84.26
Table 2: Per-category SEE Table 3 (evasion). Lower is better. Per-column best is bold , second-best underlined ; HGSS is the column-best on every super-category and on the average.
Figure 3: Qualitative analysis on four SEE prompts. Rows 1–2 show hierarchical evasion under erase target “mammal”: ESD and SS retain mammal descendants, while HGSS suppresses them. Rows 3–4 show car erasure with colour attributes under erase target “car”. SS can retain the car or transfer colour to preserved objects, while HGSS removes the car and preserves the non-target scene.
Style Unlearning
Object Unlearning
Effect.
Efficiency
Method
UA
IRA
CRA
UA
IRA
CRA
Avg.
FID
Mem.
Stor.
ESD [ Gandikota et al.(2023)Gandikota, Materzyńska, Fiotto-Kaufman, and Bau ]
98.58
80.97
93.96
92.15
55.78
44.23
77.61
65.55
17.8
4.3
FMN [ Zhang et al.(2024a)Zhang, Wang, Xu, Wang, and Shi ]
88.48
56.77
46.60
45.64
90.63
73.46
66.93
131.37
17.9
4.2
UCE [ Gandikota et al.(2024)Gandikota, Orgad, Belinkov, Materzyńska, and Bau ]
98.40
60.22
47.71
94.31
39.35
34.67
62.45
182.01
5.1
1.7
CA [ Kumari et al.(2023)Kumari, Zhang, Wang, Shechtman, Zhang, and Zhu ]
60.82
96.01
92.70
46.67
90.11
81.97
78.05
54.21
10.1
4.2
SalUn [ Fan et al.(2024)Fan, Liu, Zhang, Wong, Wei, and Liu ]
86.26
90.39
95.08
86.91
96.35
99.59
92.43
61.05
30.8
4.0
Table 3: UnlearnCanvas comparison. ↑ : UA/IRA/CRA, Avg.; ↓ : FID, Mem. (GB), Storage (GB). SS and HGSS use matched runs on the full UC set ( 50 style +20 object targets). Best bold per column; baseline FIDs are source-reported under a different reference set, so no FID best is marked.
Method
Regime
Total ↓
FID ↓
CLIP ↑
SD v1.4 [ Rombach et al.(2022)Rombach, Blattmann, Lorenz, Esser, and Ommer ]
Base
751
14.04
31.34
ESD-u [ Gandikota et al.(2023)Gandikota, Materzyńska, Fiotto-Kaufman, and Bau ]
Weight-edit
55
15.10
30.21
UCE [ Gandikota et al.(2024)Gandikota, Orgad, Belinkov, Materzyńska, and Bau ]
Weight-edit
165
14.07
30.85
MACE [ Lu et al.(2024)Lu, Wang, Li, Liu, and Kong ]
Weight-edit
123
13.42
29.41
Receler [ Huang et al.(2024)Huang, Chang, Tsai, Lai, Yang, and Wang ]
Weight-edit
159
14.10
31.02
SAFREE [ Yoon et al.(2025)Yoon, Yu, Patil, Yao, and Bansal ]
Training-free
82
–
–
Table 4: Direct lexical safety erasure on I2P (representative methods). Total = NudeNet detections over eight body-part classes on 4,703 I2P prompts ( paper8 + threshold=0.6 ; lower is safer). FID, CLIP are MS-COCO-30K. SD v1.4 is no-intervention; bold excludes SD. Full 12 -method breakdown in App. C.3, Suppl. Table 2.
Variant
Neighbor E↓
Neighbor P↑
Evasion ↓
Attr E↓
Attr P↓
Base Coarse Surgery
25.30
31.22
29.54
80.41
82.22
Base + Hierarchical Span Grounding
9.88
21.39
12.19
35.21
42.44
Base + Dynamic Attribute Binding
24.71
31.31
29.40
71.96
72.01
HGSS full
8.08
30.04
10.02
41.44
38.30
Table 5: Full SEE module ablation at β=−0.12 . Lower is better for Neighbor E , Evasion, Attr E , and Attr P ; higher is better for Neighbor P . Best per column in bold .
Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal. We argue that this is the source of much of the collateral damage in prior preservation. We introduce CARE, a closed-form concept erasure operator that replaces the raw target direction with a kept-subspace-aware direction computed from a small bank of retained concept anchors. The resulting edit is applied directly in cross-attention value space, requires no model fine-tuning, and adds only a negligible offline computation. A single shrinkage parameter controls the erase-preserve trade-off. We further show that the operator admits a minimum-disturbance interpretation and, in its projection form, leaves the kept subspace invariant. Experiments under the standard concept-erasure protocol show that our method preserves non-target concepts more faithfully while maintaining competitive erasure across instance, style, and celebrity concepts. Code: https://github.com/parthupman/care
Parth Upman, Nishita Jain, Shreyank N Gowda
School of Computer Science, University of Nottingham, Nottingham, NG8 1BB, UK · Department of Computing, Imperial College London, London, SW7 2AZ, UK
Concept erasure removes copyright-protected, privacy-sensitive, or otherwise undesirable concepts from pretrained text-to-image diffusion models to support content governance and compliance. As erasure requests arrive over time, models must remove new targets without undoing prior erasures. Existing methods do not constrain interference across edits: residual perturbations outside the retain set interact and accumulate, degrading unrelated generations and sometimes collapsing previously erased targets into noise. We propose CEASE (Continual Erasure via Adaptive Subspace Editing), a training-free method that imposes two subspace constraints on a closed-form solver. CEASE adds the token representation of the shared replacement to the solver's invariance matrix and, when interference is detected, projects the current update onto the orthogonal complement of dominant output directions extracted from cumulative past updates. A closed-form decomposition attributes the accumulated interference to repeated activation of the shared replacement and overlap between successive update directions, showing that the two constraints suppress these respective sources. Across continual erasure of celebrities, artistic styles, and instances, CEASE achieves the most consistent erase-preserve trade-off, while existing methods either degrade general generation or insufficiently erase targets.
Yongliang Wu, Haori Lu, Jinqi Luo +3
University of Illinois Urbana-Champaign · University of Pennsylvania · National University of Singapore
Text-to-image (T2I) diffusion models inevitably internalize sensitive or non-compliant concepts from large-scale pretraining data, necessitating post-hoc concept erasure. However, existing erasure methods often lack explicit constraints on parameter updates, leading to over-intervention and unintended semantic drift. In addition, many methods rely on manually crafted counterfactual supervision, such as surrogate prompts, which incurs substantial data construction costs that limit scalability to new concepts. To address these limitations, we propose GRACE, a structured concept erasure framework designed to enable localized and selective intervention. Specifically, we introduce a semantically weighted sensitive subspace estimation to precisely lock intervention directions, and employ lightweight subspace-constrained adapters to prevent global semantic disturbance. To eliminate the dependency on manual prompt engineering, we design an automatically decoupled safe-anchor mechanism. To mitigate semantic drift induced by excessive intervention, we introduce an energy-driven dynamic gating mechanism that adaptively controls the timing and strength of intervention at inference. Extensive experiments demonstrate that our method achieves a superior balance between erasure effectiveness and generation fidelity. Compared with the average performance of five state-of-the-art (SOTA) concept erasure methods, our method improves the fine-grained NSFW reduction rate by 17.86%, while reducing the macro-averaged target CLIP Score and preservation-oriented Fr'echet Inception Distance (FID) by 4.75% and 50.58%, respectively, indicating stronger concept suppression with substantially improved preservation of the original model's generative utility.
Qinghui Gong, Yihuai Liang, Yuanlun Xie +3
School of Information Science and Technology, Southwest Jiaotong University, Chengdu 611756, China · School of Electronic Information and Electrical Engineering, Chengdu University, Chengdu, China · Key Laboratory of Intelligent Control and Optimization for Industrial Equipment of Ministry of Education, Dalian University of Technology, Dalian, China +1