Text-to-image diffusion models can generate prohibited content, which motivates concept erasure through machine unlearning. Most erasure methods intervene at the text interface, through prompt modification or localized updates to text-conditioning weights, and they are evaluated by what the model outputs for given prompts. Such evaluation cannot see what the network still encodes. Latent-space auditing, which bypasses text conditioning and probes the denoising network directly, shows that erased concepts remain recoverable from internal representations. We find that this also holds for methods built to be robust against adversarial prompts, and that the problem grows with the number of erased concepts. We propose Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models (AVCE), a framework that grounds erasure in the model's latent representations. AVCE audits the embedding neighborhood of each concept and condenses the discovered vulnerable directions into an anchor at the weakest geometric point. It edits cross-attention and self-attention projections in closed form at this anchor, then fine-tunes the two pathways with pathway-level auditing losses, using orthogonal gradient projection to consolidate multiple concepts. Experiments on SD v1.5, SDXL, and Flux 1.0 across object, explicit-content, and artistic-style unlearning show that AVCE reduces attack success rates by 5.07x and improves auditing scores by 3.84x over the strongest baseline, while preserving competitive generation quality.
Figures & tables
Figure 1: The illusion of robustness on SD v1.5 (erasing the concept “golf ball”). Robust methods block adversarial text inputs (left), yet latent-space auditing, which bypasses text conditioning, shows that the visual features of the target concept remain in the model (middle, right).
Figure 2: Overview of the AVCE framework. The Verifiable Vulnerability Anchor is obtained by auditing the embedding space for its weakest geometric point. Analytical Editing rewrites cross-attention via a Semantic Modulator and shrinks self-attention via a Visual Localizer. Audit-Guided Refinement closes residual leakage through pathway-level fine-tuning with multi-concept gradient consolidation.
Metric
Method
SDXL
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
Unlearn Acc ( ↓ )
87.23
77.83
68.43
64.34
64.56
0
0
72.15
69.32
8.15
CLIP ( ↑ )
31.62
31.45
31.53
31.26
31.24
22.34
22.17
30.67
30.42
30.58
FID ( ↓ )
11.92
13.87
14.14
14.53
14.76
37.42
35.68
15.02
15.39
14.87
ASR1 ( ↓ )
99.2
84.7
81.2
78.5
80.3
–
–
68.2
69.7
14.5
ASR2 ( ↓ )
97.6
88.3
85.6
82.1
84.8
–
–
73.8
75.6
16.7
Table 1: Quantitative results of multi-concept object unlearning on SDXL. Comparison of erasure accuracy, robustness, and generation quality across methods. For UCE and RECE the edited model collapses (Fig. 3 ), so attack and auditing metrics are not reported (–).
Figure 3: Visual comparison of different methods for multi-concept unlearning on SDXL.
Method
NudeNet Detection
Metric
Armpits
Belly
Buttocks
Feet
Breasts (F)
Genitalia (F)
Breasts (M)
Genitalia (M)
Total ↓
FID ↓
CLIP ↑
ESD
119
71
7
29
85
1
5
0
317
13.87
31.45
SPM
121
68
10
33
90
2
4
0
328
14.14
31.53
MACE
156
102
21
58
139
5
9
1
491
13.53
31.26
SM
115
74
9
35
92
1
6
0
332
13.76
31.24
UCE
117
70
13
41
107
1
5
0
354
14.23
31.32
Table 2: Per-category NudeNet detections on the I2P dataset for explicit-content unlearning on SDXL. “(F)” denotes female and “(M)” male.
Figure 4: Visual comparison of different methods for explicit-content unlearning on SDXL.
Metric
Method
SDXL
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
ASR1 ( ↓ )
98.4
81.3
78.5
75.2
77.1
80.8
71.6
64.7
66.3
11.8
ASR2 ( ↓ )
96.8
85.6
82.4
79.3
81.5
84.7
76.1
70.2
72.4
13.5
ASR3 ( ↓ )
97.5
89.4
86.8
84.1
85.7
88.6
83.2
77.1
78.5
14.6
CRS ( ↑ )
–
0.13
0.17
0.20
0.21
0.15
0.18
0.22
0.11
0.79
CCS ( ↓ )
–
0.87
0.81
0.76
0.73
0.84
0.70
0.65
0.72
0.21
Table 3: Quantitative results of explicit-content unlearning on SDXL. Comparison of robustness (ASR1/2/3) and auditing scores (CRS/CCS) across methods.
Figure 5: Artistic-style unlearning on SDXL. AVCE removes stylistic traces while baselines retain visible style artifacts.
Metric
Method
SDXL
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
CLIPa ( ↓ )
31.48
30.78
30.52
30.34
30.41
29.92
29.43
29.51
29.38
24.37
CLIPc ( ↑ )
31.74
31.57
31.65
31.38
31.36
31.44
31.19
30.79
30.54
30.70
CLIPd ( ↑ )
0.26
0.79
1.13
1.04
0.95
1.52
1.76
1.28
1.16
6.33
FID ( ↓ )
11.78
13.62
14.32
14.81
14.52
14.08
13.94
15.28
15.63
14.72
ASR1 ( ↓ )
98.7
81.4
78.3
75.6
77.2
80.5
71.9
65.3
66.8
11.2
Table 4: Quantitative results of artistic-style unlearning on SDXL. Comparison of erasure effectiveness, robustness, and generation quality across methods.
Table 10
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model Architecture
Number of Forgotten Concepts ( N )
N=1
N=3
N=5
N=10
N=20
SD v1.5
6.02
6.18
6.32
6.55
6.74
SDXL
15.05
15.45
15.78
16.35
16.86
Appendix
Table 7: Computational time (in minutes) of the proposed AVCE method across different numbers of concepts ( N ). Total time grows by approximately 12% from 1 to 20 concepts.
Figure 6: Scalability analysis of multi-concept object unlearning. We report Unlearn Acc ( ↓ ), Attack Success Rate ( ↓ ), Concept Confidence Score ( ↓ ), and Concept Retrieval Score ( ↑ ) across varying numbers of forgotten concepts ( N∈{5,10,15,20} ). AVCE (red triangles) remains stable as N grows, whereas the baselines degrade.
Metric
Method
SD v1.5
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
Unlearn Acc ( ↓ )
87.23
77.83
68.43
64.34
30.45
14.62
13.48
42.15
49.32
4.15
CLIP ( ↑ )
31.62
31.45
31.53
31.26
31.24
30.95
31.12
30.67
30.42
30.58
FID ( ↓ )
11.92
13.87
14.14
14.53
14.76
15.82
15.46
15.02
15.39
14.87
ASR1 ( ↓ )
99.2
84.7
81.2
78.5
76.8
74.5
71.2
68.2
69.7
13.5
ASR2 ( ↓ )
97.6
88.3
85.6
82.1
81.5
79.8
77.1
73.8
75.6
16.7
Appendix
Table 8: Quantitative results of multi-concept object unlearning on SD v1.5. Comparison of erasure accuracy, robustness, and generation quality across methods.
Metric
Method
Flux
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
Unlearn Acc ( ↓ )
89.45
86.24
83.51
79.82
81.25
77.43
75.68
82.35
80.54
16.32
CLIP ( ↑ )
32.14
31.85
31.92
31.74
31.68
31.42
31.55
31.32
31.18
31.46
FID ( ↓ )
11.05
12.95
13.24
13.58
13.72
14.85
14.52
14.15
14.36
13.92
ASR1 ( ↓ )
99.5
92.1
89.4
85.2
87.6
84.3
81.7
79.5
80.1
15.2
ASR2 ( ↓ )
98.7
94.5
91.8
88.6
90.1
87.9
85.4
83.2
84.8
17.4
Appendix
Table 9: Quantitative results of multi-concept object unlearning on Flux 1.0. Comparison of erasure accuracy, robustness, and generation quality across methods.
Metric
Method
Flux
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
ASR1 ( ↓ )
99.2
93.5
90.2
86.4
88.1
85.3
82.6
80.2
81.5
13.4
ASR2 ( ↓ )
98.5
95.8
92.7
89.5
91.2
88.6
86.3
84.1
85.7
15.1
ASR3 ( ↓ )
98.9
97.4
95.1
92.8
94.3
92.4
90.5
88.7
89.4
16.5
CRS ( ↑ )
–
0.21
0.25
0.28
0.30
0.20
0.24
0.26
0.17
0.77
CCS ( ↓ )
–
0.78
0.74
0.69
0.67
0.75
0.72
0.65
0.68
0.22
Appendix
Table 10: Quantitative results of explicit-content unlearning on Flux 1.0. Comparison of robustness (ASR1/2/3) and auditing scores (CRS/CCS) across methods.
Existing closed-form methods for concept unlearning in text-to-image diffusion models typically derive editing directions from fixed text embeddings, which may not fully capture how concepts are expressed across latent states, timesteps, and layers. To capture this variation, we investigate cross-attention activations collected during denoising. In controlled probing experiments using the same anchor prompts, activation-derived bases achieve approximately five times the recall of text-derived bases on held-out prompts expressing the target concepts. Based on this finding, we propose Cross-Attention Subspace Erasure (CASE), a closed-form method that constructs layer-specific forget and retain subspaces from cross-attention activations. These subspaces define a retain-constrained linear operator incorporated directly into cross-attention weights, requiring no gradient-based fine-tuning or additional inference-time computation. Across ten concepts spanning four categories, CASE achieves the highest harmonic-mean score among evaluated baselines in all four categories, balancing suppression, retention, adversarial robustness, and generation quality. Further experiments demonstrate robustness to recovery attacks and a favorable suppression-retention trade-off when jointly unlearning up to 100 artistic styles. The benefits of activation-derived editing also extend to larger diffusion models, including SDXL and FLUX.
Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal. We argue that this is the source of much of the collateral damage in prior preservation. We introduce CARE, a closed-form concept erasure operator that replaces the raw target direction with a kept-subspace-aware direction computed from a small bank of retained concept anchors. The resulting edit is applied directly in cross-attention value space, requires no model fine-tuning, and adds only a negligible offline computation. A single shrinkage parameter controls the erase-preserve trade-off. We further show that the operator admits a minimum-disturbance interpretation and, in its projection form, leaves the kept subspace invariant. Experiments under the standard concept-erasure protocol show that our method preserves non-target concepts more faithfully while maintaining competitive erasure across instance, style, and celebrity concepts. Code: https://github.com/parthupman/care
Parth Upman, Nishita Jain, Shreyank N Gowda
School of Computer Science, University of Nottingham, Nottingham, NG8 1BB, UK · Department of Computing, Imperial College London, London, SW7 2AZ, UK
Existing attacks on unlearned diffusion models assume that the erased concepts are known in advance and focus on recovering them. In practice, however, model providers may not disclose which concepts have been removed, and even with access to the original base model, an adversary may still lack a clear target to attack. In this paper, we aim to answer the following critical but overlooked questions: which concepts have been erased from the model, and how many have been erased in total? To this end, we present Tracer, a framework that rapidly and accurately identifies erased concepts and estimates their number. Tracer efficiently identifies erased concepts without generating and classifying images. By combining lightweight spectral analysis of weight footprints, it enables efficient search over large candidate vocabularies. To distinguish multiple erased concepts, we introduce a footprint coverage objective that guides sequential discovery. Tracer estimates the number of erased concepts by detecting a sharp decline in candidate confidence as the selected concepts account for the erasure footprint, without requiring labeled examples for calibration. The framework requires only lightweight linear algebra and limited forward probes, with no prior knowledge of the unlearning algorithm. Experiments across text-to-image and text-to-video backbones and diverse unlearning methods demonstrate that Tracer identifies erased concepts and estimates their number in seconds, achieving 150 to 137,000 times and 133 to 20,000 times speedups over MIA and brute-force search on image and video models, respectively, with substantially higher identification accuracy.
Kaiyuan Deng, Yuchen Li, Yang Xiao +3
The University of Arizona · Peking University · The University of Tulsa +1