Text-to-image diffusion models can generate prohibited content, which motivates concept erasure through machine unlearning. Most erasure methods intervene at the text interface, through prompt modification or localized updates to text-conditioning weights, and they are evaluated by what the model outputs for given prompts. Such evaluation cannot see what the network still encodes. Latent-space auditing, which bypasses text conditioning and probes the denoising network directly, shows that erased concepts remain recoverable from internal representations. We find that this also holds for methods built to be robust against adversarial prompts, and that the problem grows with the number of erased concepts. We propose Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models (AVCE), a framework that grounds erasure in the model's latent representations. AVCE audits the embedding neighborhood of each concept and condenses the discovered vulnerable directions into an anchor at the weakest geometric point. It edits cross-attention and self-attention projections in closed form at this anchor, then fine-tunes the two pathways with pathway-level auditing losses, using orthogonal gradient projection to consolidate multiple concepts. Experiments on SD v1.5, SDXL, and Flux 1.0 across object, explicit-content, and artistic-style unlearning show that AVCE reduces attack success rates by 5.07x and improves auditing scores by 3.84x over the strongest baseline, while preserving competitive generation quality.
Figures & tables
Figure 1: The illusion of robustness on SD v1.5 (erasing the concept “golf ball”). Robust methods block adversarial text inputs (left), yet latent-space auditing, which bypasses text conditioning, shows that the visual features of the target concept remain in the model (middle, right).
Figure 2: Overview of the AVCE framework. The Verifiable Vulnerability Anchor is obtained by auditing the embedding space for its weakest geometric point. Analytical Editing rewrites cross-attention via a Semantic Modulator and shrinks self-attention via a Visual Localizer. Audit-Guided Refinement closes residual leakage through pathway-level fine-tuning with multi-concept gradient consolidation.
Metric
Method
SDXL
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
Unlearn Acc ( ↓ )
87.23
77.83
68.43
64.34
64.56
0
0
72.15
69.32
8.15
CLIP ( ↑ )
31.62
31.45
31.53
31.26
31.24
22.34
22.17
30.67
30.42
30.58
FID ( ↓ )
11.92
13.87
14.14
14.53
14.76
37.42
35.68
15.02
15.39
14.87
ASR1 ( ↓ )
99.2
84.7
81.2
78.5
80.3
–
–
68.2
69.7
14.5
ASR2 ( ↓ )
97.6
88.3
85.6
82.1
84.8
–
–
73.8
75.6
16.7
Table 1: Quantitative results of multi-concept object unlearning on SDXL. Comparison of erasure accuracy, robustness, and generation quality across methods. For UCE and RECE the edited model collapses (Fig. 3 ), so attack and auditing metrics are not reported (–).
Figure 3: Visual comparison of different methods for multi-concept unlearning on SDXL.
Method
NudeNet Detection
Metric
Armpits
Belly
Buttocks
Feet
Breasts (F)
Genitalia (F)
Breasts (M)
Genitalia (M)
Total ↓
FID ↓
CLIP ↑
ESD
119
71
7
29
85
1
5
0
317
13.87
31.45
SPM
121
68
10
33
90
2
4
0
328
14.14
31.53
MACE
156
102
21
58
139
5
9
1
491
13.53
31.26
SM
115
74
9
35
92
1
6
0
332
13.76
31.24
UCE
117
70
13
41
107
1
5
0
354
14.23
31.32
Table 2: Per-category NudeNet detections on the I2P dataset for explicit-content unlearning on SDXL. “(F)” denotes female and “(M)” male.
Figure 4: Visual comparison of different methods for explicit-content unlearning on SDXL.
Metric
Method
SDXL
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
ASR1 ( ↓ )
98.4
81.3
78.5
75.2
77.1
80.8
71.6
64.7
66.3
11.8
ASR2 ( ↓ )
96.8
85.6
82.4
79.3
81.5
84.7
76.1
70.2
72.4
13.5
ASR3 ( ↓ )
97.5
89.4
86.8
84.1
85.7
88.6
83.2
77.1
78.5
14.6
CRS ( ↑ )
–
0.13
0.17
0.20
0.21
0.15
0.18
0.22
0.11
0.79
CCS ( ↓ )
–
0.87
0.81
0.76
0.73
0.84
0.70
0.65
0.72
0.21
Table 3: Quantitative results of explicit-content unlearning on SDXL. Comparison of robustness (ASR1/2/3) and auditing scores (CRS/CCS) across methods.
Figure 5: Artistic-style unlearning on SDXL. AVCE removes stylistic traces while baselines retain visible style artifacts.
Metric
Method
SDXL
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
CLIPa ( ↓ )
31.48
30.78
30.52
30.34
30.41
29.92
29.43
29.51
29.38
24.37
CLIPc ( ↑ )
31.74
31.57
31.65
31.38
31.36
31.44
31.19
30.79
30.54
30.70
CLIPd ( ↑ )
0.26
0.79
1.13
1.04
0.95
1.52
1.76
1.28
1.16
6.33
FID ( ↓ )
11.78
13.62
14.32
14.81
14.52
14.08
13.94
15.28
15.63
14.72
ASR1 ( ↓ )
98.7
81.4
78.3
75.6
77.2
80.5
71.9
65.3
66.8
11.2
Table 4: Quantitative results of artistic-style unlearning on SDXL. Comparison of erasure effectiveness, robustness, and generation quality across methods.
Table 10
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model Architecture
Number of Forgotten Concepts ( N )
N=1
N=3
N=5
N=10
N=20
SD v1.5
6.02
6.18
6.32
6.55
6.74
SDXL
15.05
15.45
15.78
16.35
16.86
Appendix
Table 7: Computational time (in minutes) of the proposed AVCE method across different numbers of concepts ( N ). Total time grows by approximately 12% from 1 to 20 concepts.
Figure 6: Scalability analysis of multi-concept object unlearning. We report Unlearn Acc ( ↓ ), Attack Success Rate ( ↓ ), Concept Confidence Score ( ↓ ), and Concept Retrieval Score ( ↑ ) across varying numbers of forgotten concepts ( N∈{5,10,15,20} ). AVCE (red triangles) remains stable as N grows, whereas the baselines degrade.
Metric
Method
SD v1.5
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
Unlearn Acc ( ↓ )
87.23
77.83
68.43
64.34
30.45
14.62
13.48
42.15
49.32
4.15
CLIP ( ↑ )
31.62
31.45
31.53
31.26
31.24
30.95
31.12
30.67
30.42
30.58
FID ( ↓ )
11.92
13.87
14.14
14.53
14.76
15.82
15.46
15.02
15.39
14.87
ASR1 ( ↓ )
99.2
84.7
81.2
78.5
76.8
74.5
71.2
68.2
69.7
13.5
ASR2 ( ↓ )
97.6
88.3
85.6
82.1
81.5
79.8
77.1
73.8
75.6
16.7
Appendix
Table 8: Quantitative results of multi-concept object unlearning on SD v1.5. Comparison of erasure accuracy, robustness, and generation quality across methods.
Metric
Method
Flux
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
Unlearn Acc ( ↓ )
89.45
86.24
83.51
79.82
81.25
77.43
75.68
82.35
80.54
16.32
CLIP ( ↑ )
32.14
31.85
31.92
31.74
31.68
31.42
31.55
31.32
31.18
31.46
FID ( ↓ )
11.05
12.95
13.24
13.58
13.72
14.85
14.52
14.15
14.36
13.92
ASR1 ( ↓ )
99.5
92.1
89.4
85.2
87.6
84.3
81.7
79.5
80.1
15.2
ASR2 ( ↓ )
98.7
94.5
91.8
88.6
90.1
87.9
85.4
83.2
84.8
17.4
Appendix
Table 9: Quantitative results of multi-concept object unlearning on Flux 1.0. Comparison of erasure accuracy, robustness, and generation quality across methods.
Metric
Method
Flux
ESD
SPM
MACE
SM
UCE
RECE
STEREO
AdvU
AVCE
ASR1 ( ↓ )
99.2
93.5
90.2
86.4
88.1
85.3
82.6
80.2
81.5
13.4
ASR2 ( ↓ )
98.5
95.8
92.7
89.5
91.2
88.6
86.3
84.1
85.7
15.1
ASR3 ( ↓ )
98.9
97.4
95.1
92.8
94.3
92.4
90.5
88.7
89.4
16.5
CRS ( ↑ )
–
0.21
0.25
0.28
0.30
0.20
0.24
0.26
0.17
0.77
CCS ( ↓ )
–
0.78
0.74
0.69
0.67
0.75
0.72
0.65
0.68
0.22
Appendix
Table 10: Quantitative results of explicit-content unlearning on Flux 1.0. Comparison of robustness (ASR1/2/3) and auditing scores (CRS/CCS) across methods.