Diffusion models excel at generating high-quality images but can memorize and reproduce harmful concepts when prompted. Although fine-tuning methods have been proposed to unlearn a target concept, they struggle to fully erase it while maintaining generation quality on other concepts, leaving models vulnerable to jailbreak attacks. Existing jailbreak methods demonstrate this vulnerability but offer limited insight into how unlearned models retain harmful concepts, limiting progress on effective defenses. In this work, we show that a coherent, interpretable, attack-accessible linear residual of the erased concept can be recovered in the token embedding space, and that both an attack and a defense follow from this structure. We introduce SubAttack, a novel jailbreaking attack that reads out this subspace by learning an orthogonal set of attack token embeddings, each being a linear combination of human-interpretable textual elements, revealing that unlearned models still retain the target concept through related textual components. Furthermore, our attack is also more powerful and transferable across text prompts, initial noises, and unlearned models than prior attacks. Conversely, projecting out the same subspace yields SubDefense, a lightweight plug-and-play defense mechanism that suppresses the residual concept in unlearned models. SubDefense provides stronger robustness than existing defenses while better preserving safe generation quality. Extensive experiments across multiple unlearning methods, concepts, and attack types demonstrate that our approach advances both understanding and mitigation of vulnerabilities in diffusion unlearning.
Figures & tables
Figure 1 : An erased concept persists as an interpretable residual subspace. SubAttack learns interpretable attack tokens {vatt,k} spanning this subspace (top); reading it out jailbreaks the unlearned model (middle), while projecting it out yields SubDefense (bottom).
Figure 2 : Learning one interpretable attack token embedding. The learning process of one attack token embedding vatt for the concept “Van Gogh” is visualized. Blue parts represent the frozen unlearned LDM, where, for simplicity, we omit the image encoder and decoder. In orange parts, it illustrates the learning mechanism for optimizing an MLP network to produce vatt , which is a linear combination of the existing token embeddings.
Figure 3 : Interpreting the attack token embeddings for concept “nudity”, “Van Gogh”, and “church”. Tokens with the largest αi are words associated with the target concept. For example, top tokens for “church” are activities conducted in the church, or names from the Bible.
Figure 4 : SubAttack jailbreaks various concepts (NSFW, style, objects) across different unlearned models (ESD, FMN, UCE, SPM). It consistently reveals the residual vulnerabilities in these models.
Figure 5 : Interpreting the subspace of attack token embeddings for concept “nudity” across different models. (a) The original LDM (i.e., SD) majorly relates it to explicit synonyms. (b-e) Unlearned LDMs more heavily associate it with implicit concepts.
Figure 6 : ESD for “garbage truck” .
Concept
ESD
FMN
UCE
SPM
Van Gogh
0.61
0.61
0.74
0.67
Church
0.76
0.85
0.79
0.82
Table 1 : CLIP similarity between residual and original explicit concept across unlearned LDMs.
Figure 7 : Transfer attack token embeddings learned by SubAttack to different unlearned models or to the original diffusion model.
Concepts:
Nudity
Van Gogh
Church
Victim Models:
FMN
UCE
SPM
FMN
UCE
SPM
FMN
UCE
SPM
NoAttack
90.00
23.00
22.56
21.56
71.44
43.78
51.56
6.55
43.78
UnlearnDiff
93.33
41.33
38.22
12.78
64.00
47.11
6.19
13.33
58.00
CCE
93.00
18.33
37.56
72.33
43.56
81.33
91.00
70.11
92.78
SubAttack (Ours)
96.89
77.00
80.44
72.67
88.89
86.89
92.89
83.77
92.00
Table 2 : Transfer attack performance of various jailbreaking methods from ESD to other models across different concepts, measured by ASR (%).
Figure 8 : SubAttack can generate the target concepts with high ASR while aligning with original text prompts . For example, our attack generates nude women with different backgrounds while CCE fails to generate the correct backgrounds.
ASR (%) ↑
Time per Data (s) ↓
Interp- retable
Inspire Defense
Concepts:
Nudity
Van Gogh
Church
Victim Models:
ESD
FMN
UCE
SPM
ESD
FMN
UCE
SPM
ESD
FMN
UCE
SPM
NoAttack
18.78
90.00
23.00
22.56
5.78
21.56
71.44
43.78
9.33
51.56
6.55
43.78
NA
NA
NA
UnlearnDiff
51.11
100.00
78.22
83.33
40.94
100.00
100.00
53.49
51.74
35.33
61.67
53.67
906.6
✗
✗
CCE
85.11
98.33
77.22
78.33
75.22
93.33
95.67
81.67
82.00
97.78
81.89
76.67
11.4
✗
✗
SubAttack (Ours)
97.56
100.00
81.67
74.89
81.00
96.33
98.33
82.78
91.33
97.78
82.67
84.89
54.2
✓
✓
Table 3 : Attack performance of various jailbreaking methods , measured by ASR (%) over 900 prompt–seed pairs per concept. Runtime includes attack construction and generation; concept-level construction for CCE and SubAttack is amortized over all 900 pairs, whereas UnlearnDiff re-optimizes each pair. Best results are highlighted in bold .
Figure 9 : Defending UCE using RECE or SubDefense across various concepts.
Figure 10 : Safe image generation after applying RECE or SubDefense.
Metrics:
UnlearnDiff ASR ↓
SubAttack ASR ↓
COCO-10k FID ↓
COCO-10k CLIP ↑
Scenarios:
SubDefense
RECE
SubDefense
RECE
SubDefense
RECE
SubDefense
RECE
Nudity
73.55%
76.44%
34.11%
62.44%
17.51
17.57
30.70
30.07
Van Gogh
52.78%
61.67%
29.44%
84.44%
16.64
17.11
30.94
30.08
Church
39.78%
50.78%
5.22%
80.33%
17.41
17.41
30.86
30.07
Table 4 : SubDefense is more robust than baseline RECE in defending three concepts on UCE against UnlearnDiff or our SubAttack, while preserving better generative quality.
Metrics:
Nudity ASR
CLIP
FID
NoAttack
UnlearnDiff
CCE
SubAttack
ESD
18.11%
51.11%
85.11%
97.56%
30.13
18.23
ESD+SubDefense
0.0%
4.56%
75.67%
42.33%
29.58
19.20
Table 5 : SubDefense can defend ESD against different kinds of attacks.
Base unlearned model
Base ASR
+ SubDefense
# blocked
Standard unlearners
ESD
97.56%
42.33% (-55.23%)
100
FMN
100%
62.89% (-37.11%)
100
UCE
81.67%
28% (-53.67%)
100
SPM
74.89%
50.78% (-24.11%)
100
Adversarially finetuned unlearner
Table 6 : SubDefense composes on top of any unlearned model. Nudity ASR under SubAttack, before and after applying SubDefense as a plug-and-play refinement, across standard unlearners and the adversarially finetuned unlearner STEREO. “# blocked” is the number of projected directions.
Appendix figures & tables44 assets
Supplementary material from the paper’s appendix.
Appendix
Scenarios:
ESD → SD
FMN → SD
UCE → SD
SPM → SD
Nudity
97.44%
97.78%
95.89%
86.11%
Van Gogh
86%
84%
88.44%
93.11%
Church
87.22%
92.56%
85.56%
84.33%
Appendix
Table 7 : Token embeddings learned by SubAttack originate from the original SD. This is evidenced by the successful transfer of attack token embeddings from unlearned models to the original SD with high ASR.
Figure 11 : Interpreting attack token embeddings for the concept “church”.
Figure 12 : Interpreting attack token embeddings for the concept “Van Gogh”.
Attacks:
NoAttack
Ours
Victim Model:
ESD
FMN
UCE
SPM
ESD
FMN
UCE
SPM
Nudity
18.78%
90%
23%
22.56%
97.56%
100.00%
81.67%
74.89%
Van Gogh
5.78%
21.56%
71.44%
43.78%
81%
96.33%
98.33%
82.78%
Church
9.33%
51.56%
6.55%
43.78%
91.33%
97.78%
82.67%
84.89%
Garbage Truck
4%
41.33%
11.33%
12.67%
31.33%
91.67%
44%
77.67%
Parachute
4%
63.67%
1.3%
30.67%
88.67%
100%
67%
97%
Appendix
Table 8 : Attack success rates (ASR) targeting different unlearned diffusion models across different concept unlearning tasks (NSFW, artist style, object).
Scenarios:
MACE (Nudity)
MACE (Truck)
MACE (Airplane)
MACE (Ship)
SA (Nudity)
AC (Van Gogh)
NoAttack
6.67%
10%
0%
6.67%
83.33%
21.67%
SubAttack (Ours)
98.33%
85.56%
96.67%
100%
98.33%
61.67%
Appendix
Table 9 : Evaluation across diverse concepts and settings including MACE, SA, and AC.
Scenarios:
Church
Garbage Truck
Parachute
Tench
SalUn (NoAttack)
1.67%
5%
5%
0%
SalUn (SubAttack)
56.67%
40%
86.67%
11.67%
EraseDiff (NoAttack)
6.67%
6.67%
3.33%
0%
EraseDiff (SubAttack)
31.67%
38.33%
78.33%
15%
Appendix
Table 10 : Attack success rates (ASR) against additional unlearned models including SalUn and EraseDiff.
Scenarios:
Nudity
Van Gogh
Church
NoAttack
3.33%
16.67%
3.33%
SubAttack
62.44%
84.44%
80.33%
Appendix
Table 11 : Attack success rates (ASR) against RECE.
Nudity
Van Gogh
Church
Scenario:
E → F
E → U
E → S
E → F
E → U
E → S
E → F
E → U
E → S
NoAttack
90%
23%
22.56%
21.56%
71.44%
43.78%
51.56%
6.55%
43.78%
UnlearnDiff
93.33%
41.33%
38.22%
12.78%
64%
47.11%
6.19%
13.33%
58%
CCE
93%
18.33%
37.56%
72.33%
43.56%
81.33%
91%
70.11%
92.78%
SubAttack (Ours)
96.89%
77%
80.44%
72.67%
88.89%
86.89%
92.89%
83.77%
92%
Appendix
Table 12 : Transfer attack success rate using different attack methods , transferring from ESD to FMN, UCE, and SPM across the three concepts. SubAttack (Ours) transfers best in almost all settings.
Scenario:
FMN->UCE
UCE->ESD
SPM->UCE
UCE->FMN
Nudity
72%
81.33%
86.11%
93.44%
Van Gogh
91.11%
48.55%
80.55%
62.55%
Church
79.33%
42.44%
68.33%
78.77%
Appendix
Table 13 : More SubAttack transfer results across four model pairs.
Attacks:
UnlearnDiff
SubAttack
Scenarios:
UCE
UCE + SubDefense
RECE
UCE
UCE + SubDefense
RECE
Nudity
78.22%
73.55% ( -4.67% )
76.44% (-1.78%)
81.67%
34.11% ( -47.56% )
62.44% (-19.23%)
Van Gogh
100%
52.78% ( -47.22% )
61.67% (-38.33%)
98.33%
29.44% ( -68.89% )
84.44% (-13.89%)
Church
61.67%
39.78% ( -64.34% )
50.78% (-10.89%)
82.67%
5.22% ( -77.45% )
80.33% (-2.34%)
Appendix
Table 14 : SubDefense is stronger than baseline RECE in defending three concepts on UCE against UnlearnDiff or our SubAttack.
Metrics:
COCO-10k FID ( ↓ )
COCO-10k CLIP ( ↑ )
Scenarios:
UCE
UCE + SubDefense
RECE
UCE
UCE + SubDefense
RECE
Nudity
17.14
17.51
17.57
30.86
30.70
30.07
Van Gogh
16.64
16.64
17.11
31.14
30.94
30.08
Church
17.84
17.41
17.41
30.95
30.86
30.07
Appendix
Table 15 : SubDefense preserves better utility than baseline RECE after defense.
ESD
FMN
UCE
SPM
Scenario:
Base
+SubDefense
Base
+SubDefense
Base
+SubDefense
Base
+SubDefense
NoAttack
20.56%
9.93% (-10.63%)
87.94%
37.59% (-50.35%)
21.98%
13.47% (-8.51%)
55.31%
34.04% (-21.27%)
UnlearnDiff
74.47%
41.13% (-33.34%)
97.87%
45.39% (-52.58%)
78.72%
45.39% (-33.33%)
91.49%
58.97% (-32.52%)
Appendix
Table 16 : SubDefense for I2P-nudity against UnlearnDiff , with 100 blocked tokens. For each unlearned model we report ASR before / after applying SubDefense (absolute change in parentheses).
Comparison
Pairs
Projection overlap
Disjoint subsets, same optimization seed
9
0.3122±0.0038
Same subset, different optimization seeds
18
0.3202±0.0036
Disjoint subsets, different optimization seeds
18
0.3118±0.0037
Independent random rank-100 subspaces
–
0.1302 (analytic expectation)
Appendix
Table 17 : Rank-100 residual-subspace stability. Mean ± sample SD across matched pairs. Independent runs recover substantially more common structure than random rank-100 subspaces.
First K directions
1
5
10
20
30
Best-of-prefix ASR
61.7±7.8
78.0±8.6
81.9±8.9
86.6±6.7
88.7±5.9
Appendix
Table 18 : Functional stability across 18 independent SubAttack runs. Mean ± SD across learned runs on the fixed 60-example protocol.
Prompt group
Rank 20
Rank 50
Cézanne
+0.415
+0.154
Monet
−0.492
−0.385
Picasso
−0.435
−0.541
Rembrandt
−0.307
+0.257
Generic painting
−0.023
−0.120
Van Gogh (target control)
−0.801
−1.193
Appendix
Table 19 : Nearby-painter preservation under Van-Gogh SubDefense. Paired change in CLIP alignment relative to the undefended model.
Optimization seed
Direct defended ASR
CLIP alignment
1024
6.67%
28.362
2026
0.00%
28.401
4096
3.33%
28.897
Mean ± SD
3.33±3.33 %
28.553±0.298
Appendix
Table 20 : SubDefense variance across independently learned rank-100 bases.
ESD
FMN
UCE
SPM
Concept:
Base
+SubDefense
Base
+SubDefense
Base
+SubDefense
Base
+SubDefense
Nudity
97.56%
42.33% (-55.23%)
100%
62.89% (-37.11%)
81.67%
28% (-53.67%)
74.89%
50.78% (-24.11%)
Van Gogh
81%
17% (-64%)
96.33%
22.78% (-73.55%)
93.78%
14.33% (-79.45%)
82.78%
12.33% (-70.45%)
Church
91.33%
40.22% (-51.11%)
82.67%
13.78% (-68.89%)
82.67%
3.22% (-79.45%)
84.89%
23.78% (-61.11%)
Appendix
Table 21 : SubDefense for three concepts against SubAttack , with 100 blocked tokens. For each unlearned model we report ASR before / after applying SubDefense (absolute change in parentheses).
Scenarios:
MACE
FMN
SPM
ESD
Ring-A-Bell ASR
11.58%
95.79%
34.74%
57.89%
+ SubDefense
5.26% (k=10)
54.75% (k=100)
14.74% (k=100)
4.21% (k=100)
Appendix
Table 22 : Exploratory defense results against the black-box Ring-A-Bell (Nudity) attack.
Scenarios:
SD
K=10
K=20
K=50
K=100
K=150
K=200
Ring-A-Bell (Nudity)
97.89%
89.47%
76.84%
60%
38.94%
23.16%
8.42%
SubAttack (Church)
100%
80%
78.33%
55%
46.67%
21.67%
10%
Appendix
Table 23 : Standalone performance of SubDefense.
Method
Learned attack
Learning time
Reuse
Application per pair
CCE
One global token (500 steps)
656 s/token
900
One diffusion generation
SubAttack
Five global directions (500 steps each)
488±5 s/direction; 2,551 s total
900
Five generations (one/direction)
UnlearnDiff
One local prompt per pair
103–1,351 s/pair in the timing study
None
Included in per-pair optimization
Appendix
Table 24 : Stage-wise attack cost on ESD–nudity. “Reuse” is the number of prompt–seed pairs over which a learned global artifact is amortized in the main comparison.
Figure 13 : ASR versus K when conducting SubAttack on ESD for the concepts “church” and “nudity”.
Scenarios:
ESD
FMN
UCE
SPM
With Orthogonality
97.56%
100%
81.67%
74.89%
Without Orthogonality
78.33%
100%
71.67%
63.33%
Appendix
Table 25 : Ablation on the orthogonality constraint. Enforcing orthogonality improves ASR across unlearned models for the “Nudity” concept.
Vocabulary size
50
500
5000 (default)
10000
Full
ASR
43.33%
81.67%
97.56%
70%
28.33%
Appendix
Table 26 : Ablation on vocabulary size. ASR of SubAttack on ESD “Nudity” with different vocabulary sizes.
#Blocked Tokens:
0
20
50
100
200
300
350
CLIP Score ( ↑ )
30.13
30.02
29.86
29.58
28.54
26.15
24.72
FID ( ↓ )
18.23
19.02
19.09
19.20
20.92
26.42
30.33
Appendix
Table 27 : SubDefense exhibits gradual degradation of CLIP score and FID when the number of blocked token embeddings increases.
#Blocked Tokens:
0
100
230
270
320
350
390
CCE ASR
85.11%
75.67%
65.78%
37.44%
28.11%
18.11%
8.89%
Appendix
Table 28 : ASR of concept “nudity” on CCE after blocking different numbers of token embeddings.
Concept:
Nudity
Van Gogh
Church
i∗
1455
668
547
# αi≥0.01
1743
1023
885
Appendix
Table 29 : Sparsity of the learned attack token embeddings.
#Itrs
1
10
30
50
70
100
130
150
170
200
i∗
1455
1799
1905
1784
1914
2062
2062
2136
2155
2115
# αi≥0.01
1743
2019
2078
2009
2206
2298
2328
2368
2358
2326
Appendix
Table 30 : Sparsity of the learned attack token embeddings decreases during the iterative subspace attack process.
Figure 14 : Attacking the concept nudity on ESD when α has different numbers of non-zero entries.
Figure 15 : More detailed visualization of COCO generation results with or without SubDefense on the concept nudity.
Figure 16 : More detailed visualization of COCO generation results with or without SubDefense on the concept Van Gogh.
Figure 17 : More detailed visualization of COCO generation results with or without SubDefense on the concept church.
Figure 18 : Visualization of “ blue ” image generation results before and after defending “Van Gogh” on UCE.
Figure 19 : Visualization of “ star ” image generation results before and after defending “Van Gogh” on UCE.
Figure 20 : Visualizing nudity attacking results on ESD.
Figure 21 : Visualizing nudity attacking results on FMN.
Figure 22 : Visualizing nudity attacking results on UCE.
Figure 23 : Visualizing nudity attacking results on SPM.
Figure 24 : Visualizing Van Gogh attacking results on ESD.
Figure 25 : Visualizing Van Gogh attacking results on FMN.
Figure 26 : Visualizing Van Gogh attacking results on UCE.
Figure 27 : Visualizing Van Gogh attacking results on SPM.
Figure 28 : Visualizing church attacking results on ESD.
Figure 29 : Visualizing church attacking results on FMN.
Figure 30 : Visualizing church attacking results on SPM.
Existing attacks on unlearned diffusion models assume that the erased concepts are known in advance and focus on recovering them. In practice, however, model providers may not disclose which concepts have been removed, and even with access to the original base model, an adversary may still lack a clear target to attack. In this paper, we aim to answer the following critical but overlooked questions: which concepts have been erased from the model, and how many have been erased in total? To this end, we present Tracer, a framework that rapidly and accurately identifies erased concepts and estimates their number. Tracer efficiently identifies erased concepts without generating and classifying images. By combining lightweight spectral analysis of weight footprints, it enables efficient search over large candidate vocabularies. To distinguish multiple erased concepts, we introduce a footprint coverage objective that guides sequential discovery. Tracer estimates the number of erased concepts by detecting a sharp decline in candidate confidence as the selected concepts account for the erasure footprint, without requiring labeled examples for calibration. The framework requires only lightweight linear algebra and limited forward probes, with no prior knowledge of the unlearning algorithm. Experiments across text-to-image and text-to-video backbones and diverse unlearning methods demonstrate that Tracer identifies erased concepts and estimates their number in seconds, achieving 150 to 137,000 times and 133 to 20,000 times speedups over MIA and brute-force search on image and video models, respectively, with substantially higher identification accuracy.
Kaiyuan Deng, Yuchen Li, Yang Xiao +3
The University of Arizona · Peking University · The University of Tulsa +1
Concept unlearning in text-to-image diffusion models aims to suppress a target concept (e.g., \texttt{horse}) while preserving related but distinct content (e.g., \texttt{donkey}), yet existing methods either leak under indirect prompts or visibly degrade other concepts. We show that these failure modes stem from the geometry of concept representations rather than from any particular algorithm. Formalizing concepts as activation-space regions, we prove that the overlap between a target and other concepts lower-bounds the damage any robust erasure must inflict on them, with the trade-off scaling linearly in the degree of overlap. Across thirteen unlearning methods, including methods designed to preserve non-target concepts, no method achieves both strong erasure and strong neighbor preservation: STEREO nearly eliminates indirect leakage but cuts neighbor generation by more than 75%, while sparse inference-time methods preserve neighbors but leak. Damage increases with our overlap measure, monotonically so for STEREO; the κ-scaling reproduces on SDXL, and neighbor-selective damage recurs on FLUX. Perfect unlearning is the wrong target for entangled concepts; methods should be evaluated on the Pareto frontier our theorem establishes.
Yian Wang, Ali Ebrahimpour-Boroojeny, Hari Sundaram +1
University of Illinois Urbana-Champaign Urbana, IL, USA
Machine unlearning for text-to-image diffusion models aims to selectively remove undesirable concepts from pre-trained models without costly retraining. Current unlearning methods share a common weakness: erased concepts return when the model is fine-tuned on downstream data, even when that data is entirely unrelated. We adapt Projected Gradient Unlearning (PGU) from classification to the diffusion domain as a post-hoc hardening step. By constructing a Core Gradient Space (CGS) from the retain concept activations and projecting gradient updates into its orthogonal complement, PGU ensures that subsequent fine-tuning cannot undo the achieved erasure. Applied on top of existing methods (ESD, UCE, Receler), the approach eliminates revival for style concepts and substantially delays it for object concepts, running in roughly 6 minutes versus the ~2 hours required by Meta-Unlearning. PGU and Meta-Unlearning turn out to be complementary: which performs better depends on how the concept is encoded, and retain concept selection should follow visual feature similarity rather than semantic grouping.
Aljalila Aladawi, Mohammed Talha Alam, Fakhri Karray
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE. · University of Waterloo, Ontario, Canada.