Concept unlearning in text-to-image diffusion models aims to suppress a target concept (e.g., \texttt{horse}) while preserving related but distinct content (e.g., \texttt{donkey}), yet existing methods either leak under indirect prompts or visibly degrade other concepts. We show that these failure modes stem from the geometry of concept representations rather than from any particular algorithm. Formalizing concepts as activation-space regions, we prove that the overlap between a target and other concepts lower-bounds the damage any robust erasure must inflict on them, with the trade-off scaling linearly in the degree of overlap. Across thirteen unlearning methods, including methods designed to preserve non-target concepts, no method achieves both strong erasure and strong neighbor preservation: STEREO nearly eliminates indirect leakage but cuts neighbor generation by more than 75%, while sparse inference-time methods preserve neighbors but leak. Damage increases with our overlap measure, monotonically so for STEREO; the κ-scaling reproduces on SDXL, and neighbor-selective damage recurs on FLUX. Perfect unlearning is the wrong target for entangled concepts; methods should be evaluated on the Pareto frontier our theorem establishes.
Figures & tables
Figure 1: Concepts occupy coherent, partially overlapping regions in activation space. Top-left: UMAP projection of cross-attention activations at mid_block.attentions.0 (step 25/50) for 10 concepts in Stable Diffusion v1.4 (visualization pool; full evaluation pool in Section I.3 ); each point is one prompt and ellipses show covariance contours. Semantically related concepts ( horse / pony / donkey ) cluster together, while isolated concepts ( castle ) are well-separated, at the bottom left of the figure. Bottom-left: Estimated entanglement coefficient κ^ (fraction of the target’s activation region shared with semantic neighbors; see Definition 4.1 ). The Neighbors column lists representative discovered neighbors; full per-target neighbor sets are reported in Appendix C . Animal subtypes have high κ^ and dense neighbor sets; castle has low κ^ despite a comparable number of architectural neighbors. Ellipses are (z−μc)⊤Σc−1(z−μc)≤4 for the mean and covariance of each concept’s projected points; they are visualization aids only and are not used to compute κ^ , which is estimated in the original activation space.
Figure 2: Concept entanglement manifests as two coupled effects. Each row erases horse and is probed with three prompt types: Direct (mentions “ horse ”), Indirect (uses correlated context like jockey but omits “ horse ”), and Neighbor (semantically similar concept donkey ). ESD leaks horses through indirect prompts but preserves donkey; STEREO suppresses leakage but destroys donkey. We formalize this duality in Section 4 .
Erasure
Retention
Method
Type
UA ↑
IRR ↓
IRA ↑
NP ↑
CP ↑
DamageGap ↓
Original
–
0.0±0.0
58.7±38.1
70.45±12.6
100.0±0.0
100.0±0.0
–
UCE
Closed-form
80.6±18.4
10.5±23.4
67.1±12.4
88.6±25.1
99.3±3.1
10.7±25.4
RECE
Closed-form
99.4±0.9
3.0±4.0
43.3±12.1
56.1±10.6
93.5±2.0
37.4±10.2
ESD
FT
80.7±11.1
20.9±14.2
59.2±12.7
80.6±13.5
97.0±1.6
16.4±14.5
SalUn
FT
93.7±5.1
12.8±11.5
46.5±8.5
58.4±13.0
90.2±8.9
31.8±13.1
Table 1: Averaged evaluation of unlearning methods on the robustness-retention trade-off. Results are averaged over the five target concepts ( dog , bear , horse , cat , castle ) with κ^∈[0.27,0.89] . Cells report mean ± sample standard deviation across concepts. UA: Unlearning Accuracy; IRR: Indirect Recovery Rate; IRA/CP: In-domain/Cross-domain Retain Accuracy; NP: Neighbor Preservation (Eq. ( 27 )); DamageGap: composite retention-damage summary. The Original row reports baseline rates for the pretrained Dθ . A few runs use non-standard image counts; see the per-concept table captions in Appendix J . † SAeUron is averaged over four concepts ( dog , bear , horse , cat ); its released sparse autoencoder features do not cover castle (see Section J.5 ).
Figure 3: Neighbor preservation scales inversely with κ^ . Left: NP for each (method, concept) pair among the seven original methods; dashed lines connect each method’s points in order of κ^ and serve as visual guides, not fitted trends. Right: STEREO’s NP and in-domain retention ratio across the five concepts ( rs=−1.0 for both).
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Pairwise centroid cosine distance between concept activation regions at mid_block.attentions.0 , ordered by hierarchical clustering. Block structure in the matrix corresponds to the neighborhoods discovered by the thresholding procedure described above and to the κ^ values reported in Figure 1 .
Target
κ^
∣N(cu)∣
ρcu
Discovered neighbors
dog
0.89
5
0.13
puppy, labrador, husky, beagle, retriever
bear
0.82
5
0.14
grizzly, panda, polar bear, koala, cub
horse
0.81
7
0.15
pony, mare, donkey, stallion, foal, zebra, mule
cat
0.77
5
0.15
kitten, lynx, leopard, tiger, cheetah
castle
0.27
6
0.18
fortress, palace, citadel, keep, manor, cathedral
Appendix
Table 2: Per-target κ^ estimation results. For each target cu , we report the estimated entanglement coefficient κ^ , the number of discovered neighbors within the centroid-distance threshold δcu , the fattening radius ρcu (median intra-concept cosine distance), and the discovered neighbor set. castle has a discovered neighborhood of comparable size to the animal targets but κ^ roughly 3× lower, isolating semantic density rather than neighborhood cardinality as the driver of κ^ .
Layer
castle
cat
horse
bear
dog
rs vs. mid
Animal std
down_blocks.2.attentions.0
0.34
0.77
0.86
0.92
0.92
+0.97
0.061
down_blocks.2.attentions.1
0.26
0.71
0.74
0.87
0.85
+0.90
0.069
mid_block.attentions.0
0.27
0.77
0.81
0.82
0.89
—
0.043
up_blocks.1.attentions.0
0.09
0.59
0.48
0.88
0.79
+0.80
0.158
up_blocks.1.attentions.1
0.39
0.36
0.61
0.91
0.76
+0.80
0.203
up_blocks.2.attentions.0
0.00
0.78
0.63
0.89
0.88
+0.80
0.105
Appendix
Table 3: κ^ across six candidate probe layers. All five alternative layers produce a κ^ ranking strongly consistent with mid-block ( rs≥0.80 ). Mid-block additionally yields the tightest cluster of animal-subtype κ^ values (std =0.043 , vs. 0.061 to 0.203 elsewhere), making it the layer at which the “animal-subtype shared-manifold” structure is most coherently expressed.
Concept
Step 10
Step 25
Step 40
castle
0.77
0.27
0.20
cat
0.81
0.77
0.74
horse
0.80
0.81
0.84
bear
0.81
0.82
0.81
dog
0.95
0.89
0.86
sunglasses
0.99
0.99
0.99
Appendix
Table 4: κ^ at three denoising steps ( mid_block.attentions.0 ).
Figure 5: Mean activation displacement ∥Φθ^(p)−Φθ(p)∥2 for six unlearning methods erasing horse . Target prompts (blue) confirm Assumption 4.5 : erasure requires measurable displacement. Neighbor prompts (orange) are displaced more than control prompts (green) for all methods, consistent with the propagation mechanism in Theorem 4.6 .
Figure 6: Lipschitz validation via prompt variation. Each point is a (base, variant) prompt pair; the x -axis shows activation distance and the y -axis shows the change in CLIP-based concept probability. Points are colored by semantic distance from the base prompt (green: synonyms; blue: scene changes; orange: related concepts; red: unrelated concepts). All points lie below the empirical bound L^=0.174 (dotted line), supporting Assumption 4.4 .
Figure 7: Context-only prompts remain close to target prompts in CLIP text-embedding space. PCA projection of pooled prompts used for the horse case study. One group explicitly contains the concept token, while the other omits it but retains correlated context words (e.g., jockey , saddle , racetrack ). The proximity between these groups illustrates why token-level erasure can remain vulnerable to indirect contextual prompts.
Table 5: The 65 -concept classification vocabulary used for all CLIP-based metric evaluations in Section 5 . The vocabulary is the union of the five targets, their respective fine-grained subtype candidates spanning four animal category and one architectural category, and ten out-of-domain controls. See Table 2 for the per-target subset N(c) that survives the δc -filter and forms the averaging set for NP.
Target
TPR (%) ↑
FPR (%) ↓
dog
90
0.25
cat
100
0.25
bear
96
1.00
horse
99
1.25
castle
100
0.00
Average
97.0
0.55
Appendix
Table 6: CLIP classifier reliability on labeled images.
Figure 8: Per-concept operating points of the seven original methods. Each point is one (method, concept) pair from Appendix J ; color encodes method family and marker encodes the target concept. The shaded region marks strong erasure together with strong neighbor preservation (UA >90 , NP >80 ). No point for the four high- κ^ animal targets enters it; the only points that do are castle ( κ^=0.27 ) under SEOT and AdvUnlearn, consistent with the κ -scaled bound of Theorem 4.6 . SAeUron has no castle point ( Section J.5 ).
Erasure
Retention
Method
Type
UA ↑
IRR ↓
IRA ↑
NP ↑
CP ↑
DamageGap ↓
Original
–
0
27.0
78.92
100
100
–
ESD
FT
75.0
17.5
74.15
93.51
96.0
2.48
SalUn
FT
90.0
8.0
50.92
58.01
76.0
17.98
EDiff
FT
62.0
23.5
66.61
80.29
92.99
12.70
AdvUnlearn
Robust FT
96.0
20.0
74.30
63.05
99.0
35.94
Appendix
Table 7: Per-concept evaluation for dog ( κ^=0.89 ). Discovered neighbors include puppy , labrador , beagle , husky , and retriever (full list in Section I.1 ). As the highest- κ^ concept in our evaluation, dog exhibits the most severe robustness-retention trade-off: robust fine-tuning methods (STEREO, AdvUnlearn) achieve high UA but cause substantial IRA and NP degradation.
Erasure
Retention
Method
Type
UA ↑
IRR ↓
IRA ↑
NP ↑
CP ↑
DamageGap ↓
Original
–
0
85.5
62.75
100
100
–
ESD
FT
94.5
9.0
44.00
58.02
95.0
36.97
SalUn
FT
99.5
5.5
51.51
53.89
89.0
35.10
EDiff
FT
96.0
6.0
38.61
56.75
94.0
37.24
AdvUnlearn
Robust FT
99.0
1.0
42.46
67.37
100
32.62
Appendix
Table 8: Per-concept evaluation for bear ( κ^=0.82 ). Discovered neighbors include grizzly , panda , polar bear , koala , and cub (full list in Section I.1 ). Image counts differ for two runs: SalUn uses 272 neighbor images and SEOT uses 462 neighbor images.
Erasure
Retention
Method
Type
UA ↑
IRR ↓
IRA ↑
NP ↑
CP ↑
DamageGap ↓
Original
–
0
75.5
71.23
100
100
–
ESD
FT
80.0
33.5
62.92
81.12
97.0
15.87
SalUn
FT
95.0
28.5
40.92
39.42
91.99
52.57
EDiff
FT
81.0
31.0
63.84
76.48
94.0
17.51
AdvUnlearn
Robust FT
97.5
23.0
65.38
71.42
100
28.57
Appendix
Table 9: Per-concept evaluation for horse ( κ^=0.81 ). Discovered neighbors include pony , mare , donkey , zebra , stallion , foal and mule (full list in Section I.1 ). Image counts differ for two runs: Original uses approximately 480 neighbor and control images, and SEOT uses 58 direct images.
Erasure
Retention
Method
Type
UA ↑
IRR ↓
IRA ↑
NP ↑
CP ↑
DamageGap ↓
Original
–
0
96.0
85.50
100
100
–
ESD
FT
66.0
38.0
66.92
87.35
98.0
10.65
SalUn
FT
87.0
21.0
34.53
73.04
100
26.95
EDiff
FT
76.0
24.0
62.30
65.99
100
34.00
AdvUnlearn
Robust FT
100.0
0.0
50.76
72.40
98.0
25.60
Appendix
Table 10: Per-concept evaluation for cat ( κ^=0.77 ). Discovered neighbors include kitten , lynx , leopard , tiger , and cheetah (full list in Section I.1 ). The original baseline IRA = 85.50 reflects high CLIP-evaluator accuracy on cat’s in-domain object set after correcting for kitten/cat label confusion (see Appendix I ). Image counts differ for two runs: Original uses approximately 480 neighbor and control images, and SEOT uses 72 direct images.
Erasure
Retention
Method
Type
UA ↑
IRR ↓
IRA ↑
NP ↑
CP ↑
DamageGap ↓
Original
–
0
9.5
53.84
100
100
–
ESD
FT
88.0
6.5
48.15
83.03
99.0
15.98
SalUn
FT
97.0
1.0
54.76
67.58
94.0
26.41
EDiff
FT
82.0
2.5
46.46
87.35
99.0
11.65
AdvUnlearn
Robust FT
100.0
1.0
44.00
83.10
100
16.89
Appendix
Table 11: Per-concept evaluation for castle ( κ^=0.27 ). Discovered neighbors include fortress , palace , citadel , keep , cathedral and manor (full list in Section I.1 ). SAeUron’s released sparse autoencoder features do not cover castle ; we attempted to substitute its architectures feature, but the resulting ablation did not yield meaningful suppression of castle generation. Image counts differ for two runs: Original uses approximately 480 neighbor and control images, and SEOT uses 60 direct and 100 indirect images.
Target
κ^
∣N(cu)∣
Discovered neighbors
snake
0.37
7
anaconda, boa, cobra, garter snake, king snake, python, rattlesnake
birdwing, dragonfly, fritillary, monarch butterfly, moth, red admiral, swallowtail
car
0.83
7
convertible, coupe, hatchback, minivan, sedan, suv, van
Appendix
Table 12: Discovered neighbors for the five additional targets. As for the original targets ( Table 2 ), neighborhood size does not determine κ^ : fish has the most neighbors but an intermediate κ^ .
Concept
Method
κ^
UA (%) ↑
IRR (%) ↓
IRA (%) ↑
NP (%) ↑
CP (%) ↑
DamageGap (%) ↓
Snake
Original
0.37
9.2
6.0
31.1
100.0
100.0
0.0
EraseDiff
0.37
88.1
8.6
19.1
81.2
93.6
12.4
ESD
0.37
97.9
6.8
20.6
81.3
98.1
16.8
SalUn
0.37
100.0
0.0
5.5
72.4
86.2
13.8
STEREO
0.37
100.0
0.0
14.8
25.1
96.0
71.1
Fish
Original
0.52
6.2
11.0
40.5
100.0
100.0
0.0
Appendix
Table 13: Results for the five additional target concepts. Original denotes the pretrained model before unlearning. All evaluation metrics are percentages. Higher UA, IRA, NP, and CP are better; lower IRR and DamageGap are better. DamageGap is the difference between control preservation (CP) and neighbor preservation (NP).
Method
castle
cat
horse
bear
dog
rs
( κ^=0.27 )
( 0.77 )
( 0.81 )
( 0.82 )
( 0.89 )
vs. κ^
ESD
128.4
194.8
213.6
221.0
235.4
+1.00
STEREO
0 75.9
234.2
233.4
240.9
257.0
+0.90
AdvUnlearn
0 71.6
110.5
150.9
123.2
123.7
+0.70
EDiff
0 67.8
163.0
178.1
172.1
175.6
+0.70
SalUn
0 88.1
220.7
215.6
220.4
215.3
+0.10
Appendix
Table 14: Mean activation displacement at neighbor prompts. Each cell reports Δneighbor in L2 activation distance, averaged over 200 neighbor prompts. rs is computed across the five concepts. Note that n=5 concepts limits the statistical power of significance testing on individual methods; we report rs as a descriptive monotonicity measure. The consistency of positive correlations across all five methods (range +0.10 to +1.00 ; four of five at ≥0.70 ) provides aggregate evidence for the κ -scaling predicted by Theorem 4.6 .
Method
castle
cat
horse
bear
dog
Spread
STEREO
0.0035
0.0032
0.0034
0.0033
0.0034
7.8%
EDiff
0.0039
0.0046
0.0045
0.0047
0.0050
23.6%
SalUn
0.0030
0.0034
0.0037
0.0036
0.0040
29.5%
ESD
0.0021
0.0039
0.0037
0.0036
0.0037
53.4%
AdvUnlearn
0.0037
0.0068
0.0053
0.0065
0.0070
57.1%
Appendix
Table 15: Per-concept Ltight values. Ltight=κ(1−ε)/(Δneighbor+c0) is the Lipschitz constant value that places Theorem 4.6 ’s bound exactly at the observed displacement. If a method operates at the bound, Ltight should be approximately constant across concepts. STEREO is the only method exhibiting this consistency ( 7.8% spread), supporting its operation at the bound’s saturated-erasure regime.
Concept
κ^ (CLIP text)
κ^ (denoiser)
horse
1.00
0.72
pony
1.00
0.00
donkey
1.00
0.00
deer
1.00
0.00
dog
1.00
0.60
cat
1.00
0.00
Appendix
Table 16: κ^ for the ten-concept set in SDXL’s CLIP text-embedding space and in its denoiser cross-attention space.
Figure 9: UMAP projection of SDXL cross-attention activations for eight concepts, with covariance ellipses. As in SD v1.4 ( Figure 1 ), related animal concepts ( horse / pony / donkey , dog / cat ) overlap, while castle lies apart from the animal clusters. Castle’s architectural neighbors are not shown; its entanglement with them is κ^SDXL=0.72 ( Table 17 ).
Concept
κ^SDXL
UA ↑
NP ↑
DamageGap ↓
cat
0.65
99
90.9
9.1
castle
0.72
91
85.0
15.0
dog
0.92
77
77.1
22.9
Appendix
Table 17: ESD on SDXL. NP decreases monotonically with κ^SDXL .