Existing attacks on unlearned diffusion models assume that the erased concepts are known in advance and focus on recovering them. In practice, however, model providers may not disclose which concepts have been removed, and even with access to the original base model, an adversary may still lack a clear target to attack. In this paper, we aim to answer the following critical but overlooked questions: which concepts have been erased from the model, and how many have been erased in total? To this end, we present Tracer, a framework that rapidly and accurately identifies erased concepts and estimates their number. Tracer efficiently identifies erased concepts without generating and classifying images. By combining lightweight spectral analysis of weight footprints, it enables efficient search over large candidate vocabularies. To distinguish multiple erased concepts, we introduce a footprint coverage objective that guides sequential discovery. Tracer estimates the number of erased concepts by detecting a sharp decline in candidate confidence as the selected concepts account for the erasure footprint, without requiring labeled examples for calibration. The framework requires only lightweight linear algebra and limited forward probes, with no prior knowledge of the unlearning algorithm. Experiments across text-to-image and text-to-video backbones and diverse unlearning methods demonstrate that Tracer identifies erased concepts and estimates their number in seconds, achieving 150 to 137,000 times and 133 to 20,000 times speedups over MIA and brute-force search on image and video models, respectively, with substantially higher identification accuracy.
Figures & tables
Figure 1: (a) Motivation: existing unlearning attacks assume known unlearning targets, leaving the identification of targets unexplored. (b) Current problem: adapted MIA and brute-force method incur high computational costs and struggle to determine target counts and distinguish targets from affected non-targets. (c) Tracer : identifies erased concepts and estimates their number in seconds without prior target knowledge, knowledge of the unlearning algorithm, or image generation.
Figure 2: Overview of Tracer . (a) Erasure footprint assembly. Erasure imprints are grouped by their probed input representations. Spectrally reshaped and reweighted Gram matrices are fused within compatible groups into footprints Gg . (b) Candidate encoding. Candidates are encoded in each group’s input space, whitened to suppress shared structure, and normalized. (c) Unlearned concept disclosure. Tracer iteratively selects concepts explaining the remaining footprint energy. selection combines standardized gains across groups with confidence-based weights. Label-free self-calibration fixes parameters and group weights. Confidence drops determine set size without image generation or algorithm knowledge.
Method
Venue
Erased-set size K
Avg. ( ↑ )
1
3
5
10
20
Stable Diffusion v1.5
ESD
ICCV 2023
100
100
94
–
–
98.0
UCE
WACV 2024
100
100
100
100
100
100.0
MACE
CVPR 2024
100
100
100
100
–
100.0
FMN
CVPR 2024a
100
100
100
–
–
100.0
Table 1: Identification accuracy (%) of erased object concepts across three text-to-image backbones on Imagenette-20. Shaded cells report evaluated results. “–” indicates an unevaluated setting.
Figure 4Table 5
Figure 5: Runtime across three methods, showing the scalability of Tracer to large candidate sets.
Controls background suppression, from no transformation to full regularized whitening. Unit normalization is retained in both cases.
ρ
{0,0.5,1}
Controls spectral shaping within each imprint.
o
{2,1}
Specifies the token position relative to the non-padding sequence length for token-based text encodings.
Appendix
Table 6: Self-calibration search grid for the main configuration.
Parameter
Value
Role
κ
0.25
Specifies the upper-tail comparison range in the confidence statistic.
J
96
Sets the fixed search horizon.
B
8
Sets the window length for the initial-rise condition.
α
0.30
Sets the relative threshold for the first sufficiently large decline.
ε0
10−3
Sets the relative spectral cutoff in Eq. equation 1 .
εc
10−12
Stabilizes the confidence denominator in Eq. equation 4 .
Appendix
Table 7: Fixed parameters for confidence computation, erased-set size detection, and numerical stabilization in the main configuration.
Index
Object templates
Artistic-style templates
1
a photo of a {}
art by {}
2
{}
{}
3
an image of a {}
painting by {}
4
a picture of a {}
artwork by {}
5
a photo of the {}
style of {}
Appendix
Table 8: Prompt templates for probe matching and candidate encoding. Each placeholder is replaced with the complete vocabulary entry.
Erased-set size K
Method
Venue
1
3
5
10
20
Avg. ↑
Stable Diffusion v1.5
ESD
ICCV 2023
100
100
100
92
–
98.0
UCE
WACV 2024
100
100
100
100
100
100.0
MACE
CVPR 2024
100
100
100
100
–
100.0
FMN
CVPR 2024a
100
100
100
–
–
100.0
Appendix
Table 9: Identification accuracy (%) of erased artistic styles across three text-to-image backbones. K denotes the number of jointly erased styles. Shaded cells report evaluated results. “–” indicates an unevaluated setting. Avg. is the arithmetic mean over evaluated settings within each row.
Text-to-image diffusion models can generate prohibited content, which motivates concept erasure through machine unlearning. Most erasure methods intervene at the text interface, through prompt modification or localized updates to text-conditioning weights, and they are evaluated by what the model outputs for given prompts. Such evaluation cannot see what the network still encodes. Latent-space auditing, which bypasses text conditioning and probes the denoising network directly, shows that erased concepts remain recoverable from internal representations. We find that this also holds for methods built to be robust against adversarial prompts, and that the problem grows with the number of erased concepts. We propose Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models (AVCE), a framework that grounds erasure in the model's latent representations. AVCE audits the embedding neighborhood of each concept and condenses the discovered vulnerable directions into an anchor at the weakest geometric point. It edits cross-attention and self-attention projections in closed form at this anchor, then fine-tunes the two pathways with pathway-level auditing losses, using orthogonal gradient projection to consolidate multiple concepts. Experiments on SD v1.5, SDXL, and Flux 1.0 across object, explicit-content, and artistic-style unlearning show that AVCE reduces attack success rates by 5.07x and improves auditing scores by 3.84x over the strongest baseline, while preserving competitive generation quality.
Kaiyuan Deng, Yuchen Li, Gen Li +5
The University of Arizona · Peking University · Clemson University +2
The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We present eval-unlearn, an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models. eval-unlearn integrates twelve published unlearning techniques spanning fine-tuning, closed-form model editing, and inference-time intervention, alongside nine complementary evaluation metrics covering erasure efficacy, adversarial robustness, generative quality, and concept retention. Its plugin architecture lets third-party techniques and metrics self-register without modifying the core framework, and its streaming, batched pipeline supports efficient evaluation of both standard NSFW concepts and arbitrary general concepts. As a further contribution, we release a public leaderboard on HuggingFace along with an interactive tool for real-time evaluation of unlearning techniques. The leaderboard compares nudity concept erasure case study across all twelve techniques, exposing significant accuracy-quality trade-offs that are obscured by heterogeneous evaluation. eval-unlearn is released under the MIT license; the package, code, leaderboard, and documentation are all available at https://eval-unlearn.readthedocs.io.
Machine unlearning for text-to-image diffusion models aims to selectively remove undesirable concepts from pre-trained models without costly retraining. Current unlearning methods share a common weakness: erased concepts return when the model is fine-tuned on downstream data, even when that data is entirely unrelated. We adapt Projected Gradient Unlearning (PGU) from classification to the diffusion domain as a post-hoc hardening step. By constructing a Core Gradient Space (CGS) from the retain concept activations and projecting gradient updates into its orthogonal complement, PGU ensures that subsequent fine-tuning cannot undo the achieved erasure. Applied on top of existing methods (ESD, UCE, Receler), the approach eliminates revival for style concepts and substantially delays it for object concepts, running in roughly 6 minutes versus the ~2 hours required by Meta-Unlearning. PGU and Meta-Unlearning turn out to be complementary: which performs better depends on how the concept is encoded, and retain concept selection should follow visual feature similarity rather than semantic grouping.
Aljalila Aladawi, Mohammed Talha Alam, Fakhri Karray
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE. · University of Waterloo, Ontario, Canada.