Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models
Authors: Arian Komaei Koma, Seyed Amir Kasaei, Aida Aryafar, Matin Ghiasi, Ali Aghayari, Amirhossein Souri, Mohammad Mosayyebi, AmirMahdi Sadeghzadeh, +1 more
Organizations: Sharif University of Technology · Hong Kong University of Science and Technology
Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack of robustness to noise initialization. We call this phenomenon ``probabilistic forgetting'': suppressed concepts re-emerge under specific random initial noise conditions, despite appearing unlearned on other initializations. We trace this failure to the misalignment between standard Gaussian sampling during unlearning and the unlearning objective. Since the target concept manifests only in specific initial noise regions throughout the unlearning phase, uniform random sampling yields sparse, uninformative gradient updates that fail to drive robust erasure. To overcome this issue, we propose an adaptive, concept-conditioned sampling strategy that dynamically concentrates gradient updates on regions where the target concept manifests, down-weighting uninformative areas. We integrate our framework with six distinct SOTA unlearning methods across four diffusion backbones and evaluate it across safety, object, and artistic-style unlearning, as well as under black-box and white-box adversarial attacks. Our method reduces the conditional nudity re-emergence rate across random initializations by 67.2% on average over four baselines and lowers attack success rates across both adversarial evaluations. Across concept domains, Adaptive Noise Sampling strengthens adversarial robustness and non-target retention while preserving competitive generative quality and target-erasure performance.
Figures & tables
Figure 1: Inconsistency of existing unlearning methods across different noise initializations. Given a fixed prompt, the target concept is suppressed with certain noise vectors (left), but re-emerges when the noise is varied (right).
Figure 2: Classifier confidence over the noise space before and after unlearning. Left: original model. Middle: standard unlearning methods. Right: Unlearning + adaptive noise sampling. Heatmaps were generated by projecting 4000 noise latents into 2D via t-SNE and mapping their classifier-confidence scores via cubic interpolation.
Figure 3: Concept activation over the noise space. Left: The pre-trained model exhibits distinct activation peaks corresponding to the target concept. Right: Standard unlearning reduces the dominant peak but leaves residual high-activation regions. Uniform sampling (red) wastes updates on irrelevant noise points, while our adaptive sampling (green) focuses on the remaining concept-related regions.
Figure 4: Overview of the Adaptive Noise Sampling framework. While standard methods sample noise uniformly and update indiscriminately, our approach dynamically modulates the optimization landscape. (1) Evaluation: For each noise sample xT , a frozen classifier estimates the presence of the forbidden concept. (2) Reweighting: This probability is converted into an importance score Aθ(xT,c) , which acts as a weight for the unlearning loss. (3) Targeted Update: By assigning higher gradients to “active” noise regions and down-weighting safe ones, we concentrate the unlearning budget on the specific initializations that trigger the forbidden concept.
Figure 5: Qualitative comparison of unlearning robustness. While standard sampling fails to suppress the target, our adaptive strategy consistently erases the concept while preserving the original layout, visual similarity, and non-target semantics.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
∇θJ≈N1∑i∈Isg(wi)⋅∇θLfgt(θ;xti(i),c,ϵi)
Appendix
Algorithm 1 Adaptive Noise Sampling (Top- M )
Figure 6: Empirical trajectory of the batch normalizer estimate Z^θ(20)(c) during Stable Diffusion 1.4 unlearning with the prompt c=“nudity” . At each optimization step, the estimate is the mean NudeNet confidence over N=20 independently sampled initial-noise vectors. The true normalizer Zθ(c) remains intractable and is not plotted.
Method
Failure Count (20-seed) ↓
CLIP Score ↑
Standard ESD
510
0.2214
ESD + Ours (NudeNet)
201
0.2230
ESD + Ours (CLIP Zero-Shot)
210
0.2221
Standard EAP
1142
0.2358
EAP + Ours (NudeNet)
638
0.2341
EAP + Ours (CLIP Zero-Shot)
644
0.2330
Appendix
Table 8: Ablation of CLIP as the training-time classifier for nudity unlearning. We compare the primary NudeNet guidance with zero-shot CLIP guidance; all failure counts are independently measured using Open-NSFW2.
Gaussian Noise ( σ )
NudeNet Confidence
0.00
0.77
0.05
0.68
0.10
0.51
0.15
0.25
Appendix
Table 9: Classifier confidence scores under increasing Gaussian noise levels ( σ ). NudeNet exhibits graceful degradation, maintaining robust detection until heavy corruption is applied.
Method
Failure Count (20-seed) ↓
CLIP ↑
Standard ESD
510
0.221
ESD + Ours ( σ=0.00 )
201
0.223
ESD + Ours ( σ=0.05 )
233
0.218
ESD + Ours ( σ=0.10 )
282
0.229
ESD + Ours ( σ=0.15 )
308
0.234
Standard ACE
491
0.236
Appendix
Table 10: End-to-end robustness under independent Open-NSFW2 evaluation when Gaussian noise ( σ ) is applied to images during the NudeNet-guided unlearning phase. Even with a corrupted guidance signal, our adaptive method significantly outperforms the standard baselines across multiple architectures.
Scan Candidates ( N )
Time (s)
4
8.9
8
17.7
20
44.4
40
171.8
Appendix
Table 11: Average time cost (in seconds) per step for the forward-pass exploration scan across N candidates. Because this step relies solely on the frozen guidance classifier and the base diffusion model, this overhead is uniform across all underlying unlearning methods.
Method
Baseline Total Time
Ours
ACE
4h 40m
7h 20m
ESD
4h 28m
7h 32m
EAP
3h 22m
4h 15m
RECELER
4h 32m
7h 00m
Appendix
Table 12: Total wall-clock training time for the standard baselines and Adaptive Noise Sampling configured with N=20 and M=4 . The adaptive variants use fewer gradient-update steps but incur additional runtime from candidate scanning.
Baseline Method
Peak VRAM (GB)
EAP
24
ACE
24
ESD
22
RECELER
24
Appendix
Table 13: Peak VRAM utilization on a single NVIDIA A100 GPU using Adaptive Noise Sampling with N=20 and M=4 . Because candidate exploration is detached from the unlearning backward pass, memory usage remains comparable to standard uniform-sampling optimization.
Candidate Pool ( N )
Update Batch ( M )
ACE
ESD
EAP
RECELER
- (Standard Baseline)
-
491
510
1142
693
4
1
119
265
820
441
4
2
111
252
795
425
4
4
95
238
762
390
8
1
112
254
805
430
8
2
107
245
780
412
Appendix
Table 14: Ablation study on the candidate pool ( N ) and update batch ( M ) sizes using NudeNet for guidance and Open-NSFW2 for evaluation. We report the multi-seed Failure Count (lower is better) across four baselines. The standard baselines (top row) represent uniform sampling without our adaptive framework.
Method
Failure Count (1-seed) ↓
Failure Count (20-seed) ↓
CLIP Score ↑
SD 2.1 (Original)
130
2355
0.239
Standard EAP
75
1413
0.233
EAP + Ours
43
985
0.236
Standard ESD
32
673
0.221
ESD + Ours
12
250
0.223
Appendix
Table 15: Nudity unlearning performance scaled to Stable Diffusion 2.1 using NudeNet for guidance and Open-NSFW2 for evaluation. We report the Failure Count under standard (1-seed) and rigorous (20-seed) evaluations, alongside the CLIP score for generation quality.
SD3 Unlearning Method
Failures (1 seed) ↓
Failures (20 seeds) ↓
CLIP Score ↑
Standard DUO Baseline
60
1245
0.2385
DUO + Ours
46
938
0.2399
Appendix
Table 16: Proof-of-concept nudity unlearning results on Stable Diffusion 3 using NudeNet for guidance and Open-NSFW2 for evaluation. We compare standard DUO Park et al. (2024) with DUO augmented by our Adaptive Noise Sampling strategy. Lower failure counts and higher CLIP scores are better.
Figure 7: Qualitative comparison under the same prompt and identical noise seed. Our method generates images that are visually consistent with the baseline output while effectively removing the target concept, demonstrating precise and controlled unlearning without altering overall image structure.
Figure 8: Qualitative comparison across different unlearning methods under the same prompt. Our approach more consistently suppresses the target concept across stochastic noise initializations, while preserving semantic content and visual quality compared to baseline methods.