VisualErase: Dual-Branch Visual Trajectory Redirection for Robust Concept Erasure in Text-to-Image Diffusion Models
Organizations: Harbin Institute of Technology (Shenzhen) · City University of Hong Kong · Shenzhen Loop Area Institute · National University of Singapore · Shandong University
Abstract
Concept erasure is essential for the safe deployment of text-to-image diffusion models, as they may reproduce harmful, copyrighted, or privacy-sensitive content learned from unconstrained large-scale data. Existing methods typically erase unwanted concepts while preserving general generation capability by redirecting target-related text-to-image mappings. However, recent studies show that erased models may still retain visual generative trajectories of target concepts, leaving them vulnerable to adversarial recovery attacks and revealing a fundamental gap between redirecting text-to-image mappings and truly removing visual knowledge. To bridge this gap, we propose VisualErase, a new paradigm that redirects concept-bearing visual generative trajectories toward explicitly defined concept-removed outcomes. To enable this redirection, we use structure-preserving image editing to construct content-aligned, concept-removed counterparts for source images, providing explicit visual endpoints that retain non-target content. We then derive a denoising target from each source-to-counterpart pair and use a dual-branch redirection loss to align both text-conditioned and unconditional predictions with this target, since conditional supervision alone does not explicitly constrain generation without textual guidance. To mitigate the adverse effects of concept erasure on non-target generation, we jointly optimize the redirection loss with a counterpart retention loss that matches denoising predictions from the frozen pretrained model. Across style, celebrity, and nudity erasure, VisualErase limits the maximum attack success rate over seven attacks to 0%, 8%, and 0.1%, respectively, while retaining general generation quality. These results highlight the importance of visual trajectory redirection for robust concept erasure beyond text-to-image mappings alone.
Figures & tables
| Erasure method | ACC | Attack Success Rate (%) | Generation utility | |||||||
| PEZ | MMA | RAB | P4D | UDA | CCE | TINA+ | FID | CLIP | ||
| SD1.4 | 76.0 | 40.0 | 40.0 | 90.0 | 94.0 | 96.0 | 60.0 | 76.0 | 14.0 | 26.6 |
| ESD | 2.0 | 2.0 | 0.0 | 6.0 | 30.0 | 32.0 | 8.0 | 60.0 | 14.5 | 25.9 |
| FMN | 10.0 | 0.0 | 2.0 | 6.0 | 54.0 | 56.0 | 18.0 | 70.0 | 13.8 | 26.5 |
| AC | 12.0 | 0.0 | 6.0 | 14.0 | 68.0 | 77.0 | 14.0 | 72.0 | 14.0 | 26.5 |
| MACE | 4.0 | 0.0 | 0.0 | 4.0 | 42.0 | 56.0 | 26.0 | 72.0 | 12.4 | 23.5 |
| Erasure method | ACC | Attack Success Rate (%) | Generation utility | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| PEZ | MMA | RAB | P4D | UDA | CCE | TINA+ | FID | CLIP | ||
| SD1.4 | 35.6 | 31.4 | 34.8 | 70.5 | 40.7 | 39.8 | 3.5 | 42.4 | 14.0 | 26.6 |
| ESD | 2.5 | 0.8 | 0.7 | 20.0 | 3.5 | 5.6 | 46.5 | 39.0 | 13.6 | 25.6 |
| FMN | 33.9 | 30.5 | 31.1 | 69.5 | 38.0 | 38.7 | 15.5 | 44.9 | 13.8 | 26.3 |
| UCE | 0.0 | 4.2 | 6.9 | 8.4 | 7.0 | 7.7 | 14.8 | 41.5 | 14.3 | 26.4 |
| MACE | 0.0 | 1.7 | 0.1 | 0.0 | 5.6 | 7.7 | 17.6 | 43.2 | 12.8 | 24.2 |
| Variant | TINA+ ASR | FID | CLIP |
|---|---|---|---|
| Cond. | 66.0 | 16.0 | 26.8 |
| Uncond. | 2.0 | 25.3 | 26.5 |
| Cond. + Uncond. | 0.0 | 22.0 | 26.5 |
| Cond. + Uncond. + Retent. ( VisualErase ) | 0.0 | 14.7 | 26.5 |
| Trainable scope | Updated (M) | UDA ASR | TINA+ ASR | FID | CLIP |
|---|---|---|---|---|---|
| Cross-KV | 19.2 | 2.0 | 46.0 | 21.1 | 26.1 |
| Cross-Attn | 44.0 | 0.0 | 2.0 | 17.4 | 26.7 |
| Self-Attn | 49.6 | 6.0 | 0.0 | 14.2 | 26.6 |
| All-Attn ( VisualErase ) | 93.5 | 0.0 | 0.0 | 14.7 | 26.5 |
| Non-Cross | 815.4 | 2.0 | 0.0 | 14.5 | 26.5 |
| All Linear/Conv | 859.3 | 0.0 | 0.0 | 14.4 | 26.5 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Erasure method | ACC | Attack Success Rate (%) | Generation utility | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| PEZ | MMA | RAB | P4D | UDA | CCE | TINA+ | FID | CLIP | ||
| SD1.4 | 94.1 | 66.1 | 70.9 | 96.8 | 100.0 | 100.0 | 32.4 | 100.0 | 14.0 | 26.6 |
| ESD | 21.2 | 11.9 | 13.1 | 50.5 | 69.0 | 76.1 | 74.7 | 86.4 | 13.6 | 25.6 |
| FMN | 88.1 | 62.7 | 67.0 | 97.9 | 97.9 | 97.9 | 54.9 | 100.0 | 13.8 | 26.3 |
| UCE | 24.6 | 25.4 | 32.6 | 29.5 | 76.1 | 78.9 | 49.3 | 97.5 | 14.3 | 26.4 |
| MACE | 10.2 | 8.5 | 6.0 | 6.3 | 75.4 | 81.7 | 50.0 | 93.2 | 12.8 | 24.2 |
| Trainable scope | Active | Updated (M) | ACC | TINA+ | FID | CLIP |
|---|---|---|---|---|---|---|
| All-Attn | 100% | 93.5 | 0.0 | 6.8 | 15.5 | 26.8 |
| Non-Cross | 100% | 815.4 | 0.0 | 0.8 | 16.7 | 26.3 |
| All Linear/Conv | 100% | 859.3 | 0.0 | 0.8 | 16.3 | 26.4 |
| SG Non-Cross | 50% | 407.7 | 0.0 | 0.0 | 16.7 | 26.4 |
| SG Non-Cross | 25% | 203.9 | 0.8 | 0.8 | 16.0 | 26.5 |
| SG Non-Cross ( VisualErase ) | 10% | 81.5 | 0.0 | 0.0 | 15.5 | 26.6 |
| Method | Time (min) | Relative time |
|---|---|---|
| STEREO | 30.0 | 2.12 |
| VisualErase | 14.2 | 1.00 |