Concept Score Relearning: A Unified Cross-Architecture Attack on Concept Erasure
Organizations: College of Computing and Data Science, Nanyang Technological University, Singapore
Abstract
Concept erasure aims to suppress undesirable knowledge in text-to-image generative models. However, existing robustness evaluations typically rely on relearning attacks tailored to specific model architectures. We study concept reactivation across two substantially different generative paradigms: noise-prediction U-Nets and flow-matching Transformers. We introduce \textbf{Concept Score Relearning (CSR)}, a unified parameter-level framework that reactivates erased concepts by optimizing each model within its native prediction space. CSR requires no external target-concept image dataset and applies the same concept-directed objective to both U-Net-based Stable Diffusion and Transformer-based FLUX. Experiments across diverse concepts and multiple erasure methods demonstrate consistent concept reactivation across both architectures, highlighting the cross-architecture applicability of CSR and the persistent recoverability of apparently erased concepts. For strict nudity, CSR reaches average ASRs of 50.47% on FLUX and 40.29% on Stable Diffusion, consistently ranking first across all evaluated safety settings.
Figures & tables
| Method | React. | Align. | Preserv. | Quality |
| ReFLUX | 4.52 0.78 | 3.89 1.29 | 3.91 0.91 | 4.03 0.44 |
| Ours | 4.62 0.91 | 4.74 0.22 | 4.68 0.53 | 4.52 0.76 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Description |
| Image in pixel space | |
| VAE encoder and decoder | |
| Clean latent | |
| Generic perturbed latent in CSR; noisy latent at discrete diffusion timestep for U-Net-based Stable Diffusion | |
| Generic timestep/noise-level notation in CSR; discrete diffusion timestep for U-Net-based Stable Diffusion | |
| Cumulative diffusion noise-schedule coefficient for U-Net-based Stable Diffusion |
| Aspect | U-Net-based Stable Diffusion | Transformer-based FLUX |
| Backbone | Convolutional U-Net | Multimodal Transformer |
| Generative formulation | Latent diffusion | Flow matching |
| Forward perturbation | Variance-scheduled Gaussian noising | Linear data–noise interpolation |
| Time/noise parameter | Discrete timestep | Continuous noise level |
| Native prediction | Predicted noise | Predicted velocity field |
| Text conditioning | U-Net cross-attention | Multimodal Transformer attention |
| Configuration | SD-1.4 | FLUX.1-dev |
| Native prediction | Predicted noise | Predicted velocity field |
| Anchor images | 40 | 40 |
| Image sampling steps | 25 | 28 |
| Sampling guidance | 7.5 | 3.5 |
| Reference targets | Online | Cached |
| Cached tuples | – |
| Victim | ASR-Strict | ASR-Permissive |
| EAP | 53.68 (2.73 ) | 71.93 (1.39 ) |
| MACE | 58.25 (2.96 ) | 75.09 (1.46 ) |
| ESD | 33.33 (1.70 ) | 48.42 (0.94 ) |
| CP | 51.58 (2.62 ) | 70.53 (1.37 ) |
| Entity | Abstraction | Relationship |
| A photo of fruit | A scene featuring explosion | A shake hand B |
| A photo of ball | A scene featuring green bag | A kiss B |
| A photo of car | A scene featuring yellow bag | A hug B |
| A photo of airplane | A scene featuring time | A in B |
| A photo of tower | A scene featuring two cats | A on B |
| A photo of building | A scene featuring three cats | A back to back B |
| Architecture | Excluded Concept | Reference Detection Accuracy |
| SD-1.4 | A scene featuring two cats | 0% |
| A in B | 0% | |
| A on B | 0% | |
| A hold B | 10% | |
| A amidst B | 10% | |
| A scene featuring explosion | 30% |
| Setting | ASR-Strict | ASR-Permissive |
| Restoration strength | ||
| (reference matching) | 17.89 | 34.04 |
| 28.07 | 43.86 | |
| 32.98 | 49.47 | |
| 29.82 | 50.88 | |
| Number of anchors | ||
| Setting | ASR-Strict | ASR-Permissive |
| Restoration strength | ||
| (reference matching) | 10.88 | 31.23 |
| 13.33 | 33.33 | |
| 38.25 | 51.93 | |
| 30.53 | 43.51 | |
| Number of anchors | ||