cs.CVSep 27, 2026

Concept Score Relearning: A Unified Cross-Architecture Attack on Concept Erasure

Authors: Hong Xi Tae, Jiaming Zhang, Wenwen He, Xuan Wang, Wei Yang Bryan Lim

Organizations: College of Computing and Data Science, Nanyang Technological University, Singapore

Abstract

Concept erasure aims to suppress undesirable knowledge in text-to-image generative models. However, existing robustness evaluations typically rely on relearning attacks tailored to specific model architectures. We study concept reactivation across two substantially different generative paradigms: noise-prediction U-Nets and flow-matching Transformers. We introduce \textbf{Concept Score Relearning (CSR)}, a unified parameter-level framework that reactivates erased concepts by optimizing each model within its native prediction space. CSR requires no external target-concept image dataset and applies the same concept-directed objective to both U-Net-based Stable Diffusion and Transformer-based FLUX. Experiments across diverse concepts and multiple erasure methods demonstrate consistent concept reactivation across both architectures, highlighting the cross-architecture applicability of CSR and the persistent recoverability of apparently erased concepts. For strict nudity, CSR reaches average ASRs of 50.47% on FLUX and 40.29% on Stable Diffusion, consistently ranking first across all evaluated safety settings.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models

    May 19, 2026Yi Sun, Zhiqi Zhang, Xinhao Zhong +5Concept ErasureGenerative Flow Networks

  2. You Can't Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning

    Sep 28, 2026Yian Wang, Ali Ebrahimpour-Boroojeny, Hari Sundaram +1Text-To-Image Diffusion ModelsConcept Erasure

  3. GRACE: Adaptive Concept Erasure with Geometry-Guided Retention in Diffusion Models

    Sep 14, 2026Qinghui Gong, Yihuai Liang, Yuanlun Xie +3Concept ErasureText-To-Image Diffusion Models