cs.CVOct 7, 2026

Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

Authors: Gaurav Patel, Jun Fang, Greg Ver Steeg, Qiang Qiu, Sravan Sripada

Organizations: Purdue University · Amazon AGI

Abstract

Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Locality-Aware Continual Unlearning for Diffusion Models

    Dec 2, 2025Naveen George, Naoki Murata, Yuhta Takida +2Diffusion Model UnlearningMachine Unlearning

  2. TILDE: TILt-based Distributional Erasure for Concept Unlearning

    Jul 7, 2026Naveen George, Naoki Murata, Yuhta Takida +2Diffusion Model Fine-TuningDiffusion Model Unlearning