cs.AIFeb 8, 2026

Selective Fine-Tuning for Targeted and Robust Concept Unlearning

Authors: Mansi, Avinash Kori, Francesca Toni, Soteris Demetriou

Organizations: Department of Computing Imperial College London, United Kingdom

Abstract

Text guided diffusion models are used by millions of users, but can be easily exploited to produce harmful content. Concept unlearning methods aim at reducing the models' likelihood of generating harmful content. Traditionally, this has been tackled at an individual concept level, with only a handful of recent works considering more realistic concept combinations. However, state of the art methods depend on full finetuning, which is computationally expensive. Concept localisation methods can facilitate selective finetuning, but existing techniques are static, resulting in suboptimal utility. In order to tackle these challenges, we propose TRUST (Targeted Robust Selective fine Tuning), a novel approach for dynamically estimating target concept neurons and unlearning them through selective finetuning, empowered by a Hessian based regularization. We show experimentally, against a number of SOTA baselines, that TRUST is robust against adversarial prompts, preserves generation quality to a significant degree, and is also significantly faster than the SOTA. Our method achieves unlearning of not only individual concepts but also combinations of concepts and conditional concepts, without any specific regularization.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. You Can't Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning

    Sep 28, 2026Yian Wang, Ali Ebrahimpour-Boroojeny, Hari Sundaram +1Text-To-Image Diffusion ModelsConcept Erasure

  2. Projected Gradient Unlearning for Text-to-Image Diffusion Models: Defending Against Concept Revival Attacks

    Apr 22, 2026Aljalila Aladawi, Mohammed Talha Alam, Fakhri KarrayText-To-Image Diffusion ModelsDiffusion Models

  3. Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models

    Sep 29, 2026Arian Komaei Koma, Seyed Amir Kasaei, Aida Aryafar +6Text-To-Image Diffusion ModelsDiffusion Models