cs.LGSep 30, 2026

Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing

Authors: Kemou Li, Qizhou Wang, Yue Wang, Fengpeng Li, Zhuan Shi, Negar Rostamzadeh, Golnoosh Farnadi, Masashi Sugiyama, +1 more

Organizations: State Key Laboratory of Internet of Things for Smart City, University of Macau · RIKEN Center for Advanced Intelligence Project · The University of Melbourne · King Abdullah University of Science and Technology · Mila – Québec AI Institute · McGill University · Google Research · The University of Tokyo

Abstract

Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against such acquisition, motivating the problem of preemptive unlearning. Unlike retrospective unlearning, which removes capabilities already present in a fixed model, preemptive unlearning seeks to prevent their acquisition under unseen attack data and future fine-tuning procedures. Despite its practical importance, this setting remains largely unexplored, presents distinct challenges, and is therefore the central focus of our work. We first verify that existing retrospective methods provide insufficient pre-release protection. Even when forbidden capabilities are suppressed in current outputs, forbidden-domain data can still induce gradients through internal pathways, enabling later acquisition. Motivated by this finding, we propose a gradient-sealing principle that blocks these pathways by pushing relevant pre-activations into the negative region, where ReLU-family activations exhibit zero or near-zero derivatives. Experiments across multiple LLM families demonstrate our stronger resistance to downstream acquisition than retrospective baselines, validating gradient sealing as an effective mechanism for pre-release protection.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CAP: Controllable Alignment Prompting for Unlearning in LLMs

    Apr 23, 2026Zhaokun Wang, Jinyu Guo, Jingwen Pu +7Large Language Model UnlearningLarge Language Model Alignment

  2. Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

    May 26, 2026Kevin Kuo, Chhavi Yadav, Virginia SmithAttacker Large Language ModelLarge Language Model Fine-Tuning

  3. Model Unlearning Objectives Vary for Distinct Language Functions

    May 26, 2026Berk Atil, Vipul Gupta, Rebecca J. PassonneauLarge Language Model UnlearningLarge Language Model Safety