cs.LGSep 29, 2026

Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

Authors: Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen, Lixu Wang

Organizations: Southeast University (seu.edu.cn) · University of Sydney (usyd.edu.au) · Northwestern University (northwestern.edu) · Together AI (together.ai) · The Chinese University of Hong Kong Shenzhen (cuhk.edu.cn)

Abstract

Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV). We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571. The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025. The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Null-Space Constrained Low-Rank Adaptation for Response-Specified Large Language Model Unlearning

    Jun 9, 2026Bocheng Ju, Jianhua Wang, Chengliang Liu +1Large Language Model UnlearningExact Unlearning

  2. Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning

    Oct 6, 2025Kai Qin, Jiaqi Wu, Jianxiang He +8Large Language Model UnlearningForgetting

  3. Randomized Antipodal Search Done Right for Data Pareto Improvement of LLM Unlearning

    Apr 17, 2026Ziwen Liu, Huawei Lin, Yide Ran +5Large Language Model UnlearningPareto Frontier