Machine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities. Although existing methods use retention objectives or restrict where edits occur, achieving the desired forgetting level can still leave collateral changes that impair non-target behavior. Our recovery comparisons suggest that some of these changes can be reversed while preserving observed forgetting performance. In this work, we present Propose-Then-Project Unlearning (PTP-U), a framework that combines targeted forgetting with the recovery of non-target capabilities. PTP-U first applies local analytic edits to weaken target knowledge associations, then aligns non-target output distributions with those of the original model to recover capabilities while maintaining fixed forgetting constraints. Both stages serve a common goal: satisfying the forgetting requirements while preserving fluent generation and performance on non-target tasks. Across three benchmarks, PTP-U achieves the strongest forgetting-retention trade-off among evaluated methods, reaching 81.22%-91.03% forgetting while preserving 94.20% non-target utility on average. At matched forgetting, PTP-U consistently retains higher non-target utility.
Figures & tables
Figure 1: PTP-U first achieves forgetting feasibility (Stage I), then restores utility while maintaining feasibility (Stage II). Right: target-answer suppression persists during recovery.
Figure 2: Checkpoint-wise recovery across unlearning methods. Dashed lines show original checkpoints and solid lines their independently repaired copies. Shading marks paired differences. The horizontal axis records fine-tuning steps. PTP-U step 0 follows the Stage I analytic edits. Repairs are performed separately and do not feed into subsequent original checkpoints.
Figure 3: Overview of PTP-U. Stage I computes a curvature-guided closed-form proposal in an editable subspace. It updates Attention first, refreshes the operating point, then updates FFN to obtain a feasible unlearned model. Stage II constructs a soft target from current and Base policies for constrained recovery. It preserves forgetting while reducing behavioral drift and excessive refusal.
Method
Forget Perf.
Retain Perf.
BLEU ↓
R-L ↓
MMLU ↑
Flu. ↑
Original
74.76
98.68
46.39
4.03
WHP
23.55
17.93
38.49
2.52
ICUL
38.47
32.49
39.86
3.60
SCANS
19.92
17.72
42.05
3.27
ALTER
12.96
10.40
41.84
3.08
Table 1: Copyright unlearning results.
Baseline
RWKU
WMDP
FB ↓
QA ↓
AA ↓
Gen. ↑
Flu. ↑
Bio. ↓
Cyber ↓
Gen. ↑
Llama3.1-8B-Instruct ( Grattafiori et al., 2024 )
Base
63.95
66.42
69.83
68.37
4.20
72.74
47.35
68.37
ICUL ⋆
43.23
32.57
42.60
64.80
4.07
45.90
33.34
61.10
RMU ‡
19.10
23.84
21.30
54.71
2.96
32.52
32.75
53.43
MET ‡
33.64
28.47
23.54
58.52
3.18
34.63
32.47
52.19
Table 2: Main results on RWKU and WMDP. † , ‡ , and ⋆ denote output-distribution optimization, representation intervention & parameter editing, and prompt-based unlearning methods, respectively.
Figure 4: Hyperparameter sensitivity on WMDP. Editing window Ledit , indexed by its last layer; each window contains three consecutive layers. Closer to 25% indicates better forgetting, and orange diamonds show that higher MMLU accuracy is better.
Table 7
Figure 5: Post-unlearning benign relearning attack. Shading highlights the gap between the PTP-U and PTP-U’ variant.
Machine unlearning aims to eliminate the influence of sensitive data on a model. In the real world, unlearning requests arrive continually, which gives rise to two challenges. First, an unlearning intervention may redistribute target-related computation across remaining pathways, allowing previously forgotten knowledge to re-emerge. Second, repeated unlearning interventions may progressively reduce the model capacity needed to preserve retained utility. To address these challenges, we propose the Trajectory-guided Forget-Recover Network (TFR-Net). TFR-Net tracks channel-level risk across requests. It separates persistent target-related channels from transient hotspots and suppresses only the persistent ones. TFR-Net also recovers model capacity by reactivating dormant channels. These channels make strong contributions to retained utility and show low current and historical forget risk. The recovery is accepted only when retained-utility degradation remains within a predefined tolerance. Experiments on four datasets show that TFR-Net consistently achieves a more favorable trade-off between unlearning effectiveness and retained utility than representative baselines.
Zezheng Wu, Xinghe Cheng, Qinggang Zhang +4
Guilin University of Electronic Technology · Jinan University · Jilin University +2
Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full retraining. Most existing methods require a retain set of curated examples to prevent catastrophic degradation of general model utility, creating an extra data dependency that complicates deployment. We propose SHRED (Self-distillation via High-surprisal-only Retain-set-free Entropy Demotion), a retain-set-free unlearning method built on a key insight: not all tokens within a forget set instance carry memorized information equally. High-information tokens concentrate the model's memorized knowledge, while low-information tokens reflect general language competence. SHRED operates in two stages. (1) Selection: We perform a forward pass on a forget set instance, collect per-token autoregressive probabilities, and select the bottom (lowest probability, highest Shannon information) as forget positions; the remaining positions are retained as benign anchors. (2) Training: We construct modified KL targets that demote the memorized token's logit at forget positions while preserving the original distribution at benign positions. The model is then trained via a single top KL self-distillation objective that simultaneously drives forgetting and utility preservation. We evaluate SHRED across four standard unlearning benchmarks and demonstrate that it establishes a new Pareto-optimal trade-off between forget efficacy and model utility, outperforming retain-set-dependent methods. Our analysis shows that SHRED is robust against relearning attacks and membership-inference attacks, and it maintains stable utility even after many sequential unlearning runs.
Zizhao Hu, Ameya Godbole, Johnny Tian-Zheng Wei +3
University of Southern California · USC Information Sciences Institute
Machine unlearning for large language models (LLMs) remains challenging because full retraining is costly, while approximate methods often struggle to remove targeted behaviors without degrading retained utility, especially under limited post-deployment supervision. We consider a practical PEFT setting for targeted behavioral contamination removal with a small forget set, a limited retain buffer, and LoRA-only updates, and propose RapidUn, an influence-guided framework that converts cross-sample influence estimates into fixed sample-specific weights for weighted LoRA unlearning. Across Llama-3-8B on Dolly-15k and Alpaca-57k, with cross-model validation on Mistral-7B + Dolly-15k, RapidUn achieves lower seen-trigger and OOD-trigger-family ASR than Fisher, GA, and LoReUn while maintaining competitive clean utility. On Llama-3-8B + Alpaca-57k, it achieves a 77x wall-clock speedup over the clean-corpus LoRA retraining reference. Complementary TOFU, semantic LLM-judge, and IFEval evaluations further support the effectiveness of influence-guided sample reweighting beyond the controlled trigger benchmark.
Guoshenghui Zhao, Huawei Lin, Weijie Zhao
Golisano College of Computing and Information Sciences, Rochester Institute of Technology