Organizations: State Key Laboratory of Internet of Things for Smart City, University of Macau · RIKEN Center for Advanced Intelligence Project · The University of Melbourne · King Abdullah University of Science and Technology · Mila – Québec AI Institute · McGill University · Google Research · The University of Tokyo
Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against such acquisition, motivating the problem of preemptive unlearning. Unlike retrospective unlearning, which removes capabilities already present in a fixed model, preemptive unlearning seeks to prevent their acquisition under unseen attack data and future fine-tuning procedures. Despite its practical importance, this setting remains largely unexplored, presents distinct challenges, and is therefore the central focus of our work. We first verify that existing retrospective methods provide insufficient pre-release protection. Even when forbidden capabilities are suppressed in current outputs, forbidden-domain data can still induce gradients through internal pathways, enabling later acquisition. Motivated by this finding, we propose a gradient-sealing principle that blocks these pathways by pushing relevant pre-activations into the negative region, where ReLU-family activations exhibit zero or near-zero derivatives. Experiments across multiple LLM families demonstrate our stronger resistance to downstream acquisition than retrospective baselines, validating gradient sealing as an effective mechanism for pre-release protection.
Figures & tables
Figure 1: Retrospective unlearning vs. preemptive unlearning. (a) Retrospective unlearning suppresses an already learned capability but may leave its gradient pathway open. (b) Preemptive unlearning instead uses a forbidden-domain proxy set before release to seal that pathway, preventing future acquisition from fine-tuning.
Figure 2: Empirical motivation for gradient sealing. (a) Release and post-acquisition ES for unlearning methods under identical acquisition. (b) Within-layer channel-rank alignment between Df look-ahead and Da acquisition. (c) Acquisition-gain reduction across selected-channel fractions and gradient-attenuation strengths.
Figure 3: Overview of GSU. Expose reveals learning-induced changes by comparing the original model with a disposable proxy-fitted copy under matched contexts. Localize selects a fixed gate set G based on upward pre-activation shifts. Seal restarts from θo and pushes selected pre-activations into the SiLU negative tail to attenuate local gradient factors, alongside response suppression and retain supervision, yielding θrel .
Llama-3.2-1B
Llama-3.2-3B
Llama-3.1-8B
Qwen3.5-2B
Qwen3.5-4B
Qwen3.5-9B
Method
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
No Defense
9.04
16.33
50.08
9.36
19.41
50.75
9.51
19.63
53.68
6.48
6.69
6.84
9.55
26.36
66.12
19.44
90.92
98.53
Disjoint Prevention
GradDiff
8.62
15.12
51.41
9.10
20.62
35.48
8.69
18.83
44.56
5.78
6.25
6.96
9.45
26.87
80.24
19.88
93.40
98.72
NPO
8.04
13.57
47.06
8.34
19.99
4.42
7.84
20.42
95.80
4.85
5.07
5.74
9.06
25.57
64.99
19.72
93.15
98.86
RMU
9.02
16.17
49.59
9.42
19.37
50.40
9.35
19.77
86.55
6.47
6.71
6.88
9.53
26.48
68.21
19.75
92.90
98.32
WGA
8.31
14.34
48.19
8.61
20.57
48.01
8.40
21.93
89.77
4.55
5.46
6.07
9.27
25.28
69.94
19.79
93.10
99.16
Table 1: TOFU results of Llama3 and Qwen3.5 families under disjoint and identical prevention.
Bio
Cyber
Method
F3↓
F5↓
UA90↓
F3↓
F5↓
UA90↓
No Defense
65.23
65.36
65.99
43.78
43.83
44.14
GradDiff
37.54
37.72
37.87
33.21
33.71
35.51
NPO
35.44
36.39
36.48
32.64
31.08
33.19
RMU
34.04
34.83
34.24
35.09
32.48
32.57
WGA
36.89
37.01
37.18
30.99
31.72
30.11
Table 2: WMDP results under disjoint attacks.
Figure 4: Acquisition resistance, gate response, and benign learnability. (a) ES over ten attack epochs. (b) Mean absolute SiLU derivative over selected gates at release on B1 , B2 , and retain inputs. (c) Validation NLL reduction after two epochs of benign fine-tuning, interpreted relative to initial loss. Panels (a,b) use Qwen3.5-2B under disjoint TOFU; all results are single runs.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
x,y
Prompt and autoregressive response.
πθ,ℓθ
Response likelihood and response NLL.
F,R ; PF,PR
Forbidden and normal domains, and their data distributions.
Df,Da,Dr
Defender proxy, unseen attacker, and retain sets.
θo,θla,θrel,θatk
Original, look-ahead, released, and post-attack parameters.
U,A,A
Pre-release defense, acquisition attack, and attack class.
Appendix
Table 3: Core notation used in the paper.
Figure 5: A minimal ReLU counterexample. Panel (a) shows the forbidden-sensitive path in red and the retain path in gray; dashed arrows are fixed inhibitory connections. For xf=e1 and xr=e2 , both released states in (b) have identical forbidden and retain outputs. Reducing the forbidden activation from 1 to δ reduces the projected attacker-gradient norm by δ and the exact one-step acquisition gain by δ2 .
Model
K
p
Seal alpha
τ
Llama-3.2-1B
4
5%
0.01
−9
Qwen3.5-2B
4
5%
0.03
−4
Llama-3.2-3B
2
5%
0.01
−4
Appendix
Table 4: Frozen configurations for the supplementary analyses. “Seal alpha” retains the terminology of the experiment records.
Figure 6: Utility accompanying the acquisition trajectory. Left: the ES observations from Fig. 4 a , with GSU below all displayed controls at every checkpoint. Right: utility at the same checkpoints on the 100× composite-score scale, without utility filtering.
Method
ES0
ES3
ES5
ES10
U0
U10
No Defense
6.016
6.463
6.659
6.868
69.871
68.485
NPO
4.599
4.838
5.056
5.782
70.691
69.220
RMU
6.127
6.446
6.616
6.851
69.944
68.511
NPO+SAM
4.730
4.965
5.102
5.533
71.663
70.599
GSU
3.379
3.889
4.713
5.377
70.169
69.955
Appendix
Table 5: Extraction and utility during the shared attack. ES entries are percentages; ES3 averages epochs 1–3. U is shown on the 100× composite-score scale, not as task accuracy.
Figure 7: Complete component ablation on Qwen3.5-2B under disjoint TOFU. Left: mean ES over attack epochs 1–5, in percent. Right: release utility as a raw composite score. Full GSU lowers early extraction relative to the three nonzero-utility ablations; without retention, utility is zero, so low ES alone is not evidence of a successful utility-preserving defense. Results use one training seed.
Variant
ES0
ES5
ES10
U0
U10
GSU
3.3793
4.1740
5.3768
0.701687
0.699545
Without seal
4.6104
4.9598
5.7530
0.705906
0.695012
Without suppression
4.1924
5.0298
5.5130
0.683498
0.705820
Without retain
0.0000
0.1997
0.7997
0.000000
0.000000
Random gates
4.3565
5.2196
5.7881
0.707940
0.695478
Appendix
Table 6: Complete ablation measurements. ES entries are percentages; U0 and U10 are unscaled composite utility scores, not accuracies. ES5 averages attack epochs 1–5.
Figure 8: Sealing contributions across the attack trajectory. The GSU and without-seal runs from § F.2 are shown at all eleven observed checkpoints. GSU retains lower ES (left), while the narrowing gap and accompanying utility (right, 100× composite-score scale) show how the comparison evolves.
Figure 9: Gate-wise pre-activation distributions at release. Each violin contains 1,232 gate-wise means, obtained by averaging answer-token pre-activations within examples and then examples equally. White markers and dark segments show the median and interquartile range across gates; the density outlines describe between-gate variation, not uncertainty across training runs.
Figure 10: Local response and negative-tail occupancy at release. Left: mean absolute SiLU derivative, equally averaged across recorded gates. Right: mean negative-tail fraction at τ=−4 (colored), with the complement in gray. Both statistics change more on target inputs than retain inputs; they measure selected-gate responses before attack, not full-network gradients.
Mean ∣ϕ′(u)∣
Negative-tail fraction (%)
Inputs
Original
GSU
w/o seal
Original
GSU
w/o seal
B1
0.3994
0.3018
0.3881
0.411
3.684
0.487
B2
0.3997
0.3079
0.3904
0.408
3.348
0.463
Retain
0.3929
0.3484
0.3890
0.451
1.029
0.477
Appendix
Table 7: Release-stage gate statistics. Derivatives and negative-tail fractions are equally averaged across recorded gates; tail fractions use τ=−4 .
Figure 11: Selected-gate responses before and after acquisition attacks. GSU and without sealing are probed at release and after ten attack epochs on B2 ; dashed lines denote the original base model. Panels show mean pre-activation, mean absolute SiLU derivative, and occupancy at u≤−4 over the same 1,232 channels. Statistics average valid answer-prediction positions within examples, then examples and channels.
Model state
Mean u
Mean ∣ϕ′(u)∣
Tail (%)
Original base
−0.240670
0.399727
0.4084
GSU, release
−0.973011
0.307875
3.3482
GSU, attack epoch 10
−0.456793
0.373816
0.5722
Without seal, release
−0.302069
0.390441
0.4628
Without seal, attack epoch 10
−0.148507
0.419492
0.3177
Appendix
Table 8: Matched B2 gate-response measurements. All states use the same selected channels and aggregation. Tail occupancy uses u≤−4 and is reported in percent.
Figure 12: Economics accuracy during benign fine-tuning. Accuracy is equally averaged over two economics MMLU subjects. The two Llama GSU models improve and narrow their gaps to No Defense; Qwen3.5-2B declines for both initializations. The three checkpoints belong to the same benign runs used for the NLL analysis.
Figure 13: Benign validation-loss trajectories. NLL at epochs 0, 1, and 2 for the runs summarized in Fig. 4 c . All three GSU models approach the final loss of their undefended controls. GSU starts and finishes slightly higher, so larger reductions do not imply better final performance or higher learning efficiency.
Model
Initialization
NLL0
NLL2
ΔNLL
Llama-3.2-1B
No Defense
2.7635
2.4991
0.2645
GSU
2.7771
2.5033
0.2738
Qwen3.5-2B
No Defense
2.4733
2.2123
0.2610
GSU
2.5408
2.2171
0.3237
Llama-3.2-3B
No Defense
2.5707
2.2999
0.2707
GSU
2.5899
2.3030
0.2869
Appendix
Table 9: Benign validation-loss endpoints. ΔNLL is the absolute decrease over two epochs; displayed values are rounded independently.
Figure 14: Resistance after benign fine-tuning. ES (left) and utility (right, 100× composite-score scale) during direct attacks (dashed) and attacks after two epochs of benign fine-tuning (solid). GSU retains lower final extraction and higher final utility than the corresponding undefended control after benign adaptation on Qwen3.5-2B.
Large language models (LLMs) trained on unfiltered corpora inherently risk retaining sensitive information, necessitating selective knowledge unlearning for regulatory compliance and ethical safety. However, existing parameter-modifying methods face fundamental limitations: high computational costs, uncontrollable forgetting boundaries, and strict dependency on model weight access. These constraints render them impractical for closed-source models, yet current non-invasive alternatives remain unsystematic and reliant on empirical experience. To address these challenges, we propose the Controllable Alignment Prompting for Unlearning (CAP) framework, an end-to-end prompt-driven unlearning paradigm. CAP decouples unlearning into a learnable prompt optimization process via reinforcement learning, where a prompt generator collaborates with the LLM to suppress target knowledge while preserving general capabilities selectively. This approach enables reversible knowledge restoration through prompt revocation. Extensive experiments demonstrate that CAP achieves precise, controllable unlearning without updating model parameters, establishing a dynamic alignment mechanism that overcomes the transferability limitations of prior methods.
Zhaokun Wang, Jinyu Guo, Jingwen Pu +7
School of Information and Software Engineering, University of Electronic Science and Technology of China
Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmful usage without fine-tuning at all? In this paper, we show that open-weight safeguards are susceptible to simpler strategies that, despite being well known, have not been systematically evaluated against these safeguards. Specifically, we evaluate two low-cost attacks--abliteration and prefilling--that do not rely on gradient-based optimization. Across three harmfulness evaluation benchmarks (BeaverTails, HarmBench, and AdvBench), these attacks increase attack success rates against safeguarded open-weight models from below 10% to a range of 16%-96%. To mitigate this vulnerability, we introduce abliteration-resistant tuning (ART), which incorporates an abliteration-based objective into training. ART can be layered onto existing defenses and reduces the success rates of abliteration, prefilling, and their combination by 10%-20%. These findings indicate that the attack surface for open-weight models is broader than previously characterized, and that evaluations of safeguarding defenses should incorporate a more diverse set of attack strategies beyond adversarial fine-tuning.
Kevin Kuo, Chhavi Yadav, Virginia Smith
Carnegie Mellon University · Simons Institute, UC Berkeley
Large language models (LLMs) learn undesirable properties during pretraining, including dangerous knowledge and toxic text generation. Just as post-training uses different objectives to shape different behaviors, we argue that unlearning methods should be designed for the language function at issue. To study this, we consider two mechanistically distinct unlearning goals, dangerous-knowledge unlearning and toxicity unlearning. For dangerous knowledge, we introduce a cosine-based, meta-learned variant of RMU. For toxicity, we propose a multi-layer objective based on layer-specific probe directions. Across four open-source 7-8B models, our methods achieve strong results, based on distinct training objectives for the two types of unlearning. Overall, our results suggest that unlearning should be studied as a family of problems, analogous to the multiple types of LLM post-training.