Organizations: State Key Laboratory of Internet of Things for Smart City, University of Macau · RIKEN Center for Advanced Intelligence Project · The University of Melbourne · King Abdullah University of Science and Technology · Mila – Québec AI Institute · McGill University · Google Research · The University of Tokyo
Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against such acquisition, motivating the problem of preemptive unlearning. Unlike retrospective unlearning, which removes capabilities already present in a fixed model, preemptive unlearning seeks to prevent their acquisition under unseen attack data and future fine-tuning procedures. Despite its practical importance, this setting remains largely unexplored, presents distinct challenges, and is therefore the central focus of our work. We first verify that existing retrospective methods provide insufficient pre-release protection. Even when forbidden capabilities are suppressed in current outputs, forbidden-domain data can still induce gradients through internal pathways, enabling later acquisition. Motivated by this finding, we propose a gradient-sealing principle that blocks these pathways by pushing relevant pre-activations into the negative region, where ReLU-family activations exhibit zero or near-zero derivatives. Experiments across multiple LLM families demonstrate our stronger resistance to downstream acquisition than retrospective baselines, validating gradient sealing as an effective mechanism for pre-release protection.
Figures & tables
Figure 1: Retrospective unlearning vs. preemptive unlearning. (a) Retrospective unlearning suppresses an already learned capability but may leave its gradient pathway open. (b) Preemptive unlearning instead uses a forbidden-domain proxy set before release to seal that pathway, preventing future acquisition from fine-tuning.
Figure 2: Empirical motivation for gradient sealing. (a) Release and post-acquisition ES for unlearning methods under identical acquisition. (b) Within-layer channel-rank alignment between Df look-ahead and Da acquisition. (c) Acquisition-gain reduction across selected-channel fractions and gradient-attenuation strengths.
Figure 3: Overview of GSU. Expose reveals learning-induced changes by comparing the original model with a disposable proxy-fitted copy under matched contexts. Localize selects a fixed gate set G based on upward pre-activation shifts. Seal restarts from θo and pushes selected pre-activations into the SiLU negative tail to attenuate local gradient factors, alongside response suppression and retain supervision, yielding θrel .
Llama-3.2-1B
Llama-3.2-3B
Llama-3.1-8B
Qwen3.5-2B
Qwen3.5-4B
Qwen3.5-9B
Method
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
ES3↓
ES5↓
UA90↓
No Defense
9.04
16.33
50.08
9.36
19.41
50.75
9.51
19.63
53.68
6.48
6.69
6.84
9.55
26.36
66.12
19.44
90.92
98.53
Disjoint Prevention
GradDiff
8.62
15.12
51.41
9.10
20.62
35.48
8.69
18.83
44.56
5.78
6.25
6.96
9.45
26.87
80.24
19.88
93.40
98.72
NPO
8.04
13.57
47.06
8.34
19.99
4.42
7.84
20.42
95.80
4.85
5.07
5.74
9.06
25.57
64.99
19.72
93.15
98.86
RMU
9.02
16.17
49.59
9.42
19.37
50.40
9.35
19.77
86.55
6.47
6.71
6.88
9.53
26.48
68.21
19.75
92.90
98.32
WGA
8.31
14.34
48.19
8.61
20.57
48.01
8.40
21.93
89.77
4.55
5.46
6.07
9.27
25.28
69.94
19.79
93.10
99.16
Table 1: TOFU results of Llama3 and Qwen3.5 families under disjoint and identical prevention.
Bio
Cyber
Method
F3↓
F5↓
UA90↓
F3↓
F5↓
UA90↓
No Defense
65.23
65.36
65.99
43.78
43.83
44.14
GradDiff
37.54
37.72
37.87
33.21
33.71
35.51
NPO
35.44
36.39
36.48
32.64
31.08
33.19
RMU
34.04
34.83
34.24
35.09
32.48
32.57
WGA
36.89
37.01
37.18
30.99
31.72
30.11
Table 2: WMDP results under disjoint attacks.
Figure 4: Acquisition resistance, gate response, and benign learnability. (a) ES over ten attack epochs. (b) Mean absolute SiLU derivative over selected gates at release on B1 , B2 , and retain inputs. (c) Validation NLL reduction after two epochs of benign fine-tuning, interpreted relative to initial loss. Panels (a,b) use Qwen3.5-2B under disjoint TOFU; all results are single runs.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
x,y
Prompt and autoregressive response.
πθ,ℓθ
Response likelihood and response NLL.
F,R ; PF,PR
Forbidden and normal domains, and their data distributions.
Df,Da,Dr
Defender proxy, unseen attacker, and retain sets.
θo,θla,θrel,θatk
Original, look-ahead, released, and post-attack parameters.
U,A,A
Pre-release defense, acquisition attack, and attack class.
Appendix
Table 3: Core notation used in the paper.
Figure 5: A minimal ReLU counterexample. Panel (a) shows the forbidden-sensitive path in red and the retain path in gray; dashed arrows are fixed inhibitory connections. For xf=e1 and xr=e2 , both released states in (b) have identical forbidden and retain outputs. Reducing the forbidden activation from 1 to δ reduces the projected attacker-gradient norm by δ and the exact one-step acquisition gain by δ2 .
Model
K
p
Seal alpha
τ
Llama-3.2-1B
4
5%
0.01
−9
Qwen3.5-2B
4
5%
0.03
−4
Llama-3.2-3B
2
5%
0.01
−4
Appendix
Table 4: Frozen configurations for the supplementary analyses. “Seal alpha” retains the terminology of the experiment records.
Figure 6: Utility accompanying the acquisition trajectory. Left: the ES observations from Fig. 4 a , with GSU below all displayed controls at every checkpoint. Right: utility at the same checkpoints on the 100× composite-score scale, without utility filtering.
Method
ES0
ES3
ES5
ES10
U0
U10
No Defense
6.016
6.463
6.659
6.868
69.871
68.485
NPO
4.599
4.838
5.056
5.782
70.691
69.220
RMU
6.127
6.446
6.616
6.851
69.944
68.511
NPO+SAM
4.730
4.965
5.102
5.533
71.663
70.599
GSU
3.379
3.889
4.713
5.377
70.169
69.955
Appendix
Table 5: Extraction and utility during the shared attack. ES entries are percentages; ES3 averages epochs 1–3. U is shown on the 100× composite-score scale, not as task accuracy.
Figure 7: Complete component ablation on Qwen3.5-2B under disjoint TOFU. Left: mean ES over attack epochs 1–5, in percent. Right: release utility as a raw composite score. Full GSU lowers early extraction relative to the three nonzero-utility ablations; without retention, utility is zero, so low ES alone is not evidence of a successful utility-preserving defense. Results use one training seed.
Variant
ES0
ES5
ES10
U0
U10
GSU
3.3793
4.1740
5.3768
0.701687
0.699545
Without seal
4.6104
4.9598
5.7530
0.705906
0.695012
Without suppression
4.1924
5.0298
5.5130
0.683498
0.705820
Without retain
0.0000
0.1997
0.7997
0.000000
0.000000
Random gates
4.3565
5.2196
5.7881
0.707940
0.695478
Appendix
Table 6: Complete ablation measurements. ES entries are percentages; U0 and U10 are unscaled composite utility scores, not accuracies. ES5 averages attack epochs 1–5.
Figure 8: Sealing contributions across the attack trajectory. The GSU and without-seal runs from § F.2 are shown at all eleven observed checkpoints. GSU retains lower ES (left), while the narrowing gap and accompanying utility (right, 100× composite-score scale) show how the comparison evolves.
Figure 9: Gate-wise pre-activation distributions at release. Each violin contains 1,232 gate-wise means, obtained by averaging answer-token pre-activations within examples and then examples equally. White markers and dark segments show the median and interquartile range across gates; the density outlines describe between-gate variation, not uncertainty across training runs.
Figure 10: Local response and negative-tail occupancy at release. Left: mean absolute SiLU derivative, equally averaged across recorded gates. Right: mean negative-tail fraction at τ=−4 (colored), with the complement in gray. Both statistics change more on target inputs than retain inputs; they measure selected-gate responses before attack, not full-network gradients.
Mean ∣ϕ′(u)∣
Negative-tail fraction (%)
Inputs
Original
GSU
w/o seal
Original
GSU
w/o seal
B1
0.3994
0.3018
0.3881
0.411
3.684
0.487
B2
0.3997
0.3079
0.3904
0.408
3.348
0.463
Retain
0.3929
0.3484
0.3890
0.451
1.029
0.477
Appendix
Table 7: Release-stage gate statistics. Derivatives and negative-tail fractions are equally averaged across recorded gates; tail fractions use τ=−4 .
Figure 11: Selected-gate responses before and after acquisition attacks. GSU and without sealing are probed at release and after ten attack epochs on B2 ; dashed lines denote the original base model. Panels show mean pre-activation, mean absolute SiLU derivative, and occupancy at u≤−4 over the same 1,232 channels. Statistics average valid answer-prediction positions within examples, then examples and channels.
Model state
Mean u
Mean ∣ϕ′(u)∣
Tail (%)
Original base
−0.240670
0.399727
0.4084
GSU, release
−0.973011
0.307875
3.3482
GSU, attack epoch 10
−0.456793
0.373816
0.5722
Without seal, release
−0.302069
0.390441
0.4628
Without seal, attack epoch 10
−0.148507
0.419492
0.3177
Appendix
Table 8: Matched B2 gate-response measurements. All states use the same selected channels and aggregation. Tail occupancy uses u≤−4 and is reported in percent.
Figure 12: Economics accuracy during benign fine-tuning. Accuracy is equally averaged over two economics MMLU subjects. The two Llama GSU models improve and narrow their gaps to No Defense; Qwen3.5-2B declines for both initializations. The three checkpoints belong to the same benign runs used for the NLL analysis.
Figure 13: Benign validation-loss trajectories. NLL at epochs 0, 1, and 2 for the runs summarized in Fig. 4 c . All three GSU models approach the final loss of their undefended controls. GSU starts and finishes slightly higher, so larger reductions do not imply better final performance or higher learning efficiency.
Model
Initialization
NLL0
NLL2
ΔNLL
Llama-3.2-1B
No Defense
2.7635
2.4991
0.2645
GSU
2.7771
2.5033
0.2738
Qwen3.5-2B
No Defense
2.4733
2.2123
0.2610
GSU
2.5408
2.2171
0.3237
Llama-3.2-3B
No Defense
2.5707
2.2999
0.2707
GSU
2.5899
2.3030
0.2869
Appendix
Table 9: Benign validation-loss endpoints. ΔNLL is the absolute decrease over two epochs; displayed values are rounded independently.
Figure 14: Resistance after benign fine-tuning. ES (left) and utility (right, 100× composite-score scale) during direct attacks (dashed) and attacks after two epochs of benign fine-tuning (solid). GSU retains lower final extraction and higher final utility than the corresponding undefended control after benign adaptation on Qwen3.5-2B.