Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data. In this data-free regime, we analyze where forgetting occurs and why. Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere. This localization is governed by the vocabulary deficiency of the corpus rather than the training mode, allowing pre-retraining risk ranking from token counts alone within a fixed base model. Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam's second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates. We therefore propose an intervention: raising Adam's epsilon exclusively for the output projection during training. Across eight settings spanning 160M to 12B parameters and four model families, this removes 39.4% to 67.9% of forgetting across all seven stable configurations without degrading target learning or requiring per-model tuning. The defense combines additively or better with replay (79.8% on Qwen/Korean) and rescues released-head LoRA from a 23-fold forgetting surge. Because post-hoc editing of the drifted rows recovers under 5% of forgetting, the intervention must operate during training. Our findings indicate that a single-line optimizer adjustment may serve as the primary defense against catastrophic forgetting where the corpus starves the vocabulary.
Figures & tables
Figure 1: Overview: where the forgetting lives, why it lives there, and what acting on it buys.
Base model (architecture)
Primary domain
Additional domains
Pythia 160M–12B (GPT-NeoX)
Korean
math, code, Alpaca, KoAlpaca (410M and 1.4B only)
Qwen2.5-0.5B (tied)
Korean
math, code, Alpaca, KoAlpaca
TinyLlama-1.1B (Llama, untied)
Korean
–
OLMo-2-1B (olmo2, untied)
Korean
–
Table 1: The grid of settings. Korean is the primary injection domain across all base models; other domains vary vocabulary deficiency at fixed models (Section 4.2 ). All settings run replay-free for 40M tokens across three seeds (12.6M for instruction tuning; Table 4 ) unless noted.
Figure 2: (a) Causal share by (site, v^ band): bars denote forgetting reduction (up) and learning change (down) from freezing one group (Korean, replay-free). Hatched: Qwen (tied) at 1\times10−4 . Pythia-160M ( ∗ ) is at 1\times10−5 due to peak-rate collapse. (b) Pythia-160M across rates, measured with output-scoped ϵ rather than freezing (Table A2 ). Full grid: Table A3 .
Table 4
Pythia-410M
Pythia-1.4B
Qwen2.5-0.5B
Alpaca
KoAlpaca
Alpaca
KoAlpaca
Alpaca
KoAlpaca
Attention freeze
16.0 ( −1.0 )
42.7 ( −0.7 )
7.0 ( −0.5 )
30.2 ( +0.4 )
17.2 ( −1.5 )
15.5 ( +0.1 )
MLP freeze
27.9 ( −5.8 )
8.2 ( −23.1 )
39.7 ( −3.7 )
47.8 ( −16.3 )
44.8 ( −15.4 )
26.1 ( −26.3 )
Output-embedding freeze
2.7 ( −1.2 )
−3.7 ( +3.0 )
0.3 ( +0.1 )
−1.6 ( +0.5 )
11.6 ( −3.1 ) †
46.9 ( −3.9 ) †
Input-embedding freeze
1.6 ( −0.0 )
6.6 ( +1.1 )
0.8 ( +0.2 )
−1.4 ( +0.5 )
n/a
n/a
Output-scoped ϵ
14.6 ( −0.1 )
29.8 ( +6.2 )
5.6 ( −0.3 )
33.1 ( +0.9 )
6.0 ( −1.0 )
43.3 ( −1.4 )
Table 4: Instruction fine-tuning (replay-free, 12.6M tokens): forgetting reduction (learning change). Freezes hold modules at base; ϵ slows low- v^ output rows ( † freezes Qwen’s tied table). Each cell is at its own peak rate; only 410M is rate-matched across the Alpaca/KoAlpaca pair (Appendix A.1 ).
Setting
Tokens/param
Ref. forget (nats)
Interv. forget
Reduction (%)
Learning change (%)
Pythia-160M ∗
0.2469
1.448
0.388
73.2
+9.9
Pythia-410M
0.0988
0.664
0.373
43.8
+10.9
Qwen2.5-0.5B
0.0810
0.629
0.202
67.9
+5.4
TinyLlama-1.1B
0.0364
0.361
0.187
48.1
+0.4
Pythia-1.4B
0.0284
0.473
0.261
44.8
+2.2
OLMo-2-1B
0.0269
1.137
0.689
39.4
+1.8
Table 5: Output-only ϵ (Korean, replay-free, 40M tokens, peak lr; 3 seeds, 2 for 12B/OLMo baselines). ∗ 160M is at 3\times10−5 due to peak collapse (Figure 2 b, Table A2 ). The 410M arm is bimodal across seeds: at six seeds of the same configuration it reads 32.7% and −7.0% (Table A7 ).
Table 7
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
lr
Ref. forget
Out. params (M)
In. params (M)
In. zero (%)
Body params (M)
160M
3\times10−4
3.392
30.7
30.9
45.8
1.2
410M
1\times10−4
0.656
38.7
41.3
45.7
2.7
1.4B
1\times10−4
0.472
80.6
88.0
42.8
15.9
TinyLlama
1\times10−4
0.355
48.2
50.4
45.3
13.1
OLMo
3\times10−4
1.160
176.6
199.6
52.8
6.9
Qwen
1\times10−4
0.628
110.0
n/a
n/a
1.0
Appendix
Table A1: Reference forgetting and the composition of the lowest band ( v^<1\text{\times}{10}^{-6}$$ ) per site: parameters held (M) and, for the input embedding, the share of that set whose v^ is exactly zero. Reference forgetting is this campaign’s own pair; at 160M, whose runs scatter widely at 3\times10−4 , the six-seed mean of Table A2 is 3.108 nats against the 3.392 here.
lr
Ref. forget (nats)
Interv. forget
Reduction (%)
Learning change (%)
n (ref/interv)
1\times10−5
0.491
0.126
74.2
−2.5
3/3
3\times10−5
1.448
0.388
73.2
+9.9
3/3
1\times10−4
2.068
1.855
10.3
−7.5
6/6
3\times10−4
3.108
3.795
−22.1
−17.3
6/6
1\times10−3
5.233
4.681
10.5
+20.5
2/3
Appendix
Table A2: Pythia-160M across learning rates: output-only ϵ , Korean, replay-free, 40M tokens. The 1\times10−3 uptick sits inside the seed spread. Cells here are matched pairs drawn from one campaign, with the same seed count on both arms, whereas the 160M row of Table A11 pools every available reference at each rate into a deeper but unbalanced denominator; that is why the two differ by up to four points at the shared rates.
Site
Band
160M
160M †
410M
1.4B
TinyLlama
OLMo
Qwen
Output emb.
<1\times10−6
-6.2 (-13.4)
68.8 (+3.1)
31.4 (+5.2)
33.7 (+2.0)
37.0 (+0.1)
31.3 (+1.6)
51.0 (+4.4)
1\times10−6 – 1\times10−5
-13.6 (-9.7)
-20.1 (-0.8)
2.3 (+6.3)
1.1 (+0.1)
3.6 (-0.4)
0.3 (-0.1)
10.7 (+1.3)
1\times10−5 – 1\times10−4
-10.9 (-7.7)
-17.8 (+2.2)
-5.6 (+0.6)
-1.0 (-0.2)
-3.5 (-0.5)
1.3 (-0.4)
1.1 (-0.8)
Input emb.
<1\times10−6
-3.6 (-4.6)
7.3 (-1.8)
1.4 (+9.0)
-1.5 (+0.1)
-0.6 (-0.5)
0.4 (-0.2)
n/a
1\times10−6 – 1\times10−5
-17.9 (-6.5)
6.8 (-2.4)
3.3 (+8.1)
-1.5 (+0.2)
-0.3 (-0.4)
1.9 (-0.2)
n/a
1\times10−5 – 1\times10−4
-18.7 (-9.3)
12.5 (+0.2)
7.2 (+9.9)
-0.5 (+0.1)
1.3 (-0.2)
-1.3 (-0.2)
n/a
Appendix
Table A3: Causal share by (site, v^ band): forgetting reduction (learning change in parentheses, %) from freezing that group. Korean, replay-free, 40M tokens, peak lr, two seeds per arm except Qwen at three. Every column is at that setting’s peak rate except † , which repeats the 160M grid at 1\times10−5 : the 160M peak collapses and every band row in its column goes negative (the two count-matched controls do not), so † is what Figure 2 a plots for that model and what (iv) of Section 4.1 reads. Each column divides by its own campaign’s three-seed reference trained at that column’s rate. Dashes are arms not run. Freezing a band and freezing by token count are different membership rules, not two estimates of one quantity. Count-matched rows freeze the output lowest band’s mass from the named site’s low end. Stats: Table A1 .
Setting
Exactly 0 (%)
<1\text{\times}{10}^{-8}$$ (%)
<1\text{\times}{10}^{-6}$$ (%)
Zero, input emb. (%)
Pythia-160M
8.7
15.4
32.8
36.6
Pythia-410M
4.7
4.6
16.5
36.6
Pythia-1.4B
2.7
3.5
10.7
36.6
TinyLlama-1.1B
2.1
3.0
8.3
34.8
OLMo-2-1B
7.1
11.2
20.1
51.2
Qwen2.5-0.5B
0.0
5.0
22.5
n/a
Appendix
Table A4: Distribution of v^ at the end of the reference run (peak lr, Korean, replay-free, 40M tokens, seed 0). Zero coordinates never received a gradient (unfed input rows; tied Qwen has none) and carry no update for ϵ to damp; the two threshold columns exclude them. The last column shares the zeros of the input table alone (Table A1 ’s In-zero column instead shares them of the input’s lowest band).
Setting
lr
ϵ=1\text{\times}{10}^{-5}$$
ϵ=1\text{\times}{10}^{-4}$$
ϵ=1\text{\times}{10}^{-3}$$
Pythia-160M
3\times10−4
−6.2 ( −10.7 )
−22.1 ( −17.3 )
−30.0 ( −24.4 )
Pythia-410M
1\times10−4
41.5 ( +13.2 )
43.8 ( +10.9 )
32.7 ( +2.5 )
Pythia-1.4B
1\times10−4
40.9 ( +2.1 )
44.8 ( +2.2 )
45.6 ( +2.2 )
Appendix
Table A5: Sensitivity to the ϵ value, Korean injection, replay-free, 40M tokens: forgetting reduction (learning change in parentheses, both %) against the same reference. At 1.4B the response is flat across the hundredfold span; 410M gives up eleven points at the 1\times10−3 edge; in 160M’s collapse regime no value rescues the run. The sweep exists only for these three Pythia checkpoints, so the flatness claim is not tested on the other three families.
Method
Setting
Forgetting reduction (%)
Learning change (%)
Output-only ϵ (ours)
ϵ=1\text{\times}{10}^{-4}$$
69.5
+0.4
Under-exposed freeze
K=1,000
70.7
−1.4
Step-wise gating
70.2
−0.4
L2 -to-base
λ=10
11.2
−1.3
L2 -to-base
λ=30
29.3
−4.4
L2 -to-base
λ=100
64.3
−15.4
Appendix
Table A6: No baseline reaches the free point. Qwen / Korean, replay-free, 40M tokens, three seeds per arm, all measured against one reference trained at lr 3\times10−5 (forget +0.1373 nats, learn +0.3189 ). Arms whose own sweep peaks elsewhere carry their rate in the Setting column. Qwen’s reference learns most at 1\times10−4 ; 3\times10−5 is the rate at which this baseline campaign was run. Qwen ties its input and output embeddings, so on this model “output-only” scopes ϵ to the one shared table and the input/output separation of Table A18 cannot be drawn here.
Base model
wd
Reduct.
Learn.
n
Pythia-410M
0
38.4%
−5.7 %
6
Pythia-410M
0.1
32.7%
−7.0 %
6
Qwen2.5-0.5B
0
68.0%
+5.7 %
3
Qwen2.5-0.5B
0.1
67.9%
+5.6 %
3
Appendix
Table A7: Disabling weight decay leaves forgetting and reduction unchanged, confirming drift is not a decay artifact (Korean, replay-free, 40M tokens, lr 10−4 ). The decoupled term λθ of Eq. ( 2 ) is therefore inert at this site, which is why Section 3.2 reads the adaptive step alone. The 410M gap reverses sign as seeds are added: at three seeds wd=0.1 led, 39.2% against 29.0% for wd=0 , while at six seeds wd=0 leads, 38.4% against 32.7%. A gap that changes sign with three more seeds is noise rather than an effect.
Figure A1: Three instruments stopping the same rows achieve equivalent reduction (Korean, replay-free; rates in Table A8 ).
Forgetting reduction (%)
Setting
lr
Output-only ϵ
Freeze, K=1,000
Step-wise gating
Max gap (pp)
Qwen2.5-0.5B
3\times10−5
69.5
70.7
70.2
1.2
Pythia-1.4B
1\times10−4
44.8
45.1
43.7
1.5
TinyLlama-1.1B
1\times10−4
48.1
50.1
48.2
2.0
OLMo-2-1B
3\times10−4
40.7
40.7
38.1
2.6
Pythia-410M
1\times10−4
43.8
38.1
43.7
5.7
Appendix
Table A8: Comparison of three equivalence instruments (output-specific ϵ=1×10−4 , freezing rows with c(t)<1,000 , and step-wise gating) under Korean injection (replay-free, 40M tokens, two seeds per arm except Qwen and every ϵ arm at three). The rate column is the campaign rate, which is the peak rate except at Qwen, whose reference learns most at 1\times10−4 , and at 160M, which has no stable peak. Cells keep this campaign’s seed pool; main-table rows pool more seeds, hence small offsets (OLMo 40.7 here vs. 39.4 in Table 5 ).
Arm
Forgetting reduction (%)
Learning change (%)
Freeze the whole output projection
74.4
−4.4
Freeze a random 96.6% of the embeddings
66.4
−4.4
Freeze by count, K=100 (96.6%)
68.2
−0.2
Output-only ϵ=1\text{\times}{10}^{-4}$$
72.1
+0.3
Count-freeze and ϵ together
73.2
−0.8
Appendix
Table A9: Retention under different selection criteria at matched mass: every arm below the first freezes 96.6% of the output embedding, while the first row freezes the whole projection, so only the lower four rows are mass-matched with each other (Qwen/Korean, 40M tokens, lr 3\times10−5 , three seeds, one reference with forget +0.0993 nats). This campaign carries 1% replay, which is why its ϵ row is the 72.1% share of Table 7 and not the replay-free 69.5% of Table A6 .
Arm
Forgetting reduction (%)
Learning change (%)
embedding learning rate ×0.10
67.5
+0.7
embedding ϵ=5\text{\times}{10}^{-5}$$
66.7
+2.2
embedding ϵ=2\text{\times}{10}^{-5}$$
64.2
+2.8
embedding ϵ=5\text{\times}{10}^{-6}$$
57.9
+3.0
embedding ϵ=2\text{\times}{10}^{-6}$$
52.0
+2.8
embedding learning rate ×0.36
51.3
+1.4
Appendix
Table A10: Equivalence of embedding learning-rate scaling and ϵ elevation, alongside body negative controls across v^ bands (Qwen/Korean, replay-free, three seeds). All arms share one reference, whose learning is 0.3507 and forgetting 0.1924 nats; that reference belongs to this campaign and is not the one behind Table A6 .
Setting
1\times10−5
3\times10−5
1\times10−4
3\times10−4
1\times10−3
Pythia-160M
73.8 ( −2.9 )
71.4 ( +6.5 )
10.0 ( −7.8 )
−17.7 ( −16.5 )
8.0 ( +6.1 )
Pythia-410M
50.7 ( −1.0 )
50.6 ( +0.7 )
43.5 ( +10.3 )
15.3 ( −25.6 )
8.8 ( −49.1 )
Pythia-1.4B
57.6 ( −0.1 )
52.1 ( +0.6 )
44.9 ( +2.0 )
32.5 ( +3.8 )
6.7 ( +16.0 )
TinyLlama-1.1B
n/q ( −0.9 )
111.6 ( −0.2 )
47.7 ( +0.4 )
27.9 ( +2.3 )
6.1 ( +11.8 )
OLMo-2-1B
44.5 ( −3.1 )
46.0 ( −0.8 )
41.9 ( +0.5 )
34.2 ( +1.5 )
26.9 ( +4.5 )
Qwen2.5-0.5B
75.4 ( −3.8 )
69.5 ( +0.4 )
67.7 ( +5.7 )
60.9 ( +15.9 )
30.8 ( −77.5 )
Appendix
Table A11: Forgetting reduction and learning change (parentheses, both %) across learning rates (Korean, replay-free, 40M tokens; 3–10 seeds, independently paired vs. Table 5 ). Values >100% denote old loss below base. n/q marks a ratio with no readable meaning, and the reason differs by column: in the reduction position the baseline forgets too little to divide by ( <0.025 nats), while in the learning position the reference itself ends below its base model on the new domain ( −0.45 and −0.63 nats at 12B), so change in Eq. ( 4 ) divides by a negative quantity and returns +13.3% and +50.9% for arms that in fact learned less.
Pythia-6.9B
Pythia-1.4B
Task
untouched
1\times10−8
1\times10−5
untouched
1\times10−8
1\times10−5
LAMBADA
61.05
60.00
65.17
60.78
52.53
56.79
HellaSwag
63.24
57.89
62.06
52.06
48.08
49.90
PIQA
76.22
72.09
73.23
71.00
67.19
69.04
ARC-easy
67.26
64.02
65.32
60.73
57.00
58.33
WinoGrande
61.64
59.35
61.01
56.99
55.46
56.64
Appendix
Table A12: Capability by benchmark rather than loss, at ϵ=1\text{\times}{10}^{-5}$$ against the default 1\times10−8 : five likelihood-scored tasks at two Pythia checkpoints. Both campaigns run at lr 1\times10−4 under 1% replay on a math injection, 6.9B to 40M tokens and 1.4B to 60M, three seeded training runs per arm; the 1.4B columns average three benchmark evaluations per arm and the 6.9B columns one. Accuracy in percent; higher is better.
Setting
Rows
Top-1, absent (%)
Top-1, control (%)
RMS, absent
RMS, control
Qwen-0.5B / ko
85,143
66.1
3.2
2.17\times10−3
5.53\times10−4
OLMo-2-1B / ko
47,492
72.6
8.8
5.97\times10−3
2.38\times10−3
Pythia-1.4B / ko
15,542
80.8
12.9
6.30\times10−3
1.67\times10−3
Pythia-1.4B / math
1,242
72.0
3.7
3.08\times10−3
5.87\times10−4
Appendix
Table A13: Drift geometry of the output embedding, replay-free checkpoints. Rows: the size of the absent set ( c(t)=0 ), which is not the number fed to the SVD (Table A14 ); that runs on a size-matched subsample of this set and of the control ( c(t)≥1,000 ). Top-1: energy share on the group’s first singular direction; RMS: mean displacement per row. Rates and budgets, shared with Tables 7 and A14 : Qwen 3\times10−5 at 40M, OLMo 1\times10−4 at 60M, Pythia-1.4B/ko 1\times10−4 at 40M, Pythia-1.4B/math 3\times10−5 at 60M.
Figure A2: Per-token displacement on the top-2 drift plane (Qwen / Korean, 400 rows per group): absent rows (red) move together, well-fed rows (blue) stay near the origin. Inset: drift energy by mode.
Shared removed
Residual removed
Setting
Rows
Top-1 (%)
Recovery (%)
Learning (%)
Recovery (%)
Learning (%)
Qwen-0.5B / ko
85,143
66.1
−1.7
−18.7
−138.9
−1.8
OLMo-2-1B / ko
42,037
74.6
−9.6
−133.6
−74.4
−1.7
Pythia-1.4B / ko
15,542
80.3
+1.6
−7.7
−60.4
−0.9
Pythia-1.4B / math
1,242
72.0
−0.0
−0.1
−1.2
−0.0
Appendix
Table A14: Component ablation: the shared (top-1) or residual component is subtracted from the trained weights and the model re-evaluated; negative recovery means forgetting grows. Rows and top-1 differ from Table A13 because that table subsamples both groups to the smaller one (794 rows at OLMo) while the ablation acts on every absent row, and because the OLMo row counts c(t) over a 40M window there against the run’s full 60M here.
Pythia-6.9B
Pythia-12B
Budget
Ref. forget
Interv.
Reduction
Ref. forget
Interv.
Reduction
40M
0.2383
0.1100
53.9%
0.0443
0.0105
76.2%
80M
0.3106
0.1283
58.7%
0.0703
0.0149
78.9%
120M
0.3114
0.1192
61.7%
0.0757
0.0125
83.5%
160M
0.3051
0.1152
62.3%
0.0755
0.0120
84.1%
Appendix
Table A15: The reduction does not shrink with budget: evaluation checkpoints taken inside 160M token runs, Korean, replay-free; three seeded runs per arm at 6.9B at its peak rate 3\times10−5 , one at 12B at 1\times10−5 . Learning changes at the same checkpoints stay within +2.2% (6.9B) and +0.6% (12B). The 40M row is a checkpoint of a run still scheduled to 160M, so its learning rate has not yet decayed; this is why it forgets more than the independent 40M runs of Table 5 and why neither those nor the matching rate column of Table A11 are comparable with it.
Setting
Arm
Forgetting reduction (%)
Learning change (%)
Qwen2.5-0.5B / Alpaca
output ϵ=1\text{\times}{10}^{-4}$$
12.1
+0.8
body ϵ=1\text{\times}{10}^{-5}$$
14.5
+3.0
body ϵ=1\text{\times}{10}^{-4}$$
35.6
−0.3
Pythia-1.4B / Alpaca
output ϵ=1\text{\times}{10}^{-4}$$
5.0
−0.3
body ϵ=1\text{\times}{10}^{-4}$$
32.5
−6.7
Appendix
Table A16: Scoped ϵ on in-distribution Alpaca across sites (replay-free, 12.6M tokens, lr 1\times10−5 , 3 seeds). Each cell is measured against its own reference: Qwen forget +0.0620 nats and learn +0.1685 ; Pythia-1.4B forget +0.1356 and learn +0.3256 .
Arm
Forgetting
Learning
n
Qwen2.5-0.5B, tied
Full fine-tuning reference
+0.6285
+0.3764
3
Full fine-tuning + output ϵ
+0.2019
+0.3968
3
Adapter, output head frozen
+0.1210
+0.2326
3
Adapter, output head released
+2.8204
+0.1938
3
Adapter, head released + output ϵ
+0.2517
+0.2660
3
Appendix
Table A17: Low-rank adapters at two base models on Korean, replay-free, 40M tokens, r=16 , adapter arms at their swept peak lr 1\times10−3 , three seeds per arm; the head-freeze controls carry no adapter. Qwen ties its embeddings and TinyLlama does not. The Qwen full fine-tuning rows are at lr 1\times10−4 and are shown for scale only. Forgetting and learning in nats.
Figure A3: Reduction–learning plane (Qwen/Korean, replay-free, 3 seeds). Shaded band is reference seed spread; methods below the axis are marked at bottom and quantified in legend.
Method
Setting
Forgetting reduction (%)
Learning change (%)
Output-only ϵ (ours)
ϵ=1\text{\times}{10}^{-4}$$
58.3
+0.7
Under-exposed freeze
K=1,000
58.4
+0.6
Step-wise gating
57.6
+0.7
Input-only ϵ
ϵ=1\text{\times}{10}^{-4}$$
−0.5
−0.0
L2 -to-base
λ=0.1
0.5
−0.1
L2 -to-base
λ=1
5.8
−0.3
Appendix
Table A18: The baseline suite at OLMo-2-1B / Korean, lr 1\times10−4 , 60M tokens, 1% replay, three seeds per arm, against one reference (forget +0.2455 nats, learn +0.7909 ). Reductions above 100% mean the arm left the old-domain loss below the base model’s.
Setting
lr
Budget
Ref. forget
Interv. forget
Reduction (%)
Learning change (%)
Pythia-160M
1\times10−5
40M
+0.1876
−0.0287
115.3
−0.6
Pythia-410M
3\times10−5
40M
+0.0803
+0.0356
55.7
+0.4
Pythia-410M
1\times10−4
40M
+0.3248
+0.1938
40.3
−0.9
Pythia-1.4B
3\times10−5
40M
+0.0832
+0.0391
53.1
+4.5
Qwen2.5-0.5B
3\times10−5
40M
+0.1806
+0.0493
72.7
+18.4
Pythia-410M
1\times10−4
400M
+1.2701
+0.5557
56.2
+10.1
Appendix
Table A19: Code injection, replay-free, output-only ϵ=1\text{\times}{10}^{-4}$$ , arms paired within campaign. Rows above the rule are at each campaign’s own rate. Above 100% the old-domain loss ended below the base model’s.
Setting
v^ range
Share
Dist. ( ϵ=1\text{\times}{10}^{-8}$$ )
Dist. ( ϵ=1\text{\times}{10}^{-4}$$ )
Damping
Net/total
Pythia-410M / ko
<1\text{\times}{10}^{-6}$$
6.0%
1.75\times10−3
3.59\times10−5
48.7
0.702
1\times10−6 – 1\times10−5
14.5%
2.25\times10−3
9.95\times10−5
22.6
0.453
1\times10−5 – 1\times10−4
77.5%
2.49\times10−3
5.60\times10−4
4.4
0.320
>1\text{\times}{10}^{-4}$$
2.0%
2.43\times10−3
1.38\times10−3
1.8
0.288
Pythia-1.4B / ko
<1\text{\times}{10}^{-6}$$
3.3%
1.57\times10−3
3.27\times10−5
48.0
0.673
1\times10−6 – 1\times10−5
28.9%
2.36\times10−3
1.55\times10−4
15.2
0.353
Appendix
Table A20: Mean path length per parameter over 120 steps, whole-model v^ bands, replay-free, at the campaign rate (410M and 1.4B 1\times10−4 , Qwen 3\times10−5 ; Qwen’s own peak is 1\times10−4 ). Damping is the ratio of the two distances; net/total is at the default ϵ (1.0 = pushed the same way every step, 0.09 = random walk).
Setting
Site
Params (M)
Zero
<1\text{\times}{10}^{-6}$$
1\times10−6 – 1\times10−5
1\times10−5 – 1\times10−4
Median
Pythia-160M
output emb.
38.6
0.0
79.4
16.3
3.9
2.6\times10−8
input emb.
38.6
36.6
79.9
14.9
4.7
1.1\times10−7
body
85.1
0.0
1.4
8.1
76.5
4.2\times10−5
Pythia-410M
output emb.
51.5
0.0
75.2
18.2
6.2
6.8\times10−8
input emb.
51.5
36.6
80.1
15.8
3.8
1.0\times10−7
body
302.3
0.0
0.9
8.6
89.7
2.6\times10−5
Appendix
Table A21: Distribution of v^ within each site at the end of the reference run (Korean, replay-free, 40M tokens, seed 0, campaign rate, which is 3\times10−5 at Qwen against its peak 1\times10−4 , so its band shares do not reproduce the peak-rate mass held in Table 3 ). Band columns are percentages of that site, computed from the float32 dump of the final second moment; the Zero column counts coordinates that are exactly zero and is contained in the <1\text{\times}{10}^{-6}$$ column beside it, so the two must not be added. Qwen ties its embeddings and has a single table. Medians are over nonzero coordinates.
Figure A4: The same embeddings at two moments, Qwen / Korean under 1% replay. Governed during training, 72.1% of the forgetting is removed at no cost (dashed, the ϵ arm of Table A9 ); restored after training, the edit is a loss from the first checkpoint on and degrades monotonically thereafter.
Models trained on a new task typically degrade on prior tasks, a phenomenon known as forgetting. Traditionally, mitigating forgetting has required replaying stored exemplars from prior tasks, which is often impractical. By contrast, language models can sample from their own training distribution, and we show that these self-generated samples serve as effective replay data, nearly eliminating forgetting. We find that forgetting nonetheless persists when the model has little remaining capacity: models pretrained close to saturation cannot absorb new information without overwriting prior knowledge. When capacity is not the limiting factor, low learning rates reduce forgetting but require substantially more training steps. Replay breaks this tradeoff, enabling fast, high-learning-rate finetuning without forgetting.
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.
Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.
Vedant Palit, Florent Draye, Nicolas Zucchet +2
MPI for Intelligent Systems, Tübingen · Jinesis Lab, University of Toronto & Vector Institute · EuroSafeAI +3