A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks
Organizations: Seoul National University of Science and Technology · Korea Institute of Land & Infrastructure Safety Technology
Abstract
Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data. In this data-free regime, we analyze where forgetting occurs and why. Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere. This localization is governed by the vocabulary deficiency of the corpus rather than the training mode, allowing pre-retraining risk ranking from token counts alone within a fixed base model. Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam's second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates. We therefore propose an intervention: raising Adam's epsilon exclusively for the output projection during training. Across eight settings spanning 160M to 12B parameters and four model families, this removes 39.4% to 67.9% of forgetting across all seven stable configurations without degrading target learning or requiring per-model tuning. The defense combines additively or better with replay (79.8% on Qwen/Korean) and rescues released-head LoRA from a 23-fold forgetting surge. Because post-hoc editing of the drifted rows recovers under 5% of forgetting, the intervention must operate during training. Our findings indicate that a single-line optimizer adjustment may serve as the primary defense against catastrophic forgetting where the corpus starves the vocabulary.
Figures & tables
| Base model (architecture) | Primary domain | Additional domains |
|---|---|---|
| Pythia 160M–12B (GPT-NeoX) | Korean | math, code, Alpaca, KoAlpaca (410M and 1.4B only) |
| Qwen2.5-0.5B (tied) | Korean | math, code, Alpaca, KoAlpaca |
| TinyLlama-1.1B (Llama, untied) | Korean | – |
| OLMo-2-1B (olmo2, untied) | Korean | – |
| Pythia-410M | Pythia-1.4B | Qwen2.5-0.5B | ||||
|---|---|---|---|---|---|---|
| Alpaca | KoAlpaca | Alpaca | KoAlpaca | Alpaca | KoAlpaca | |
| Attention freeze | 16.0 ( ) | 42.7 ( ) | 7.0 ( ) | 30.2 ( ) | 17.2 ( ) | 15.5 ( ) |
| MLP freeze | 27.9 ( ) | 8.2 ( ) | 39.7 ( ) | 47.8 ( ) | 44.8 ( ) | 26.1 ( ) |
| Output-embedding freeze | 2.7 ( ) | ( ) | 0.3 ( ) | ( ) | 11.6 ( ) † | 46.9 ( ) † |
| Input-embedding freeze | 1.6 ( ) | 6.6 ( ) | 0.8 ( ) | ( ) | n/a | n/a |
| Output-scoped | 14.6 ( ) | 29.8 ( ) | 5.6 ( ) | 33.1 ( ) | 6.0 ( ) | 43.3 ( ) |
| Setting | Tokens/param | Ref. forget (nats) | Interv. forget | Reduction (%) | Learning change (%) |
|---|---|---|---|---|---|
| Pythia-160M ∗ | 0.2469 | 1.448 | 0.388 | 73.2 | |
| Pythia-410M | 0.0988 | 0.664 | 0.373 | 43.8 | |
| Qwen2.5-0.5B | 0.0810 | 0.629 | 0.202 | 67.9 | |
| TinyLlama-1.1B | 0.0364 | 0.361 | 0.187 | 48.1 | |
| Pythia-1.4B | 0.0284 | 0.473 | 0.261 | 44.8 | |
| OLMo-2-1B | 0.0269 | 1.137 | 0.689 | 39.4 |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | lr | Ref. forget | Out. params (M) | In. params (M) | In. zero (%) | Body params (M) |
|---|---|---|---|---|---|---|
| 160M | 3.392 | 30.7 | 30.9 | 45.8 | 1.2 | |
| 410M | 0.656 | 38.7 | 41.3 | 45.7 | 2.7 | |
| 1.4B | 0.472 | 80.6 | 88.0 | 42.8 | 15.9 | |
| TinyLlama | 0.355 | 48.2 | 50.4 | 45.3 | 13.1 | |
| OLMo | 1.160 | 176.6 | 199.6 | 52.8 | 6.9 | |
| Qwen | 0.628 | 110.0 | n/a | n/a | 1.0 |
| lr | Ref. forget (nats) | Interv. forget | Reduction (%) | Learning change (%) | (ref/interv) |
|---|---|---|---|---|---|
| 0.491 | 0.126 | 74.2 | 3/3 | ||
| 1.448 | 0.388 | 73.2 | 3/3 | ||
| 2.068 | 1.855 | 10.3 | 6/6 | ||
| 3.108 | 3.795 | 6/6 | |||
| 5.233 | 4.681 | 10.5 | 2/3 |
| Site | Band | 160M | 160M † | 410M | 1.4B | TinyLlama | OLMo | Qwen |
|---|---|---|---|---|---|---|---|---|
| Output emb. | -6.2 (-13.4) | 68.8 (+3.1) | 31.4 (+5.2) | 33.7 (+2.0) | 37.0 (+0.1) | 31.3 (+1.6) | 51.0 (+4.4) | |
| – | -13.6 (-9.7) | -20.1 (-0.8) | 2.3 (+6.3) | 1.1 (+0.1) | 3.6 (-0.4) | 0.3 (-0.1) | 10.7 (+1.3) | |
| – | -10.9 (-7.7) | -17.8 (+2.2) | -5.6 (+0.6) | -1.0 (-0.2) | -3.5 (-0.5) | 1.3 (-0.4) | 1.1 (-0.8) | |
| Input emb. | -3.6 (-4.6) | 7.3 (-1.8) | 1.4 (+9.0) | -1.5 (+0.1) | -0.6 (-0.5) | 0.4 (-0.2) | n/a | |
| – | -17.9 (-6.5) | 6.8 (-2.4) | 3.3 (+8.1) | -1.5 (+0.2) | -0.3 (-0.4) | 1.9 (-0.2) | n/a | |
| – | -18.7 (-9.3) | 12.5 (+0.2) | 7.2 (+9.9) | -0.5 (+0.1) | 1.3 (-0.2) | -1.3 (-0.2) | n/a |
| Setting | Exactly 0 (%) | 1\text{\times}{10}^{-8}$$ (%) | 1\text{\times}{10}^{-6}$$ (%) | Zero, input emb. (%) |
|---|---|---|---|---|
| Pythia-160M | 8.7 | 15.4 | 32.8 | 36.6 |
| Pythia-410M | 4.7 | 4.6 | 16.5 | 36.6 |
| Pythia-1.4B | 2.7 | 3.5 | 10.7 | 36.6 |
| TinyLlama-1.1B | 2.1 | 3.0 | 8.3 | 34.8 |
| OLMo-2-1B | 7.1 | 11.2 | 20.1 | 51.2 |
| Qwen2.5-0.5B | 0.0 | 5.0 | 22.5 | n/a |
| Setting | lr | 1\text{\times}{10}^{-5}$$ | 1\text{\times}{10}^{-4}$$ | 1\text{\times}{10}^{-3}$$ |
|---|---|---|---|---|
| Pythia-160M | ( ) | ( ) | ( ) | |
| Pythia-410M | 41.5 ( ) | 43.8 ( ) | 32.7 ( ) | |
| Pythia-1.4B | 40.9 ( ) | 44.8 ( ) | 45.6 ( ) |
| Method | Setting | Forgetting reduction (%) | Learning change (%) |
|---|---|---|---|
| Output-only (ours) | 1\text{\times}{10}^{-4}$$ | 69.5 | |
| Under-exposed freeze | 70.7 | ||
| Step-wise gating | 70.2 | ||
| -to-base | 11.2 | ||
| -to-base | 29.3 | ||
| -to-base | 64.3 |
| Base model | wd | Reduct. | Learn. | |
|---|---|---|---|---|
| Pythia-410M | 0 | 38.4% | % | 6 |
| Pythia-410M | 0.1 | 32.7% | % | 6 |
| Qwen2.5-0.5B | 0 | 68.0% | % | 3 |
| Qwen2.5-0.5B | 0.1 | 67.9% | % | 3 |
| Forgetting reduction (%) | |||||
|---|---|---|---|---|---|
| Setting | lr | Output-only | Freeze, | Step-wise gating | Max gap (pp) |
| Qwen2.5-0.5B | 69.5 | 70.7 | 70.2 | 1.2 | |
| Pythia-1.4B | 44.8 | 45.1 | 43.7 | 1.5 | |
| TinyLlama-1.1B | 48.1 | 50.1 | 48.2 | 2.0 | |
| OLMo-2-1B | 40.7 | 40.7 | 38.1 | 2.6 | |
| Pythia-410M | 43.8 | 38.1 | 43.7 | 5.7 | |
| Arm | Forgetting reduction (%) | Learning change (%) |
|---|---|---|
| Freeze the whole output projection | 74.4 | |
| Freeze a random 96.6% of the embeddings | 66.4 | |
| Freeze by count, (96.6%) | 68.2 | |
| Output-only 1\text{\times}{10}^{-4}$$ | 72.1 | |
| Count-freeze and together | 73.2 |
| Arm | Forgetting reduction (%) | Learning change (%) |
|---|---|---|
| embedding learning rate | 67.5 | |
| embedding 5\text{\times}{10}^{-5}$$ | 66.7 | |
| embedding 2\text{\times}{10}^{-5}$$ | 64.2 | |
| embedding 5\text{\times}{10}^{-6}$$ | 57.9 | |
| embedding 2\text{\times}{10}^{-6}$$ | 52.0 | |
| embedding learning rate | 51.3 |
| Setting | |||||
|---|---|---|---|---|---|
| Pythia-160M | 73.8 ( ) | 71.4 ( ) | 10.0 ( ) | ( ) | 8.0 ( ) |
| Pythia-410M | 50.7 ( ) | 50.6 ( ) | 43.5 ( ) | 15.3 ( ) | 8.8 ( ) |
| Pythia-1.4B | 57.6 ( ) | 52.1 ( ) | 44.9 ( ) | 32.5 ( ) | 6.7 ( ) |
| TinyLlama-1.1B | n/q ( ) | 111.6 ( ) | 47.7 ( ) | 27.9 ( ) | 6.1 ( ) |
| OLMo-2-1B | 44.5 ( ) | 46.0 ( ) | 41.9 ( ) | 34.2 ( ) | 26.9 ( ) |
| Qwen2.5-0.5B | 75.4 ( ) | 69.5 ( ) | 67.7 ( ) | 60.9 ( ) | 30.8 ( ) |
| Pythia-6.9B | Pythia-1.4B | |||||
|---|---|---|---|---|---|---|
| Task | untouched | untouched | ||||
| LAMBADA | 61.05 | 60.00 | 65.17 | 60.78 | 52.53 | 56.79 |
| HellaSwag | 63.24 | 57.89 | 62.06 | 52.06 | 48.08 | 49.90 |
| PIQA | 76.22 | 72.09 | 73.23 | 71.00 | 67.19 | 69.04 |
| ARC-easy | 67.26 | 64.02 | 65.32 | 60.73 | 57.00 | 58.33 |
| WinoGrande | 61.64 | 59.35 | 61.01 | 56.99 | 55.46 | 56.64 |
| Setting | Rows | Top-1, absent (%) | Top-1, control (%) | RMS, absent | RMS, control |
|---|---|---|---|---|---|
| Qwen-0.5B / ko | 85,143 | 66.1 | 3.2 | ||
| OLMo-2-1B / ko | 47,492 | 72.6 | 8.8 | ||
| Pythia-1.4B / ko | 15,542 | 80.8 | 12.9 | ||
| Pythia-1.4B / math | 1,242 | 72.0 | 3.7 |
| Shared removed | Residual removed | |||||
|---|---|---|---|---|---|---|
| Setting | Rows | Top-1 (%) | Recovery (%) | Learning (%) | Recovery (%) | Learning (%) |
| Qwen-0.5B / ko | 85,143 | 66.1 | ||||
| OLMo-2-1B / ko | 42,037 | 74.6 | ||||
| Pythia-1.4B / ko | 15,542 | 80.3 | ||||
| Pythia-1.4B / math | 1,242 | 72.0 | ||||
| Pythia-6.9B | Pythia-12B | |||||
|---|---|---|---|---|---|---|
| Budget | Ref. forget | Interv. | Reduction | Ref. forget | Interv. | Reduction |
| 40M | 0.2383 | 0.1100 | 53.9% | 0.0443 | 0.0105 | 76.2% |
| 80M | 0.3106 | 0.1283 | 58.7% | 0.0703 | 0.0149 | 78.9% |
| 120M | 0.3114 | 0.1192 | 61.7% | 0.0757 | 0.0125 | 83.5% |
| 160M | 0.3051 | 0.1152 | 62.3% | 0.0755 | 0.0120 | 84.1% |
| Setting | Arm | Forgetting reduction (%) | Learning change (%) |
|---|---|---|---|
| Qwen2.5-0.5B / Alpaca | output 1\text{\times}{10}^{-4}$$ | 12.1 | |
| body 1\text{\times}{10}^{-5}$$ | 14.5 | ||
| body 1\text{\times}{10}^{-4}$$ | 35.6 | ||
| Pythia-1.4B / Alpaca | output 1\text{\times}{10}^{-4}$$ | 5.0 | |
| body 1\text{\times}{10}^{-4}$$ | 32.5 |
| Arm | Forgetting | Learning | |
| Qwen2.5-0.5B, tied | |||
| Full fine-tuning reference | 3 | ||
| Full fine-tuning output | 3 | ||
| Adapter, output head frozen | 3 | ||
| Adapter, output head released | 3 | ||
| Adapter, head released output | 3 | ||
| Method | Setting | Forgetting reduction (%) | Learning change (%) |
|---|---|---|---|
| Output-only (ours) | 1\text{\times}{10}^{-4}$$ | 58.3 | |
| Under-exposed freeze | 58.4 | ||
| Step-wise gating | 57.6 | ||
| Input-only | 1\text{\times}{10}^{-4}$$ | ||
| -to-base | 0.5 | ||
| -to-base | 5.8 |
| Setting | lr | Budget | Ref. forget | Interv. forget | Reduction (%) | Learning change (%) |
|---|---|---|---|---|---|---|
| Pythia-160M | 40M | 115.3 | ||||
| Pythia-410M | 40M | 55.7 | ||||
| Pythia-410M | 40M | 40.3 | ||||
| Pythia-1.4B | 40M | 53.1 | ||||
| Qwen2.5-0.5B | 40M | 72.7 | ||||
| Pythia-410M | 400M | 56.2 |
| Setting | range | Share | Dist. ( 1\text{\times}{10}^{-8}$$ ) | Dist. ( 1\text{\times}{10}^{-4}$$ ) | Damping | Net/total |
|---|---|---|---|---|---|---|
| Pythia-410M / ko | 1\text{\times}{10}^{-6}$$ | 6.0% | 48.7 | 0.702 | ||
| – | 14.5% | 22.6 | 0.453 | |||
| – | 77.5% | 4.4 | 0.320 | |||
| 1\text{\times}{10}^{-4}$$ | 2.0% | 1.8 | 0.288 | |||
| Pythia-1.4B / ko | 1\text{\times}{10}^{-6}$$ | 3.3% | 48.0 | 0.673 | ||
| – | 28.9% | 15.2 | 0.353 |
| Setting | Site | Params (M) | Zero | 1\text{\times}{10}^{-6}$$ | – | – | Median |
|---|---|---|---|---|---|---|---|
| Pythia-160M | output emb. | 38.6 | 0.0 | 79.4 | 16.3 | 3.9 | |
| input emb. | 38.6 | 36.6 | 79.9 | 14.9 | 4.7 | ||
| body | 85.1 | 0.0 | 1.4 | 8.1 | 76.5 | ||
| Pythia-410M | output emb. | 51.5 | 0.0 | 75.2 | 18.2 | 6.2 | |
| input emb. | 51.5 | 36.6 | 80.1 | 15.8 | 3.8 | ||
| body | 302.3 | 0.0 | 0.9 | 8.6 | 89.7 |