When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting
Organizations: MPI for Intelligent Systems, Tübingen · Jinesis Lab, University of Toronto & Vector Institute · EuroSafeAI · Hector Foundation · Stanford University · ELLIS Institute Tübingen
Abstract
Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.
Figures & tables
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Minimal | Transformer | OLMo 2 | ||
|---|---|---|---|---|
| Analysis | model | attention-only | with MLP blocks | 1B |
| Collapse, recovery and erosion | Fig. 2 | Fig. 1 | App. C.5 | Fig. 5 |
| New answers in both regions | Fig. 2 | Fig. 1 | App. C.5 | – |
| Common shift and fact-specific drift | App. B.10 | Fig. 4 | App. C.5 | – |
| Removing the common shift | App. B.9 | Fig. 4 | App. C.5 | – |
| Ordering within the region | App. B.8 | App. C.1 | App. C.5 | – |
| Weights restored | Share of the common shift |
|---|---|
| Attention value and output | |
| Attention query and key | |
| Input embedding | |
| LayerNorms | |
| Unembedding (control) |
| Old facts | New facts | |||||
|---|---|---|---|---|---|---|
| New facts | Step | trained | top removed | random removed | trained | top removed |
| Synthetic individuals | 80 | 0.28 | 0.77 | 0.28 | 0.03 | 0.01 |
| 400 | 0.38 | 0.70 | 0.38 | 0.90 | 0.40 | |
| Real entities | 80 | 0.73 | 0.88 | 0.73 | 0.46 | 0.25 |
| 400 | 0.55 | 0.79 | 0.55 | 1.00 | 0.66 | |