Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models
Authors: Florian Eichin, Philipp Mondorf, Andrei Mircea, Yupei Du, Barbara Plank, Michael A. Hedderich
Organizations: MaiNLP, Center for Information and Language Processing, LMU Munich & Munich Center for Machine Learning (MCML), Germany · Mila – Quebec AI Institute & University of Montreal, Canada · Saarland University, Germany
Memorization has been proposed as a mechanism to explain how language models fit the tail of their training distributions, but its training dynamics are not understood well. In this work, we take a fine-grained look at memorization by decomposing the loss trajectory of memorized sequences over training and model parameters. Across the Pythia family, we study memorization of duplicated training sequences (recitation) and rare ones (recollection). We find that memorization in both cases is characterized by sequence-level gradient alignment, though recitation suffers from misalignment with other training influences which causes forgetting, explaining the necessity for higher duplication of these examples. We further show that the lower model layers are the most involved in memorization and forgetting. Predicting memorization, our decomposition improves over a cross-entropy baseline, especially in larger models and early in training. Intervening on a small set of highly influential parameters we are able to ablate memorization in the final model. Together, these findings advance our understanding of how memorization develops during training and offer insights for predicting and intervening on it.
Figures & tables
Figure 1: Pythia-1B training dynamics. We plot the loss and extractability of our sampled sequences over time and find that a lot of memorization happens early in training. Other models in App. Fig. 9 .
Figure 2: Pythia-1B update magnitude, test sensitivity, and alignment. We plot the mean and standard deviation over the samples. See Eq. 5 for definitions. Other models in App. Fig. 14 and 12
Figure 3: Memorization classification. We plot the mean accuracy on held-out set and mean weights of the decomposition features of the decomposition and CE+decomposition. Others in App. Fig. 19 .
Figure 4: Pythia-1B update magnitude and test sensitivity by layer depth. Others in App. Fig. 14 .
Figure 5: Pythia-1B alignment by layer depth and type. Other models, control in App. Fig. 15 , 16 .
Figure 6: Pythia-1B train/weight decay alignment by model region. Other models in App. Fig. 17
Figure 7: Pythia-1B memorization ablation. Left: Number of 32-extractable sequences in the memorized sample, right: CE loss on the control sample. Other models in App. Fig. 20 .
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Alignment with other injected instances. Other models in App. Fig. 18
Figure 9: Pythia training dynamics across scales. We plot the loss over our sampled sequences and 32-extractability over time.
Figure 10: Pythia BLiMP accuracy across scales.
Figure 11: Pythia update magnitudes and test sensitivities across scales. We plot each at early and later pretraining.
Figure 12: Pythia self alignment across scales. We plot early and later training stages.
Figure 13: Pythia weight decay alignment across scales. We plot the alignment of the instances we decompose with the parameter update introduced by weight decay, early and late in training.
Figure 14: Pythia update magnitudes and test sensitivity by model region across scales.
Figure 15: Pythia self alignment by model region across scales.
Figure 16: Pythia self alignment by layer type across scales.
Figure 17: Pythia train and regularization alignment by model region across scales.
Figure 18: Alignment with other injected instances across scales.
Figure 19: Memorization vs. not memorization classification across model scales. We also plot the weights of the decomposition features for the decomposition only and CE+decomposition model.
Figure 20: Pythia memorization ablations across scales. Left: Number of 32-extractable sequences in the memorized sample, right: CE loss on the control sample.
Figure 21: Pythia memorization ablations affected parameters, embeddings included. Ablation at step 10,000 with 0.00004 of all model parameters ablated.
Figure 22: Pythia memorization ablations affected parameters, embeddings excluded. Ablation at step 10,000 with 0.00004 of all model parameters ablated.
Figure 23: Attention head ablation in Pythia (1/2). We ablate each attention head, one by one, and plot number of 32-extractable memorized sequences and CE on control examples in intervened model.
Figure 24: Attention head ablation in Pythia (2/2). We ablate each attention head, one by one, and plot number of 32-extractable memorized sequences and CE on control examples in intervened model.
Models trained on a new task typically degrade on prior tasks, a phenomenon known as forgetting. Traditionally, mitigating forgetting has required replaying stored exemplars from prior tasks, which is often impractical. By contrast, language models can sample from their own training distribution, and we show that these self-generated samples serve as effective replay data, nearly eliminating forgetting. We find that forgetting nonetheless persists when the model has little remaining capacity: models pretrained close to saturation cannot absorb new information without overwriting prior knowledge. When capacity is not the limiting factor, low learning rates reduce forgetting but require substantially more training steps. Replay breaks this tradeoff, enabling fast, high-learning-rate finetuning without forgetting.
Language models (LMs) can "memorize" information, i.e., encode training data in their weights in such a way that inference-time queries can lead to verbatim regurgitation of that data. This ability to extract training data can be problematic, for example, when data are private or sensitive. In this work, we investigate methods to mitigate memorization: three regularizer-based, three finetuning-based, and eleven machine unlearning-based methods, with five of the latter being new methods that we introduce. We also introduce TinyMem, a suite of small, computationally-efficient LMs for the rapid development and evaluation of memorization-mitigation methods. We demonstrate that the mitigation methods that we develop using TinyMem can successfully be applied to production-grade LMs, and we determine via experiment that: regularizer-based mitigation methods are slow and ineffective at curbing memorization; fine-tuning-based methods are effective at curbing memorization, but overly expensive, especially for retaining higher accuracies; and unlearning-based methods are faster and more effective, allowing for the precise localization and removal of memorized information from LM weights prior to inference. We show, in particular, that our proposed unlearning method BalancedSubnet outperforms other mitigation methods at removing memorized information while preserving performance on target tasks.
Mansi Sakarvadia, Aswathy Ajith, Arham Khan +6
University of Chicago · Argonne National Laboratory · Lawrence Berkeley National Laboratory +3
As new knowledge rapidly accumulates, language models (LMs) with pretrained knowledge quickly become obsolete. A common approach to updating LMs is fine-tuning them directly on new knowledge. However, recent studies have shown that fine-tuning for memorization may be ineffective in storing knowledge or may exacerbate hallucinations. In this work, we introduce a setting we call continual memorization, where a model must memorize and retain a set of factoids through multiple stages of fine-tuning on subsequent datasets. We characterized the forgetting patterns through extensive experiments and show that LMs widely suffer from forgetting, especially when needing to memorize factoids in the second stage. We posit that forgetting can be alleviated by modifying training dynamics: (1) protecting the memorization process when learning factoids or (2) reducing interference from subsequent training stages. Intriguingly, we find that mixing randomly generated word sequences or generic data sampled from pretraining corpora at different training stages effectively mitigates forgetting REMIX: Random and Generic Data Mixing). REMIX can recover performance from severe forgetting, outperforming replay methods and other continual learning baselines. We analyze how REMIX influences the learning process and find that robust memorization follows a distinct pattern: the model stores factoids in earlier layers than usual and diversifies the layers that retain them, which results in easier recall and manipulate of the learned factoids.
Howard Chen, Jiayi Geng, Adithya Bhaskar +2
Princeton Language and Intelligence (PLI), Princeton University