cs.LGSep 28, 2026

Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models

Authors: Florian Eichin, Philipp Mondorf, Andrei Mircea, Yupei Du, Barbara Plank, Michael A. Hedderich

Organizations: MaiNLP, Center for Information and Language Processing, LMU Munich & Munich Center for Machine Learning (MCML), Germany · Mila – Quebec AI Institute & University of Montreal, Canada · Saarland University, Germany

Abstract

Memorization has been proposed as a mechanism to explain how language models fit the tail of their training distributions, but its training dynamics are not understood well. In this work, we take a fine-grained look at memorization by decomposing the loss trajectory of memorized sequences over training and model parameters. Across the Pythia family, we study memorization of duplicated training sequences (recitation) and rare ones (recollection). We find that memorization in both cases is characterized by sequence-level gradient alignment, though recitation suffers from misalignment with other training influences which causes forgetting, explaining the necessity for higher duplication of these examples. We further show that the lower model layers are the most involved in memorization and forgetting. Predicting memorization, our decomposition improves over a cross-entropy baseline, especially in larger models and early in training. Intervening on a small set of highly influential parameters we are able to ablate memorization in the final model. Together, these findings advance our understanding of how memorization develops during training and offer insights for predicting and intervening on it.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mitigating Memorization In Language Models

    Oct 3, 2024Mansi Sakarvadia, Aswathy Ajith, Arham Khan +6MemorizationLarge Language Model Unlearning

  2. Continual Memorization of Factoids in Language Models

    Nov 11, 2024Howard Chen, Jiayi Geng, Adithya Bhaskar +2MemorizationModel Fine-Tuning