Replay on Demand: An Emergent Curriculum for Balancing Adaptation and Forgetting in Continued Pretraining
Authors: Lukas Thede, Shengzhuang Chen, Stefan Winzeck, Matthias Bethge, Zeynep Akata, Jonathan Richard Schwarz
Organizations: University of Tübingen, Tübingen AI Center · Helmholtz Munich · Munich Center for Machine Learning (MCML) · Thomson Reuters Foundational Research · Imperial College London · Technical University of Munich
Continued pretraining enables language models to adapt to new domains and knowledge, but often at the cost of forgetting previously acquired capabilities. Replay can mitigate this trade-off, but fixed replay mixtures allocate training independently of the model's actual retention needs. We introduce Replay on Demand (RoD), which instead derives the replay allocation from the model's learning dynamics. RoD jointly prioritizes adaptation samples by their remaining learning potential and replay samples by their observed forgetting. Their competition for a shared training budget yields an online curriculum that determines what to train on at each step. Across models, scales, and adaptation domains, RoD reaches or improves upon the adaptation-forgetting frontier of tuned fixed-replay baselines and model merging without prescribing a replay allocation in advance. Replay concentrates on sources that are more vulnerable to forgetting and dynamically increases and redistributes as forgetting emerges during training. Together, our results show that replay can be allocated online from the model's evolving state, targeting what is needed, when it is needed.
Figures & tables
Figure 1 : RoD dynamically balances adaptation and replay through joint data selection. Adaptation candidates are scored by their remaining learning potential and replay candidates by their forgetting. Joint top- k selection lets both compete for a shared training budget, dynamically determining the replay share and composition. As adaptation learning potential decreases and forgetting emerges, replay becomes increasingly competitive and receives a larger share of the training budget.
Figure 2 : RoD reaches or improves upon the adaptation–forgetting frontier. Adaptation loss is shown against general validation-loss forgetting (top) and task-based forgetting (bottom); lower is better on both axes. Fixed-replay CPT traces the frontier as replay share varies, while RoD reaches or improves upon it without specifying a replay ratio in advance. Nemotron uses native replay from its pretraining data, whereas Qwen uses the Nemotron data as proxy replay because its pretraining data are unavailable. No-replay CPT and model merging provide additional baselines; max replay is a higher-budget reference trained with approximately 2× the trained-token budget. Insets magnify the operating region around RoD and the strongest baselines.
Figure 3 : RoD allocates replay according to forgetting demand. On German adaptation with Nemotron-12B, RoD concentrates protection on the knowledge categories most vulnerable to forgetting (a) by allocating more replay to these categories during training (b).
Figure 4 : RoD dynamically determines what to replay, how much, and when. Stacked areas show the fraction of each training batch allocated to different knowledge categories, while the line shows the total replay share. Replay emerges as forgetting develops and is rebalanced throughout training, while its composition evolves across knowledge categories and model–domain settings. RoD thus adapts both the amount and composition of replay over training.
Figure 5 : Joint adaptation–replay competition drives RoD’s gains. Joint competition improves over fixed-allocation selection, while trade-off and replay allocation stabilize when unconstrained.
Method
Adapt. ↓
Val. forget. ↓
Task forget. ↓
Qwen3.5-9B target (4B → 9B; 2.25× )
No-replay
1.793
0.419
0.108
Fixed replay (32%)
1.807
0.019
0.039
RoD (native)
1.790
0.028
0.061
RoD (4B ref.)
1.796
0.013
0.047
RoD (4B curr.)
1.795
0.019
0.052
Table 1: Cross-scale RoD. Larger target models use native RoD or components constructed by a smaller model ( ref. : adaptation reference; curr. : curriculum). No-replay and fixed replay provide reference points.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Capability
Evaluation tasks
Parametric knowledge
MMLU World Religions, Prehistory, High School Geography, High School Government and Politics, US Foreign Policy, Security Studies, Global Facts, Miscellaneous, Clinical Knowledge, Medical Genetics, Professional Medicine, Anatomy, Management, Marketing, Business Ethics, and Professional Accounting; SciQ; OpenBookQA; ARC-Easy; MedQA; MMLU-Pro.
Reasoning and problem solving
BBH; MuSR; MMLU Formal Logic, Logical Fallacies, Econometrics, High School Microeconomics, High School Macroeconomics, High School Mathematics, College Mathematics, Elementary Mathematics, Abstract Algebra, and High School Statistics; ARC-Challenge; AQuA-RAT; SAT Math.
Commonsense and robustness
HellaSwag; PIQA; WinoGrande; CommonsenseQA.
Multilingual capabilities
multilingual MMLU in German, French, Spanish, Russian, Chinese, and Arabic; BeleBele in German, French, Spanish, Russian, Chinese, Arabic, and English.
Appendix
Table 2: Task-based capability evaluation. Benchmarks used to evaluate the four pretraining-time capability groups considered in our experiments.
Setting
Method
Replay share
Adapt. ↓
Val. forget. ↓
Task forget. ↓
Tokens (B)
Legal ⋅ Nemotron-12B
Base
–
1.492
0.000
0.000
0.00
No-replay
0%
0.972
0.553
0.122
12.41
Fixed replay (10%)
10%
0.967
0.125
0.067
12.41
Fixed replay (20%)
20%
0.970
0.082
0.055
12.41
Fixed replay (50%)
50%
0.992
0.024
0.036
12.41
Max-replay (off-budget)
50%
0.976
0.057
0.046
24.31
Appendix
Table 3: Full adaptation–forgetting frontier, all settings. Adaptation loss, general validation-loss forgetting, and task-based forgetting (capped mean accuracy drop over four capabilities) for every arm plotted in fig. 2 . All arms within a setting are compute-matched (same optimizer steps from base) except max-replay , trained to its own epoch target at ≈2× the tokens. Replay share is the realized fraction of the trained batch for fixed-replay arms and the emergent fraction for RoD; model merging has no replay share (weight-space interpolation). Val. forget. and task forget. are ↓ (0 = no forgetting); adapt. is ↓ . Legal · Qwen3.5-9B has no off-budget max-replay arm (never trained).
Figure 6 : Fixed-replay CPT trained to convergence. Adaptation–forgetting trade-offs for German adaptation when fixed-replay CPT is trained beyond the matched trained-token ( 1× ) budget until convergence. Filled circles show the trained-token-matched fixed-replay runs used in the main comparison, while open circles show the corresponding runs trained to convergence; annotations indicate their training budget relative to the trained-token-matched runs. RoD (star) is shown at the 1× compute budget. Extending fixed-replay training improves adaptation but does not recover a consistently better adaptation–forgetting trade-off than RoD, while requiring up to 2.27× the training budget.
Setting
Budget
Adapt. ↓
Val. forget. ↓
German ⋅ Nemotron-12B
1.00×
1.7878
0.0731
1.19×
1.7882
0.0734
Legal ⋅ Nemotron-12B
1.00×
0.9670
0.0714
1.20×
0.9705
0.0725
Appendix
Table 4: RoD performance beyond the compute-matched training budget. Continuing RoD beyond the 1× budget does not improve either adaptation or retention, indicating that RoD has already converged at the budget used in our main comparison.
Setting
World knowl.
Reason./math
Multiling.
Code
Overall
Legal ⋅ Nemotron-12B
1.30
0.49
1.07
0.40
0.81
German ⋅ Nemotron-12B
1.36
0.49
1.08
0.43
0.90
Legal ⋅ Qwen3.5-9B
1.47
0.60
1.51
0.53
1.03
German ⋅ Qwen3.5-9B
1.46
0.60
1.51
0.53
1.02
German ⋅ Qwen3.5-4B
1.55
0.64
1.61
0.56
1.08
Appendix
Table 5: Base-model loss on the replay distribution. We report held-out language-modeling loss on the Nemotron replay corpus overall and across its four broad source categories. Nemotron is evaluated on replay data from its own pretraining distribution, whereas the same corpus serves as proxy replay for Qwen3.5. The higher Qwen losses indicate that the Qwen base models are less converged on the proxy replay distribution.
Figure 7 : Task-based forgetting with proxy replay. Change in capability accuracy relative to the pretrained Qwen3.5-9B base model after German adaptation; values closer to zero indicate stronger retention. RoD substantially reduces forgetting relative to no-replay CPT across all capability groups. Its remaining gap to fixed replay is concentrated primarily in reasoning and mathematics and, to a lesser extent, multilingual evaluations, many of which contain translated reasoning and mathematical tasks.
Figure 8 : RHO-score dynamics across model–domain settings. Mean adaptation RHO score ρA and replay RHO score ρR over training, together with the resulting replay share. Across settings, adaptation scores decrease as the model learns the adaptation distribution, while replay scores initially increase as forgetting accumulates. The resulting competition between both signals dynamically adjusts the amount of replay throughout training. Vertical changes in replay share coincide with transitions between data passes and the learning-rate annealing phase.
Figure 9 : Source-level forgetting profiles across model–domain settings. Each point represents one replay-data source. The x -axis measures source-level forgetting under no-replay CPT, capturing how vulnerable a source is to forgetting, while the y -axis measures the forgetting that remains after applying replay. Lines show linear fits for fixed-replay CPT and RoD; the reported slopes summarize how strongly remaining forgetting depends on the source’s original vulnerability. A slope near one indicates approximately uniform reduction of the no-replay forgetting profile, whereas a flatter slope indicates stronger relative protection of sources that would otherwise forget most. Across settings, RoD consistently produces the flattest forgetting profile.
Figure 10 : Replay allocation follows source-level forgetting across model–domain settings. Each point represents one replay-data source. The x -axis measures source vulnerability as forgetting under no-replay CPT, while the y -axis shows its sampling ratio under RoD relative to its prevalence in the replay candidate distribution; values above 1 indicate oversampling and values below 1 undersampling. Colors denote broad pretraining-data categories. Dashed lines show linear fits and r denotes the Pearson correlation between vulnerability and RoD sampling ratio. Across settings, RoD preferentially allocates replay to sources that are more susceptible to forgetting.