cs.LGOct 7, 2026

Sequential Pretraining Favors Large Models

Authors: Mohnish Harwani, Yujia Zheng

Organizations: Purdue University · University of Illinois Urbana–Champaign

Abstract

Large neural networks often acquire capabilities that small models fail to learn. Does this stem from large models learning more representative features, or from being more robust to unaccounted-for adverse effects introduced during training? We define and quantify one such adverse effect, primacy bias, as the extent to which exposure to early data distributions impairs later learning. We show that small models can allocate learning capacity inefficiently toward early distributions, whereas sufficiently overparameterized models are robust to this effect. This inefficiency is particularly consequential in pretraining, where foundation models often encounter heterogeneous data distributions sequentially rather than jointly. As a result, small foundation models can struggle to learn distributions encountered late in training, which is particularly harmful when later data emphasizes desirable capabilities such as code, mathematics, and reasoning. Motivated by these findings, we introduce Exposure Therapy (ET), a simple regularization that promotes more efficient allocation of learning capacity during sequential pretraining. We demonstrate that ET improves foundation models' performance on late data distributions as well as overall capability in models up to the billion-parameter scale. Overall, our results suggest that some benefits of large foundation models may arise from greater robustness to adverse training effects, rather than from learning more representative features, and that improved training algorithms can recover some of these advantages in smaller models.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Early Data Exposure Improves Robustness to Subsequent Fine-Tuning

    May 12, 2026Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer +2Model Fine-TuningRetention

  2. Small Initialization Matters for Large Language Models

    Jun 16, 2026Liangkai Hang, Junjie Yao, Zhiyu Li +3CapacityDevelopmental Trajectories

  3. Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

    May 28, 2026Jing Huang, Daniel Wurgaft, Rachit Bansal +6Model SizeCapacity