cs.CLOct 4, 2026

Selecting Repetition Counts Across Model Scales in Data-Constrained Pretraining

Authors: Ziyue WANG, T. Kanamori

Organizations: Institute of Science Tokyo

Abstract

The repetition count that works best for a small language model may not remain best at a larger scale. We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction. On Wikipedia-derived data and Proof-Pile-2, the ranking of measured repetition counts changes with model size, and a 520M Proof-Pile-2 experiment confirms that reducing repetition from sixteen to eight improves loss while using fewer training tokens. We use loss curves from several smaller models to retain a short list of promising repetition counts for evaluation at a larger scale. On PubMed and Caselaw, candidate sets fixed before target-model training retain the lowest-loss measured count on the original evaluation grids at both 200M and 520M. This supports candidate retention as a practical alternative to exact point prediction. We also relate the pruning regression to an empirical scaling model with two opposing repetition-dependent loss terms. A first-order expansion in log model size yields the linear form used by the selection rule, providing a scaling-based interpretation of the candidate-selection procedure.

Figures & tables

Appendix figures & tables51 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Internal Data Repetition Destroys Language Models

    Jun 23, 2026Jessica Chudnovsky, Joshua Kazdan, Noam Levi +6RepetitionScaling Laws

  2. Scaling Laws for Mixture Pretraining Under Data Constraints

    May 12, 2026Anastasiia Sedova, Skyler Seto, Natalie Schluter +1PretrainingScaling Laws

  3. What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining

    Oct 4, 2026Yekun Chai, Haoyi XiongEpochPretraining