Selecting Repetition Counts Across Model Scales in Data-Constrained Pretraining
Organizations: Institute of Science Tokyo
Abstract
The repetition count that works best for a small language model may not remain best at a larger scale. We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction. On Wikipedia-derived data and Proof-Pile-2, the ranking of measured repetition counts changes with model size, and a 520M Proof-Pile-2 experiment confirms that reducing repetition from sixteen to eight improves loss while using fewer training tokens. We use loss curves from several smaller models to retain a short list of promising repetition counts for evaluation at a larger scale. On PubMed and Caselaw, candidate sets fixed before target-model training retain the lowest-loss measured count on the original evaluation grids at both 200M and 520M. This supports candidate retention as a practical alternative to exact point prediction. We also relate the pruning regression to an empirical scaling model with two opposing repetition-dependent loss terms. A first-order expansion in log model size yields the linear form used by the selection rule, providing a scaling-based interpretation of the candidate-selection procedure.
Figures & tables
| PubMed | Caselaw | |||
|---|---|---|---|---|
| Prediction and evaluation grid | Excess loss | Excess loss | ||
| Original choice, seven counts | 8 | 0.038 | 12 | 0 |
| Original choice, twelve counts | 8 | 0.044 | 12 | 0 |
| Same function on integer , twelve counts | 10 | 0.003 | 11 | 0.008 |
| New-count RMSE | Selection loss | |||
|---|---|---|---|---|
| Model | PubMed | Caselaw | PubMed | Caselaw |
| Adapted Lovelace model, Eq. ( 19 ) | 0.238 | 0.224 | 0.003 | 0.008 |
| Same repetition terms, free | 0.159 | 0.184 | 0.003 | 0.008 |
| Saturated extension, Eq. ( 9 ) | 0.108 | 0.074 | 0.006 | 0.018 |
Appendix figures & tables51 assets
Supplementary material from the paper’s appendix.
Appendix
| Item | Setting |
|---|---|
| Target pool | 100M English-Wikipedia-derived tokens: two 50M-token subpools, one clean-prose construction and one redundancy-injected construction |
| Generic training corpus | FineWeb-derived training stream; generic tokens fill the fraction of each main training configuration |
| Tokenizer and context | Fixed 32,768-token byte-level byte-pair encoding (BPE); context length 1024 |
| Target evaluation | Mean over the two clean held-out target splits; each constituent split uses at most 2,097,152 tokens |
| Held-out generic loss | Fixed held-out generic data, used only for comparisons between configurations within this study |
| Candidate sets | 35M: ; 50M: ; 80M: |
| Item | 35M/50M/80M experiments | 520M experiments |
|---|---|---|
| Target corpus and tokenizer | Fixed 100M-token Proof-Pile-2 formal-mathematics/code target, validation split, and 32,768-token tokenizer | Same corpus, split, and tokenizer |
| Generic training corpus | Fixed FineWeb-derived training stream | Same corpus |
| Target evaluation | Fixed held-out target split; at most 2,097,152 tokens | Same split and maximum |
| Held-out generic loss | Fixed held-out generic data used for within-corpus comparisons | Same evaluation data |
| Candidate sets | 35M: ; 50M: ; 80M: | First experiment: ; second experiment: |
| Seeds | Four per 35M model/count combination; three per 50M and 80M combination | Six paired seeds in each nonoverlapping seed set |
| 1 | 4 | 12 | 24 | 36 | 48 | 60 | |
| Mean | 6.274 | 4.307 | 3.229 | 3.005 | 2.941 | 2.927 | 2.929 |
| Seed SD | 0.026 | 0.019 | 0.011 | 0.004 | 0.003 | 0.002 | 0.004 |
| Seeds | 4 | 4 | 4 | 3 | 4 | 4 | 4 |
| 1 | 4 | 8 | 16 | 24 | 32 | 48 | |
| Mean | 6.256 | 4.118 | 3.330 | 3.014 | 2.930 | 2.910 | 2.943 |
| Seed SD | 0.036 | 0.018 | 0.005 | 0.004 | 0.004 | 0.002 | 0.002 |
| Seeds | 3 | 3 | 3 | 3 | 3 | 3 | 3 |
| 1 | 4 | 8 | 12 | 16 | 24 | 36 | |
| Mean | 6.225 | 3.818 | 3.155 | 2.977 | 2.912 | 2.905 | 3.044 |
| Seed SD | 0.015 | 0.062 | 0.013 | 0.003 | 0.001 | 0.005 | 0.009 |
| Seeds | 3 | 3 | 3 | 3 | 3 | 3 | 3 |
| Scale | Mean loss | Scale | Mean loss | ||
|---|---|---|---|---|---|
| 35M | 6 | 0.913 | 50M | 4 | 1.021 |
| 35M | 12 | 0.783 | 50M | 8 | 0.795 |
| 35M | 24 | 0.804 | 50M | 16 | 0.788 |
| 35M | 48 | 0.950 | 50M | 32 | 0.978 |
| 35M | 64 | 1.070 | 50M | 48 | 1.243 |
| Scale | Mean loss | Seed SD | |
|---|---|---|---|
| 80M | 4 | 0.898 | 0.009 |
| 80M | 8 | 0.762 | 0.004 |
| 80M | 12 | 0.796 | 0.005 |
| 80M | 16 | 0.866 | 0.007 |
| 80M | 24 | 1.101 | 0.007 |
| 80M | 36 | 1.497 | 0.015 |
| Rule | Output at 80M | Interpretation |
|---|---|---|
| Direct reuse of 50M selection | Reuse the 50M minimum | |
| Two-point log-linear extrapolation | 24.22, mapped to | Extrapolate the two selected counts, then use the nearest evaluated count |
| Nearby candidate set | Extrapolated center and its adjacent evaluated counts |
| Seed | ||||
|---|---|---|---|---|
| 1342 | 0.820 | 0.864 | 1.357 | 0.537 |
| 1343 | 0.846 | 0.942 | 1.374 | 0.528 |
| 1344 | 0.878 | 0.884 | 1.368 | 0.490 |
| 1345 | 0.856 | 0.941 | 1.377 | 0.521 |
| 1346 | 0.822 | 0.915 | 1.354 | 0.532 |
| 1347 | 0.834 | 0.955 | 1.373 | 0.539 |
| Seed | |||
|---|---|---|---|
| 1348 | 0.852 | 1.373 | 0.521 |
| 1349 | 0.848 | 1.364 | 0.516 |
| 1350 | 0.854 | 1.362 | 0.509 |
| 1351 | 0.853 | 1.350 | 0.497 |
| 1352 | 0.846 | 1.366 | 0.520 |
| 1353 | 0.880 | 1.369 | 0.489 |
| Seed | Loss order | ||
|---|---|---|---|
| 1342 | 0.493 | 0.044 | |
| 1343 | 0.432 | 0.095 | |
| 1344 | 0.484 | 0.007 | |
| 1345 | 0.436 | 0.084 | |
| 1346 | 0.439 | 0.093 | |
| 1347 | 0.418 | 0.121 |
| Change | 50M | 80M | 520M first experiment | 520M second experiment |
|---|---|---|---|---|
| 0.104 | 0.524 | 0.509 | ||
| – |
| Target loss | Generic loss | ||||||
|---|---|---|---|---|---|---|---|
| Matched | Additional | Final | Matched | Additional | Final | ||
| 35M | |||||||
| 35M | |||||||
| 35M | |||||||
| 50M | |||||||
| 50M | |||||||
| Configuration | Target fraction | Repetitions | Target loss | Generic loss |
|---|---|---|---|---|
| High target fraction | 0.978 | 4.262 | ||
| Middle target fraction | 0.748 | 3.904 | ||
| Low target fraction | 0.748 | 3.835 |
| PubMed | Caselaw | |
|---|---|---|
| 4 | ||
| 8 | ||
| 12 | ||
| 16 | ||
| 24 | ||
| 32 |
| Comparison | PubMed | Caselaw |
| Reuse the 80M winner | 0.854 | 2.492 |
| Keep the three lowest-loss 80M counts | 0.213 | 0.658 |
| Use the fixed set | 0 | 0 |
| Exchange the selected sets between corpora | 0 | 0 |
| PubMed mean SD | Caselaw mean SD | Stage | |
|---|---|---|---|
| 4 | Original | ||
| 6 | Planned | ||
| 8 | Original | ||
| 9 | (SD ) | Exploratory | |
| 10 | (SD ) | Planned | |
| 11 | Exploratory |
| Corpus | Frozen fit | Local RMSE | ||||
|---|---|---|---|---|---|---|
| PubMed | Primary, | 9.659 | 8 | 0.038 | 0.044 | 0.021 |
| Free | 7.730 | 8 | 0.038 | 0.044 | 0.075 | |
| Saturated | 14.827 | 16 | 0.100 | 0.106 | 0.048 | |
| 35M/50M, | 10.534 | 12 | 0 | 0.006 | 0.015 | |
| Caselaw | Primary, | 11.431 | 12 | 0 | 0 | 0.019 |
| Free | 10.112 | 12 | 0 | 0 | 0.044 |
| Target (M) | Sources (M) | PubMed | Caselaw |
|---|---|---|---|
| 520 | 35, 50, 80 | ||
| 520 | 35, 50 | ||
| 520 | 35, 80 | ||
| 520 | 50, 80 | ||
| 80 | 35, 50 |
| Corpus | Target (M) | Pruned set | Fixed set | Fixed minus pruned |
|---|---|---|---|---|
| PubMed | 80 | 2.311 | 2.354 | 0.043 |
| Caselaw | 80 | 2.541 | 2.603 | 0.062 |
| PubMed | 520 | 2.339 | 2.339 | 0 |
| Caselaw | 520 | 2.649 | 2.649 | 0 |
| Corpus | Sizes | ||||
|---|---|---|---|---|---|
| Wikipedia | 35M 80M | 0.203 | 0.064 | 0.139 | |
| Proof-Pile-2 | 50M 80M | 0.110 | 0.007 | 0.104 |
| Target corpus | Nominal | Realized | 80M continuous minimum | Evaluated choice |
|---|---|---|---|---|
| Wikipedia-derived | ||||
| Proof-Pile-2 |
| Target corpus | Ordered bootstrap resamples | Available fits | Unavailable fits | Same 80M choice |
|---|---|---|---|---|
| Wikipedia-derived | 6912 | 6885 | 27 | 6885/6885 select |
| Proof-Pile-2 | 6912 | 6912 | 0 | 6912/6912 select |
| Fit RMSE | Held-out RMSE | |||
|---|---|---|---|---|
| Corpus | Cap 4 | Cap 8 | Cap 4 | Cap 8 |
| Wikipedia-derived | 0.020 | 0.020 | — | — |
| Proof-Pile-2 | 0.015 | 0.015 | — | — |
| PubMed | 0.089 | 0.088 | 0.107 | 0.113 |
| Caselaw | 0.076 | 0.071 | 0.093 | 0.095 |
| Corpus | |||||
|---|---|---|---|---|---|
| Proof-Pile-2 | 6.026 | 1.938 | 2.383 | 0.275 | 0.233 |
| PubMed | 9.490 | 2.279 | 3.734 | 0.602 | 0.175 |
| Caselaw | 9.984 | 1.536 | 4.000 | 0.779 | 0.207 |
| Corpus | Fitted | (cap 4) | Retained-set regret | Unit-grid regret |
|---|---|---|---|---|
| PubMed | 9.46–9.49 | 5.124 | 0.066–0.069 | 0.006–0.007 |
| Caselaw | 9.98–10.32 | 4.784 | 0.060–0.094 |
| Corpus | Lovelace shape | Only freed | Saturating extension |
|---|---|---|---|
| Wikipedia-derived | 0.066 | 0.034 | 0.020 |
| Proof-Pile-2 | 0.087 | 0.068 | 0.015 |
| PubMed | 0.258 | 0.167 | 0.089 |
| Caselaw | 0.263 | 0.202 | 0.076 |
| Corpus | Lovelace shape | Only freed | Saturating extension |
|---|---|---|---|
| PubMed | 0.336 | 0.236 | 0.107 |
| Caselaw | 0.371 | 0.340 | 0.093 |
| Free | |||||||
|---|---|---|---|---|---|---|---|
| Corpus | RMSE | RMSE | |||||
| Wikipedia-derived | 0.996 | 0.071 | 0.099 | 1.579 | 0.046 | 0.099 | |
| Proof-Pile-2 | 1.280 | 0.034 | 0.163 | 1.360 | 0.034 | 0.060 | |
| PubMed | 1.309 | 0.015 | 0.093 | 1.292 | 0.015 | 0.077 | |
| Caselaw | 1.175 | 0.021 | 0.088 | 1.170 | 0.021 | 0.067 | |
| Corpus | Fitted sizes (M) | Free | |
|---|---|---|---|
| PubMed | 35, 50 | [1.103, 1.384] | [0.894, 1.082] |
| 35, 50, 80 | [1.285, 1.334] | [1.263, 1.320] | |
| All | [0.728, 0.772] | [0.647, 0.718] | |
| Caselaw | 35, 50 | [0.860, 0.998] | [0.688, 0.941] |
| 35, 50, 80 | [1.132, 1.213] | [1.123, 1.212] | |
| All | [0.845, 0.887] | [0.819, 0.887] |
| Corpus | Fitted sizes (M) | Best | RMSE | Best RMSE |
|---|---|---|---|---|
| Wikipedia-derived | 35, 50 | 0.659 | 0.016 | 0.016 |
| 35, 50, 80 | 0.826 | 0.020 | 0.041 | |
| Proof-Pile-2 | 35, 50, 80 | 0.012 | 0.055 | |
| All | 3.523 | 0.015 | 0.065 | |
| PubMed | 35, 50, 80 | 0.003 | 0.074 | |
| All | 0.089 | 0.215 |
| Corpus | Source target (M) | Predicted | Measured best | Excess loss |
|---|---|---|---|---|
| Wikipedia-derived | 24 | 24 | 0 | |
| Proof-Pile-2 | 8 | 8 | 0 | |
| PubMed | 16 | 16 | 0 | |
| Caselaw | 24 | 24 | 0 | |
| Proof-Pile-2 | 4 | 8 | 0.074 | |
| PubMed | 4 | 8 | 1.435 |
| PubMed | Caselaw | |||
|---|---|---|---|---|
| Predicted | Measured | Predicted | Measured | |
| 8 | ||||
| 10 | ||||
| 11 | ||||
| 16 | ||||
| 48 | ||||
| Corpus | Predicted interval | Measured counts inside | Measured low-loss counts |
|---|---|---|---|
| PubMed | |||
| Caselaw |
| Corpus | Earlier fit | Choice | Excess loss | Full RMSE | Local RMSE | |
|---|---|---|---|---|---|---|
| PubMed | Bound 4 | 11.214 | 11 | 0 | 0.375 | 0.048 |
| Bound 8 | 9.525 | 10 | 0.003 | 0.598 | 0.171 | |
| Caselaw | Bound 4 | 13.181 | 14 | 0.018 | 0.317 | 0.070 |
| Bound 8 | 11.639 | 12 | 0 | 0.598 | 0.068 |
| Check | PubMed | Caselaw |
| Joint fit, all 40 means: training RMSE | 0.115 | 0.089 |
| Joint fit, all 12 target means: 200M training RMSE | 0.129 | 0.083 |
| Joint fit, seven target counts: absolute prediction RMSE on five added counts | 0.108 | 0.074 |
| Leave one count out at every size: RMSE over 30 predictions | 0.127 | 0.118 |
| Same leave-count-out fits: RMSE on ten 200M interior counts | 0.111 | 0.095 |
| Same leave-count-out fits: RMSE on five added 200M counts | 0.089 | 0.077 |
| PubMed | Caselaw | |||||
|---|---|---|---|---|---|---|
| Lovelace | Saturated | Lovelace | Saturated | |||
| 6 | 0.219 | 0.203 | 0.069 | 0.162 | 0.208 | 0.108 |
| 9 | 0.226 | 0.103 | 0.206 | 0.148 | ||
| 10 | 0.235 | 0.115 | 0.222 | 0.157 | ||
| 11 | 0.242 | 0.134 | 0.236 | 0.169 | ||
| 14 | 0.267 | 0.207 | 0.279 | 0.225 | ||
| Work | Selected decision | From inexpensive evidence to intended use | Run construction and repetition |
|---|---|---|---|
| Kaplan et al. (2020) ; Hoffmann et al. (2022) | Model size and token allocation | Model size, data, and compute | Processed tokens are treated as fresh |
| Repeated-data scaling ( Hernandez et al., 2022 ; Xue et al., 2023 ; Muennighoff et al., 2025 ; Xu et al., 2026 ) | Useful epochs or finite-data allocation | Unique data, repeats, and model size | Repetition is explicit in an effective-data law |
| Liu et al. (2025) | Data-mixture weights | Small proxy models to a larger run | Proxy regression over independently trained mixtures |
| Sedova et al. (2026) | Mixture and repetition-aware minimum | Model size, token count, target size, and mixture | Candidate losses come from intermediate states of one run |
| Zhou et al. (2026) | Mixture at a target token horizon | Token horizon and model size | Repetition is matched across training runs |
| DataDecide; AutoScale ( Magnusson et al., 2025 ; Kang et al., 2025 ) | Corpus ranking or domain weights | Model size or data scale | Completed corpus runs or domain-weight extrapolation |
| Question | Observation | Supported conclusion |
|---|---|---|
| Does a better 35M curve fit ensure transfer? | Conditions with small leave-one-out fit error still lose 0.032–0.035 nat at 50M. | Fit quality at one scale does not determine the transferred count (Section G.4 ). |
| Does gradient alignment predict the loss increase? | The fitted change is 0.001 nat per standard deviation, with an approximate 95% interval of . | No large association is observed in this twelve-condition design; an exactly zero effect is not established (Section G.3 ). |
| Do simple text-overlap and duplication statistics predict the loss increase? | Lexical overlap is driven by two high-leverage conditions; longer -gram overlap is nearly constant; 8-gram redundancy is correlated with the corpus construction. | None supplies a reliable prediction rule here (Section G.3 ). |
| Does changing the generic corpus remove the failure? | All twelve conditions select 48 at 35M and 32 at 50M, with 0.032–0.042 nat loss increases from reuse. | The tested generic-corpus substitutions do not remove the transfer failure (Section G.3 ). |
| Do additional natural corpora extend the result? | Their gradient-alignment values cover the constructed range, but they have no loss curves across repetition counts. | They contextualize the measurement only (Section G.6 ). |
| Does higher repetition degrade performance on both corpus categories? | Held-out generic loss continues to improve in the matched-exposure comparisons. | The target-loss increase is not accompanied by an increase on the measured generic data (Section B.4 ). |
| Condition | Gradient alignment | Lexical overlap | Longer -gram overlap | 8-gram redundancy | Seeds | |
|---|---|---|---|---|---|---|
| Semantic 0.00 | 0.00 | 0.209 | 0.228 | 0.005 | 0.008 | 4 |
| Semantic 0.25 | 0.25 | 0.217 | 0.178 | 0.005 | 0.008 | 2 |
| Semantic 0.50 | 0.50 | 0.222 | 0.179 | 0.005 | 0.008 | 2 |
| Semantic 0.75 | 0.75 | 0.235 | 0.180 | 0.005 | 0.008 | 2 |
| Semantic 1.00 | 1.00 | 0.239 | 0.178 | 0.005 | 0.008 | 4 |
| Entity 0.00 | 0.00 | 0.393 | 0.189 | 0.005 | 0.007 | 4 |
| Condition | Seeds | 35M selection | 50M selection | Loss increase from reusing 48 |
|---|---|---|---|---|
| Semantic 0.00 | 4 | 48 | 32 | 0.042 |
| Semantic 0.25 | 2 | 48 | 32 | 0.034 |
| Semantic 0.50 | 2 | 48 | 32 | 0.033 |
| Semantic 0.75 | 2 | 48 | 32 | 0.032 |
| Semantic 1.00 | 4 | 48 | 32 | 0.037 |
| Entity 0.00 | 4 | 48 | 32 | 0.036 |
| Term | Coefficient | Standard error | statistic |
|---|---|---|---|
| Intercept | -0.135 | 0.135 | -1.00 |
| Construction strength | 0.001 | 0.002 | 0.75 |
| Gradient alignment | 0.019 | 0.016 | 1.18 |
| Lexical overlap | 0.233 | 0.088 | 2.64 |
| Longer -gram overlap | 24.54 | 23.31 | 1.05 |
| 8-gram redundancy | -0.078 | 0.037 | -2.11 |
| Rule | Count at 50M | Loss increase | Count at 80M | Loss increase |
|---|---|---|---|---|
| Fixed | 4 | 1.208 | 4 | 0.914 |
| Fixed | 8 | 0.419 | 8 | 0.251 |
| Fixed | 16 | 0.104 | 16 | 0.008 |
| Fixed | 24 | 0.020 | 24 | 0 |
| Fixed | 32 | 0 | 36 † | 0.139 |
| Reuse 35M selection | 48 | 0.033 | 36 | 0.139 |
| minibatches | 1 | 8 | 32 | 64 | 128 |
|---|---|---|---|---|---|
| Target-to-target alignment | 0.059 | 0.320 | 0.619 | 0.770 | 0.868 |
| Unpreconditioned cosine | 0.41 | 0.85 | 0.96 | 0.98 | 0.99 |
| FineWeb-to-target alignment | 0.025 | 0.124 | 0.229 | 0.278 | 0.298 |
| Corpus | Mean alignment | Seed SD | Relation to the constructed ranges |
|---|---|---|---|
| Project Gutenberg | 0.105 | 0.005 | Below semantic range |
| OpenAssistant-1 | 0.184 | 0.003 | Below semantic range |
| FineWeb generic | 0.222 | 0.007 | Inside semantic range |
| Unmodified base corpus | 0.232 | 0.004 | Near semantic range |
| Harvard USPTO Patent Dataset (HUPD) claims | 0.271 | 0.013 | Between constructed ranges |
| Raw Wikipedia | 0.307 | 0.004 | Between constructed ranges |