What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining
Organizations: FloatAI · ETH Zurich · Independent Researcher
Abstract
As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters. We answer them by pricing a repeated token against two references: one epoch on the same data, which gives its value, and fresh data at equal compute, which gives its cost. Against fresh data, the cost of repetition follows a single variable, the number of extra epochs divided by the unique tokens per parameter. Against the same data, a second epoch is worth nearly as much as a fresh one, and repeated tokens fall to half the value of fresh ones after a critical epoch count that grows with the training budget per parameter but hardly with model size. With unique data fixed, the predicted compute-optimal run grows model size and epochs together until loss stops improving, near the critical epoch count. The same variable accounts for the direction of size trends that appear to conflict: larger models tolerate fewer epochs when the corpus is fixed, from about 15 at 127M to 4 at 2B parameters, but not when unique data grow with the model. Counts alone do not determine loss: at identical counts, replaying shards consecutively raises loss by up to 0.46~bits per byte, concentrating repeats on fewer samples also raises it, lower-entropy sources degrade faster with repetition, and re-tokenizing repeats helps only under heavy repetition. These results offer an empirical guide to pretraining when unique data, rather than compute, are the binding constraint.
Figures & tables
| Design | Baseline | with size | Predicted | Observed |
|---|---|---|---|---|
| Fixed (§ 5.1 ) | data-matched | minimum at fewer epochs | 15.3 to 3.6 epochs, 127M to 2B | |
| Fixed (§ 5.2 ) | data-matched | minimum nearly fixed | 3.7–4.4 epochs on CommonCrawl | |
| Fixed , , (§ 5.3 ) | compute-matched | 22–39% less at 1.2B | 33% and 21% less | |
| Fixed , T5 ( Xue et al., 2023 ) | compute-matched | larger models degrade more | larger models degrade more | |
| , ( Muennighoff et al., 2023 ) | compute-matched | bpb | nearly as good as fresh data | |
| Small subset, many repeats ( Hernandez et al., 2022 ) | compute-matched | large | large excess loss | marked degradation |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Experiment | Models | Controlled | Varied | Sources |
|---|---|---|---|---|
| Single-run sweeps | ||||
| grid (§ 4 ) | 23M–361M | , | , ( ) | CC |
| Fixed (§ 5.1 ) | 14M–2B | ; (14M–59M) | CC | |
| (§ 5.2 ) | 23M–1.2B | , | CC, GitHub | |
| Order (§ 6.2 ) | 59M–511M | , , | order | CC |
| Source entropy (§ 6.2 ) | 59M | , | source, | 7 sources |
| Field | 45M model | 185M model |
|---|---|---|
| Exact non-embedding parameters | 45,084,900 | 184,623,360 |
| Tied embedding parameters | 32,164,480 | 51,463,168 |
| Exact total parameters | 77,249,380 | 236,086,528 |
| Rounded total size | 77.2M | 236.1M |
| Width / depth | 640 / 10 | 1024 / 16 |
| Width/depth ratio | 64 | 64 |
| Study/model | LR tuple | Fixed budget / updates | ||
|---|---|---|---|---|
| Source / 45M | 0.01/0.01/0.05 | 131,072 | 1 | ; |
| / 45M | 0.02/0.02/0.05 | 131,072 | 1 | ; |
| / 185M | 0.02/0.02/0.05 | 524,288 | 4 | ; |
| / 45M | 0.02/0.02/0.05 | 131,072 | 1 | ; |
| / 185M | 0.02/0.02/0.05 | 524,288 | 4 | ; |
| heads (KV) | (M) | batch (tok) | ||||
| Main CommonCrawl model-size sweep | ||||||
| 384 | 8 | 48.00 | 6 (3) | 14M | 12.990 | 262,144 |
| 512 | 10 | 51.20 | 8 (4) | 31M | 28.859 | 524,288 |
| 640 | 12 | 53.33 | 10 (5) | 59M | 54.102 | 524,288 |
| 768 | 18 | 42.67 | 12 (6) | 127M | 116.848 | 524,288 |
| 1024 | 20 | 51.20 | 16 (8) | 252M | 230.779 | 524,288 |
| 1 | 1.112899 | 1.152695 | 0.000 | 0.000 |
|---|---|---|---|---|
| 2 | 1.114845 | 1.153794 | 1.945 | 1.100 |
| 4 | 1.124383 | 1.157711 | 11.484 | 5.016 |
| 8 | 1.160769 | 1.189120 | 47.869 | 36.426 |
| 16 | 1.279985 | 1.283226 | 167.086 | 130.531 |
| 32 | 1.867810 | 1.760766 | 754.911 | 608.072 |
| Contrast | Stratum | Mean (bpb) | 95% CI |
|---|---|---|---|
| Concentration | Source mean | ||
| CommonCrawl | |||
| GitHub | |||
| Books | |||
| Two sizes, | Source mean | ||
| Two sizes, | Source mean |
| Full corpus | Partial | Partial | |||||
|---|---|---|---|---|---|---|---|
| Size | CC | GH | CC | GH | CC | GH | |
| 23M | 23,087,168 | 3.72 | 3.06 | – | – | – | – |
| 45M | 45,084,900 | 3.79 | 2.84 | 0.0322 0.0021 | 0.0337 0.0060 | 0.0038 0.0033 | 0.0043 0.0013 |
| 78M | 77,898,384 | 4.03 | 2.75 | – | – | – | – |
| 185M | 184,623,360 | 4.22 | 2.68 | 0.0249 0.0012 | 0.0296 0.0007 | 0.0024 0.0023 | 0.0041 0.0010 |
| 361M | 360,563,600 | 4.30 | 2.66 | 0.0245 0.0012 | 0.0270 0.0023 | 0.0027 0.0014 | 0.0029 0.0004 |
| B | ||||||
|---|---|---|---|---|---|---|
| 14M | 1.2029 | |||||
| 31M | 1.1642 | |||||
| 59M | 1.1343 | |||||
| 127M | 1.1060 | |||||
| 252M | 1.0837 | |||||
| , | , | |||
|---|---|---|---|---|
| 59M | 1.1343 | 1.0605 | 0.95 [0.76, 1.05] | 0.81 [0.75, 0.83] |
| 252M | 1.0837 | 0.9955 | 0.93 [0.75, 1.03] | 0.72 [0.64, 0.75] |
| 511M | 1.0951 | 0.9788 | 1.01 [0.82, 1.12] | 0.72 [0.65, 0.75] |
| 1B | 1.0696 | 0.9553 | 0.94 [0.76, 1.05] | 0.59 [0.51, 0.63] |
| Form | Penalty | Parameters | MAE (bpb) | |
|---|---|---|---|---|
| Benefit-only | none | 6 | 0.360 | 0.0884 |
| Separable | 9 | 0.985 | 0.1142 | |
| Load-augmented | 8 | 0.982 | 0.0311 |
| Primary contrast | Mean bpb | 95% CI |
|---|---|---|
| 20%-off minus On | ||
| 50%-off minus 20%-off |