Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
Authors: Atsuki Yamaguchi, Tatsuro Inaba, Joel Niklaus, Michal Štefánik, Aline Villavicencio, Nikolaos Aletras
Organizations: University of Sheffield, UK · Mohamed bin Zayed University of Artificial Intelligence, UAE · Hugging Face · National Institute of Informatics, Japan · University of Exeter, UK · Federal University of Rio Grande do Norte, Brazil
Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.
Figures & tables
Figure 1: (a) PPT is claimed to induce a grammatical prior that transfers to natural language. (b) Prior work sits at or below 1B parameters and short training horizons. We span four scales, four PT mixtures, five PPT tasks, and up to 100B PT tokens.
Category (%)
C4
SmolLM3
OLMo3
Marin
Natural Lang.
100.0
85.0
89.5
92.6
└ Web
100.0
79.8
76.9
92.6
Code
0.0
12.0
7.1
6.1
Math
0.0
3.0
3.4
1.3
Table 1: Comparison of PT data mixtures, as reported by their creators. Web is a subset of natural language, with the remainder drawn from non-web sources such as books, academic text, and encyclopedic text. Our PT runs follow these ratios.
Figure 2: Downstream score changes from PPT relative to PT-Only. Bars show mean paired differences across 5K–10K checkpoints; whiskers denote ± 1 SD. Full scores are in Appendix Table 7 .
PT-Only
Δ NLL vs. PT-Only
Data Mix
NLL
Control
k -Shuffle Dyck
500M
C4
3.358 .041
+0.060 .044
-0.064 .047
SmolLM3
3.391 .058
+0.030 .033
-0.020 .036
OLMo3
3.650 .063
-0.046 .053
-0.066 .072
Marin
3.341 .065
-0.006 .016
-0.084 .033
1B
C4
3.307 .052
-0.036 .031
-0.058 .027
Table 2: Verbatim retrieval as mean NLL ( ↓ ) averaged over checkpoints at 5K–10K steps ( Δ = PPT − PT-Only). Subscripts are standard deviations (SD) across checkpoints. Green marks an improvement over the baseline.
Figure 3: BLiMP score changes from PPT relative to PT-Only. Bars show mean paired differences across 5K–10K checkpoints; whiskers denote ± 1 SD. Full scores are in Appendix Table 8 .
Downstream (%)
BLiMP
Verbatim
RC
Sci. QA
CR
LM
Avg
(%)
NLL ↓
n/4
(a)
Marin (full)
+0.6 0.6
+2.3 0.6
+1.8 0.8
+3.1 1.3
+1.9 0.5
-0.3 1.0
-0.073 0.024
–
17% Math
+1.0 0.4
+2.2 1.1
+1.8 0.4
+3.1 0.8
+2.0 0.3
-0.3 1.6
-0.024 0.014
–
DCLM only
+0.8 0.2
+1.6 0.5
+1.6 0.9
+2.1 0.7
+1.5 0.3
+0.0 2.0
-0.054 0.043
–
FineWeb-Edu only
+0.5 0.6
+0.0 1.3
+1.9 1.2
+1.5 1.4
+1.0 0.4
-1.5 3.9
-0.042 0.050
–
Marin\DCLM
+0.4 0.2
+0.6 0.9
-0.9 1.3
+0.6 0.9
+0.2 0.5
+0.4 1.4
-0.012 0.041
–
Table 3: Ablations at 3B: (a) Marin composition and (b) PPT task. Scores are mean paired differences from PT-Only over checkpoints at 5K–10K steps. In (a), subscripts are standard deviations across checkpoints. In (b), scores are averaged over the four mixtures, except for PT-Only, which gives absolute scores, and n/4 counts the mixtures in which the downstream average improves (per-mixture breakdown in Appendix Table 14 ). Green marks an improvement.
C4 (3B)
Marin (3B)
Marin (7B)
Tokens
Avg
BLiMP
NLL ↓
Avg
BLiMP
NLL ↓
Avg
BLiMP
NLL ↓
21B
+1.2
+3.0
-0.039
+1.1
-3.1
-0.015
+0.5
-0.2
+0.008
42B
+0.7
+3.1
+0.014
+1.6
-10.3
-0.014
+0.3
+1.8
+0.002
63B
+0.7
-2.2
-0.033
+1.5
-1.2
-0.019
+0.4
-0.6
-0.004
75.5B
–
–
–
–
–
–
+1.0
-1.1
-0.029
84B
+0.8
-0.7
-0.021
+1.3
+1.3
-0.035
–
–
–
Table 4: Scaling. Each entry is the difference between PPT ( k -Shuffle Dyck) and the PT-Only baseline at the same PT budget. Score breakdowns are in Appendix Tables 15 and 16 . Results at 21B tokens are not comparable to § 5 , as the two use different learning rate schedules.
Task
Symbol inventory
Ints / doc
k -Shuffle Dyck
k=64 , popen=0.50 , depth ∈[1,8]
2,048
MP-Struct Core
kstruct=1 , kdep=4 , 3 heads, 1 leaf
2,048
Set
999 values, separator 999
2,048
NCA
12×12 grid, 10 states, 2×2 patches, T=10−4
1,406
Table 5: Generation parameters for the pre-pretraining corpora.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
500M
1B
3B
7B
Architecture
Hidden size ( dmodel )
1,024
1,536
2,560
3,584
Intermediate size ( dff )
2,816
4,096
6,912
18,944
Number of layers ( nlayer )
32
36
40
28
Attention heads ( nhead )
16
12
20
28
Key-Value heads ( nkv )
4
4
4
4
Appendix
Table 6: Architectural and optimization hyperparameters across model scales.
RC
Science QA
CR
LM
Approach
RACE
ReCoRD
SciQ
ARC-Easy
OpenBookQA
COPA
PIQA
SIQA
HellaSwag
LAMBADA
500M
C4 (PT-Only)
29.3
66.8
61.5
35.9
27.4
71.2
67.8
35.1
37.7
30.8
Control
-0.5 0.6
-0.8 0.5
-0.4 1.6
+1.4 1.2
+0.2 1.1
-4.5 2.8
-1.0 0.6
+0.6 0.7
-0.6 0.6
-0.4 1.2
k -Shuffle Dyck
-0.0 1.0
+0.9 0.7
+1.1 1.3
+0.6 1.1
+0.7 0.9
-2.3 2.1
-1.0 1.1
+1.0 0.6
+1.7 0.7
+0.1 1.1
SmolLM3 (PT-Only)
29.6
65.1
67.8
53.5
30.0
66.0
64.9
38.3
36.3
35.7
Control
-0.2 0.8
-0.3 0.7
-1.6 1.5
+0.0 0.9
-2.2 1.5
+2.8 2.4
-0.5 1.3
+0.6 0.5
-0.8 0.4
+0.1 1.2
Appendix
Table 7: Task-level downstream performance. Gray rows give the PT-Only mean over checkpoints at 5K–10K steps; the rows below give the mean paired difference from that baseline, with standard deviations across checkpoints as subscripts. Green marks an improvement.
Semantics
Morphology
Syntax
Approach
Quant
NPI
Ana Agr
Irregul
DN Agr
SV Agr
Arg Str
Bind
Ctrl Rais
Ellips
Fill Gap
Island
Overall
500M
C4 (PT-Only)
71.0
64.5
97.5
93.9
94.8
85.8
80.3
77.2
81.4
87.2
78.7
68.8
79.6
Control
-1.3 4.2
-0.6 9.3
-1.3 3.1
+0.7 2.0
-1.4 1.3
-2.0 1.8
-1.6 2.0
+2.2 2.2
-2.2 2.4
-0.2 1.6
-1.7 1.8
-6.4 2.2
-1.6 1.5
k -Shuffle Dyck
+3.7 5.9
-2.1 3.9
-2.4 2.9
-1.0 1.6
-0.4 0.4
-2.0 3.2
-0.5 0.8
+2.9 1.2
-0.8 1.0
+0.2 1.8
+0.6 1.0
-3.9 3.4
-0.6 1.0
SmolLM3 (PT-Only)
73.8
57.0
96.0
94.8
95.6
86.7
80.9
79.7
82.4
88.3
80.7
67.7
79.7
Control
-3.2 4.6
+12.1 7.7
+0.8 3.3
+0.4 2.0
-0.1 1.1
+1.3 2.5
-0.1 1.3
+1.7 1.2
-2.7 1.0
-0.1 1.1
+1.3 1.0
-0.0 3.0
+1.3 1.7
Appendix
Table 8: Linguistic competence on BLiMP by phenomenon. Gray rows give the PT-Only mean over checkpoints at 5K–10K steps; the rows below give the mean paired difference from that baseline, with standard deviations across checkpoints as subscripts. Green marks an improvement.
Scale
Data
Δ10K
Δˉ
sd
Ckpts
Downstream Average
500M
C4
+0.66
+0.28
0.23
6/6
500M
Marin
+1.89
+1.39
0.57
6/6
500M
OLMo3
+1.38
+0.66
0.62
4/6
500M
SmolLM3
+2.64
+0.98
0.94
5/6
1B
C4
+2.94
+2.23
0.73
6/6
Appendix
Table 9: Single-checkpoint versus multi-checkpoint estimates of the k -Shuffle Dyck effect. Δ10K denotes the paired difference at step 10K, while Δˉ and sd represent the mean and standard deviation across checkpoints from 5K to 10K steps. The column Ckpts reports the ratio of checkpoints favoring PPT, and † marks configurations with directional disagreement between the two estimates.
Downstream average
BLiMP
Verbatim
Window
Cells >0
Flips
500M
1B
3B
Flips
Cells >0
2K–10K
11/12
0/12
+0.86
+1.13
+1.23
5/12
12/12
3K–10K
11/12
0/12
+0.84
+1.29
+1.22
5/12
12/12
4K–10K
11/12
0/12
+0.84
+1.36
+1.22
5/12
12/12
5K–10K
11/12
0/12
+0.83
+1.37
+1.27
6/12
12/12
6K–10K
11/12
0/12
+0.80
+1.43
+1.29
6/12
12/12
Appendix
Table 10: Sensitivity of downstream gains and linguistic benchmarks to the choice of averaging window. The metric Cells >0 records the proportion of configurations where k -Shuffle Dyck yields improvements, whereas Flips denotes directional disagreement relative to the 10K snapshot across all 12 scale and mixture combinations. Bold text denotes the baseline 5K–10K window adopted throughout this study.
RC
Science QA
CR
LM
Approach
RACE
ReCoRD
SciQ
ARC-Easy
OpenBookQA
COPA
PIQA
SIQA
HellaSwag
LAMBADA
Seed 1
Marin 3B (PT-Only)
33.7
75.1
75.9
60.5
32.0
73.3
69.8
42.9
47.4
48.5
k -Shuffle Dyck
-0.6 0.7
+1.7 0.5
+2.4 1.3
+2.6 0.7
+1.9 1.4
+0.8 2.9
+2.3 0.6
+0.4 0.6
+3.6 0.2
+3.1 1.3
Seed 2
Marin 3B (PT-Only)
33.0
74.6
74.9
60.5
33.4
71.7
69.7
42.4
47.0
48.1
k -Shuffle Dyck
-0.4 0.9
+1.7 1.0
+3.6 1.7
+2.1 0.4
+0.2 0.9
+1.2 3.7
+1.4 0.9
+1.1 0.8
+3.4 0.3
+3.3 1.3
Seed 3
Marin 3B (PT-Only)
32.4
75.2
75.4
61.3
34.0
72.8
70.0
41.4
47.7
49.4
Appendix
Table 11: Task-level downstream performance across random seeds (3B Marin). Gray rows give the PT-Only mean over checkpoints at 5K–10K steps; the rows below give the mean paired difference from that baseline, with standard deviations across checkpoints as subscripts. Green marks an improvement.
Semantics
Morphology
Syntax
Approach
Quant
NPI
Ana Agr
Irregul
DN Agr
SV Agr
Arg Str
Bind
Ctrl Rais
Ellips
Fill Gap
Island
Overall
Seed 1
Marin 3B (PT-Only)
77.8
58.5
93.1
88.9
96.0
86.2
80.7
82.2
79.4
89.5
77.9
68.2
79.7
k -Shuffle Dyck
-8.7 5.6
+6.5 6.8
+2.9 3.2
+1.0 3.8
-0.6 0.9
-0.1 2.8
-2.0 1.3
-1.9 1.1
+1.1 1.6
-2.3 0.4
-0.5 1.0
+0.3 3.2
-0.3 1.0
Seed 2
Marin 3B (PT-Only)
72.4
66.1
90.9
90.0
95.3
87.2
80.0
82.3
80.4
89.6
78.7
68.3
80.2
k -Shuffle Dyck
+3.9 9.3
-3.0 8.2
+7.9 6.8
+4.0 6.4
-0.2 0.4
-1.5 2.4
-0.7 1.3
-0.3 0.9
+0.5 1.0
-1.9 1.0
-0.4 1.3
+1.5 3.3
+0.1 1.9
Seed 3
Marin 3B (PT-Only)
79.3
66.5
94.1
87.7
96.2
85.5
81.8
82.2
80.2
88.9
79.3
71.6
81.3
Appendix
Table 12: Linguistic competence on BLiMP across random seeds (3B Marin). Gray rows give the PT-Only mean over checkpoints at 5K–10K steps; the rows below give the mean paired difference from that baseline, with standard deviations across checkpoints as subscripts. Green marks an improvement.
Approach
Mean NLL ↓
Seed 1
Marin 3B (PT-Only)
3.150
k -Shuffle Dyck
-0.073 0.024
Seed 2
Marin 3B (PT-Only)
3.101
k -Shuffle Dyck
+0.014 0.022
Seed 3
Marin 3B (PT-Only)
3.113
k -Shuffle Dyck
-0.040 0.044
Appendix
Table 13: Verbatim retrieval performance across random seeds (3B Marin). Gray rows give the PT-Only baseline mean over checkpoints at 5K–10K steps; the rows below give the mean paired difference ( Δ ) from that baseline, with checkpoint standard deviations as subscripts. Lower is better ( Δ<0 is green ).
Downstream (%)
BLiMP
Verbatim
Approach
RC
Sci. QA
CR
LM
Avg
(%)
NLL ↓
C4
PT-Only (abs.)
53.1
45.9
57.2
36.7
48.2
78.3
3.227
k -Shuffle Dyck
+1.6 0.5
+0.4 0.6
+0.5 0.7
+2.8 1.1
+1.3 0.6
+1.7 1.8
-0.093 0.028
MP-Struct Core
+1.1 0.7
+1.3 1.4
+1.1 1.0
+3.2 0.7
+1.7 0.6
+2.3 1.9
-0.077 0.026
NCA
+0.5 0.3
+1.9 0.7
+0.8 1.1
+1.4 0.9
+1.2 0.4
+2.1 1.8
-0.101 0.033
Set
-8.3 4.6
-6.4 3.8
-7.2 4.4
-14.0 11.4
-9.0 6.0
-4.4 10.8
+0.527 0.469
Appendix
Table 14: Pre-pretraining task ablation at 3B, per data mixture. Entries are mean paired differences from PT-Only over checkpoints at 5K–10K steps, with standard deviations across checkpoints as subscripts. Higher is better except for verbatim NLL.
Setting
Tokens
Approach
RC
Science QA
CR
LM
Avg (%)
BLiMP (%)
Verbatim NLL ↓
C4 (3B)
21B
PT-Only
55.0
46.0
59.0
39.0
49.7
77.8
3.165
PPT ( k -Shuffle Dyck)
55.5
47.8
59.4
41.2
51.0
80.8
3.127
42B
PT-Only
57.0
51.0
60.5
44.2
53.2
77.9
3.062
PPT ( k -Shuffle Dyck)
58.5
49.2
61.8
46.0
53.9
81.0
3.077
63B
PT-Only
58.3
53.0
61.8
44.6
54.4
78.2
3.057
PPT ( k -Shuffle Dyck)
58.8
52.9
63.0
46.1
55.2
75.9
3.025
Appendix
Table 15: Absolute scores for the extended-budget runs. Baselines ( PT-Only ) are shaded gray; green and bold mark improvements from PPT ( k -Shuffle Dyck).
RC
Science QA
CR
LM
Setting
Tokens
Approach
RACE
ReCoRD
SciQ
ARC-Easy
OpenBookQA
COPA
PIQA
SIQA
HellaSwag
LAMBADA
C4 (3B)
21B
PT-Only
33.2
76.8
67.8
41.3
28.8
76.0
71.7
36.9
51.2
39.0
PPT
33.2
77.8
69.5
39.8
34.0
75.0
73.2
35.9
53.7
41.2
42B
PT-Only
34.3
79.7
72.9
49.0
31.0
75.0
72.6
37.7
56.6
44.2
PPT
36.0
81.0
74.3
40.1
33.2
76.0
75.0
37.4
58.8
46.0
63B
PT-Only
35.9
80.8
77.2
50.8
31.0
73.0
74.9
40.7
58.6
44.6
Appendix
Table 16: Task-level breakdown for the extended-budget runs. Baselines ( PT-Only ) are shaded gray; green and bold mark improvements from PPT ( k -Shuffle Dyck).
Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and ultimately degrade model performance. Data curation mitigates but cannot eliminate such noise, so pre-training corpora remain noisy in practice. We therefore study whether a lightweight pre-pre-training (PPT) stage based on synthetic data with learnable temporal structure helps resist noisy data during the pre-training (PT) stage. Across various corruption settings, our method consistently improves robustness to noise during PT, with larger relative gains at higher noise levels. For a 1B-parameter model, a synthetic PPT stage with only 65M tokens achieves the same final loss as the baseline while using up to 49% fewer natural-text PT tokens across different noise levels. Mechanistic analyses suggest PPT does not immediately suppress attention to noisy tokens. Rather, PPT-initialized models gradually downweight attention between corrupted tokens during noisy PT. This indicates that synthetic PPT inhibits noise self-modeling and shapes the subsequent optimization trajectory. Code is available at https://github.com/guox18/formal-language-prepretraining.
Xu Guo, Runyu Peng, Jian Tong +4
Shanghai AI Laboratory · Fudan University · Shanghai Innovation Institute
Large Language Models (LLMs) remain substantially less data-efficient than humans. Pre-pretraining (PPT) on synthetic languages has been proposed to close this gap, with prior work emphasizing highly expressive formal languages such as k-Shuffle Dyck. Inspired by the Language Acquisition Device (LAD) hypothesis, which posits that innate constraints preemptively restrict the learner's hypothesis space to natural-language-like structure, we propose LAD-inspired PPT: pre-pretraining on MP-STRUCT, a formal language whose strings encode hierarchical composition, feature-based dependencies, and long-distance displacement via MERGE, AGREE, and MOVE. A brief 500-step PPT with MP-STRUCT matches strong formal-language baselines in token efficiency while additionally imparting a human-like resistance to structurally implausible languages (e.g., REVERSE). Analyzing simplified variants, we find that MP-STRUCT CORE outperforms k-Shuffle Dyck despite not being definable in C-RASP (a formal bound on transformer expressivity), challenging the prior hypothesis that effective PPT languages must be both hierarchically expressive and circuit-theoretically learnable. We show that functional landmarks, which reduce dependency resolution ambiguity, are a key driver, suggesting that effective PPT design depends not only on expressivity but also on the accessibility of dependency resolution.
Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33% of training tokens needed to reach a certain performance. We validate this prior result for English on a larger set of natural languages across four language families, using two different tokenizers and varying model sizes. We also relate the observed gains (or losses) in token efficiency to quantified linguistic properties of the languages, such as sentence length, morphological richness, and features of dependency syntactic trees (tree depth, number of children, number of crossing dependencies). Our empirical results indicate that the reported gains depend heavily on the experiment setup and the choice of random seed, although we can confirm the trend of stable gains with 128-Dyck pretraining of small models with the Llama tokenizer for most of the examined languages. On a general note, we argue that multiple training runs should be carried out at least for a subset of experiments to avoid the community adopting unstable approaches.
Sofiia Riazhskykh, Nam Luu, Ondřej Bojar
Charles University, Faculty of Mathematics and Physics