It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
Authors: Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, +2 more
Organizations: PleIAs · Sorbonne Center for Artificial Intelligence · Sciences Po Médialab · KU Leuven · EPFL · CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt · Lattice, ENS-PSL · TU Munich, Munich Center for Machine Learning · Paris Dauphine-PSL
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present SYNTH, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate SYNTH by training a suite of models: a 56M tiny model (Monad), 0.3B-0.6B dense models (Baguettotron), and a 13B / 1B-active MoE. At iso-compute, SYNTH outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because SYNTH is back-translated from grounded passages, SYNTH-trained models achieve high factual precision despite 10-140x fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our SYNTH dataset, are capable of producing competitive generalist models from a fraction of the training data, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our SYNTH dataset and the suite of Baguettotron models under a permissive license, thus supporting open-source language model development.
Figures & tables
Figure 1 : Synth pipeline for the memorization task. At Stage 1, auxiliary models are fine-tuned on LLM-generated data. At Stage 2, these models are used to produce synthetic training data at scale.
Figure 2 : Log-scale word counts for Synth tasks and languages.
Figure 3 : Synth excels in the quality/value pairings by a wide margin compared to other open pre-training corpora, scored by Propella-1. We present integrity/safety evaluations in Appendix D .
Figure 4 : Baguettotron sits on the token-efficiency frontier, matching or trailing baselines trained on 80–700 × more tokens. Average accuracy vs. pre-training tokens (log scale) on 22 multiple-choice (left) and 8 open-ended (right) tasks. Diamonds: Baguettotron (ours); circles: open baselines. Token budgets are publicly reported pre-training token counts (post-training tokens excluded). Top-left is better, i.e. more accurate per token of training data.
Model
Mode
Tokens
Sup %
Con %
Inc %
S/(S+C)
macro
∼ 50M parameters
Monad (56M, ours)
chat
180B
16.3
21.7
61.9
42.9% ± 3.8
16.3% ± 1.8
∼ 300–400M parameters
SmolLM2-360M-IT
chat
4T
26.0
16.0
58.0
61.8% ± 3.4
26.6% ± 2.1
LFM2.5-350M
chat
28T
22.0
15.1
62.9
59.3% ± 3.1
22.0% ± 1.8
Gemma-3-270M-IT
chat
2T
18.4
13.6
68.0
57.6% ± 4.0
19.4% ± 1.8
Table 1 : Factual precision on Wikipedia entities ( n=500 ), grouped by parameter tier. S/(S+C) = fraction of decomposed atomic facts Supported by Wikipedia. macro = per-entity Supported / total facts (Inconclusive in the denominator). Bold = best in tier. Pill colour encodes a paired-bootstrap test (B = 10 000) of each baseline against our in-tier reference: green = ours (reference); red = other model in-tier significantly worse compared to ours; gray = not significant ( α=0.05 ).
Figure 5 : Epistemic markers in reasoning traces predict factual quality and output volume. Per-topic FActScore (a) and atomic-fact count (b), conditioned on whether the trace is dominated by confident or uncertain markers (majority count; ties dropped). Diamonds mark conditional means.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Category
GPU-hours
Inference (generation)
10,804
Training
4,936
Fine-tuning
195
Evaluation
71
Embeddings
39
Total
16,044
Appendix
Table 2: Cluster compute. GPU-hours of all Synth -related jobs (H100), including development. Excludes the FineWiki and FinePDFs-Edu 600M runs (broken down in Table 4 ).
Table 4: End-to-end GPU-hours of the controlled 600M runs ( § 4.2 ), excluding Stage-1 frontier supervision.
Figure 6 : Propella evaluation on safety and integrity axes.
Corpus
Mean score
Synth
2.23 [2.10, 2.35]
Nemotron-CC
1.72 [1.56, 1.88]
FineWiki
1.64 [1.49, 1.79]
Common Corpus
1.49 [1.36, 1.62]
FinePDFs
1.34 [1.19, 1.49]
FineWeb
1.21 [1.09, 1.33]
Appendix
Table 5: Educational value by an independent classifier ( fineweb-edu-classifier , scale 0–5, higher is better), 100 English documents per corpus, with 95% CIs.
Model
Entities
n
Abstain (%)
S/(S+C), attempted (%)
Baguettotron -MoE
in-seed
500
19.8 [16.5, 23.5]
81.8 [79.5, 84.0]
held-out
600
66.8 [63.0, 70.5]
61.7 [56.2, 66.8]
no title mention
302
74.2 [69.0, 78.8]
55.2 [46.1, 63.9]
OLMoE-1B-7B-Instruct
in-seed
500
0.8 [0.3, 2.0]
82.7 [81.0, 84.4]
held-out
600
7.2 [5.4, 9.5]
81.1 [79.6, 82.6]
no title mention
302
11.9 [8.7, 16.1]
77.5 [75.0, 79.8]
Appendix
Table 6: Abstention and precision on held-out entities. Abstention rate (LLM judge) and S/(S+C) over attempted responses, with 95% CIs (Wilson for abstention; entity-level bootstrap, B=10,000 , for precision). No title mention : held-out entities whose title never occurs in the Synth training text.
Model
Responses
n
S
C
I
S/(S+C) (%)
Baguettotron -MoE
attempt
199
960
596
1,944
61.7 [56.2, 66.8]
abstain
401
1,988
299
2,429
86.9 [85.0, 88.7]
total
600
2,948
895
4,373
76.7 [74.0, 79.3]
OLMoE-1B-7B-Instruct
attempt
557
10,581
2,459
11,763
81.1 [79.6, 82.6]
abstain
43
401
87
559
82.2 [77.3, 86.5]
total
600
10,982
2,546
12,322
81.2 [79.7, 82.6]
Appendix
Table 7: Atomic-fact counts on the 600 held-out entities , split by the judge’s abstain/attempt label. S , C , I : Supported, Contradicted, Inconclusive facts. 95% CIs by entity-level bootstrap ( B=10,000 ).
Figure 7 : Average accuracy per grouped capability. Multiple-choice tasks are grouped by knowledge domain, open-ended tasks by the capability they probe; the number of tasks per group is in parentheses. Dotted line: 4-way random baseline.
Pre-training data
Post-training
MCQ (22)
MCQ w/o MMLU (21)
Open-ended (8)
FineWiki
SmolTalk + MMLU-aux
26.1
26.2
10.2
FinePDFs-Edu
SmolTalk + MMLU-aux
25.1
25.1
13.5
Synth
none
42.2
42.2
24.3
Appendix
Table 8 : Pre-training data ablation at 600M. Identical architecture, tokenizer, and training steps. Only the Synth model is evaluated without post-training. w/o MMLU : 21-task average, since the MMLU training split is in the web models’ post-training mix.
Group
Benchmark
With traces
Without
Δ
Factual recall
NQ-Open
17.7
16.6
+1.1
PopQA
10.0
9.8
+0.2
TriviaQA
24.0
25.3
−1.3
WikiFact
14.3
10.9
+3.4
SimpleQA
1.9
2.7
−0.8
Truthfulness
TruthfulQA
42.6
33.5
+9.1
Appendix
Table 9: Reasoning-trace ablation. Baguettotron -600M trained on Synth with and without reasoning traces ( § 4.2 ), grouped by capability. Matched on tokens, optimizer steps, unique examples, and seed; one run each.
Figure 8 : Per-benchmark accuracy on the 22 multiple-choice tasks. Dotted line: 4-way random baseline. The Overall row (highlighted) is the simple average across the 22 tasks. Tasks are ordered by knowledge domain (right margin).
Figure 9 : Per-benchmark CORRECT answers on the 8 open-ended tasks. Tasks are ordered by capability group (right margin).
Model
Training tokens
Micro
Macro
Baguettotron -600M + tool calling
158B + 1.5B
53.1
53.1
Baguettotron -350M + tool calling
200B + 1.5B
47.7
46.9
FunctionGemma-270M [ Google DeepMind, 2025 ]
6T
49.2
45.3
Baguettotron -600M (base)
158B
28.0
11.1
Appendix
Table 10 : Tool calling on BFCL v2. Micro: accuracy over all items; macro: mean over categories. The base model emits no tool calls and scores only on the irrelevance categories.
Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and ultimately degrade model performance. Data curation mitigates but cannot eliminate such noise, so pre-training corpora remain noisy in practice. We therefore study whether a lightweight pre-pre-training (PPT) stage based on synthetic data with learnable temporal structure helps resist noisy data during the pre-training (PT) stage. Across various corruption settings, our method consistently improves robustness to noise during PT, with larger relative gains at higher noise levels. For a 1B-parameter model, a synthetic PPT stage with only 65M tokens achieves the same final loss as the baseline while using up to 49% fewer natural-text PT tokens across different noise levels. Mechanistic analyses suggest PPT does not immediately suppress attention to noisy tokens. Rather, PPT-initialized models gradually downweight attention between corrupted tokens during noisy PT. This indicates that synthetic PPT inhibits noise self-modeling and shapes the subsequent optimization trajectory. Code is available at https://github.com/guox18/formal-language-prepretraining.
Xu Guo, Runyu Peng, Jian Tong +4
Shanghai AI Laboratory · Fudan University · Shanghai Innovation Institute
Due to limited supervised training data, large language models (LLMs) are typically pre-trained via a self-supervised "predict the next word" objective on a vast amount of unstructured text data. To make the resulting model useful to users, it is further trained on a far smaller amount of "instruction-tuning" data comprised of supervised training examples of instructions and responses. To overcome the limited amount of supervised data, we propose a procedure that can transform the knowledge in internet-scale pre-training documents into billions of synthetic instruction and answer training pairs. The resulting dataset, called FineInstructions, uses ~18M instruction templates created from real user-written queries and prompts. These instruction templates are matched to and instantiated with human-written source documents from unstructured pre-training corpora. With "supervised" synthetic training data generated at this scale, an LLM can be pre-trained from scratch solely with the instruction-tuning objective, which is far more in-distribution with the expected downstream usage of LLMs (responding to user prompts). We conduct controlled token-for-token training experiments and find pre-training on FineInstructions outperforms standard pre-training and other proposed synthetic pre-training techniques on standard benchmarks measuring free-form response quality. Our resources can be found at https://huggingface.co/fineinstructions .
Ajay Patel, Colin Raffel, Chris Callison-Burch
University of Pennsylvania · University of Toronto · Vector Institute +1
LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. However, reaching the data-bound regime does not mean the model has fully utilized its organic corpus. In this paper, we introduce SynPro, a synthetic data generation framework that helps LLMs more thoroughly learn from limited organic data. SynPro applies two operations, rephrasing and reformat, that present the same organic source in diverse forms to facilitate deeper learning without introducing external information. Both generators are optimized via reinforcement learning with quality, faithfulness, and data influence rewards, and are continuously updated as pretraining plateaus to target content the model has yet to absorb. We pretrain 400M and 1.1B models with 10% of their Chinchilla-optimal tokens (0.8B and 2.2B) from DCLM-Baseline, reflecting a realistic data-bound regime in frontier pretraining. Our results reveal that organic data is significantly underutilized by standard repetition: SynPro unlocks 3.7-5.2x the effective tokens of repetition, even surpassing the non-data-bound oracle that trains on equivalent unique data at the 1.1B scale. Analyses confirm that faithful, model-aware synthesis sustains data-bound scaling without causing distribution collapse. We open-source our code at https://github.com/cxcscmu/SynPro.
Zichun Yu, Chenyan Xiong
Language Technologies Institute, Carnegie Mellon University · Xlue