It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
Organizations: PleIAs · Sorbonne Center for Artificial Intelligence · Sciences Po Médialab · KU Leuven · EPFL · CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt · Lattice, ENS-PSL · TU Munich, Munich Center for Machine Learning · Paris Dauphine-PSL
Abstract
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present SYNTH, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate SYNTH by training a suite of models: a 56M tiny model (Monad), 0.3B-0.6B dense models (Baguettotron), and a 13B / 1B-active MoE. At iso-compute, SYNTH outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because SYNTH is back-translated from grounded passages, SYNTH-trained models achieve high factual precision despite 10-140x fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our SYNTH dataset, are capable of producing competitive generalist models from a fraction of the training data, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our SYNTH dataset and the suite of Baguettotron models under a permissive license, thus supporting open-source language model development.
Figures & tables
| Model | Mode | Tokens | Sup % | Con % | Inc % | S/(S+C) | macro |
|---|---|---|---|---|---|---|---|
| 50M parameters | |||||||
| Monad (56M, ours) | chat | 180B | 16.3 | 21.7 | 61.9 | 42.9% 3.8 | 16.3% 1.8 |
| 300–400M parameters | |||||||
| SmolLM2-360M-IT | chat | 4T | 26.0 | 16.0 | 58.0 | 61.8% 3.4 | 26.6% 2.1 |
| LFM2.5-350M | chat | 28T | 22.0 | 15.1 | 62.9 | 59.3% 3.1 | 22.0% 1.8 |
| Gemma-3-270M-IT | chat | 2T | 18.4 | 13.6 | 68.0 | 57.6% 4.0 | 19.4% 1.8 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | GPU-hours |
|---|---|
| Inference (generation) | 10,804 |
| Training | 4,936 |
| Fine-tuning | 195 |
| Evaluation | 71 |
| Embeddings | 39 |
| Total | 16,044 |
| Stage-1 category | Calls | Output tokens |
|---|---|---|
| Query model | 7,267 | 13.3M |
| Memorization | 6,006 | 10.2M |
| RAG | 7,108 | 23.2M |
| Arithmetic | 6,927 | 6.9M |
| Creative writing | 14,584 | 29.6M |
| Editing | 15,963 | 27.4M |
| Model | Pre-training | Post-training | Generation share | Total |
|---|---|---|---|---|
| Baguettotron -600M ( Synth ) | 843 | – | 1,536 | 2,379 |
| FineWiki 600M | 941 | – | ||
| FinePDFs-Edu 600M | 1,015 | – |
| Corpus | Mean score |
|---|---|
| Synth | 2.23 [2.10, 2.35] |
| Nemotron-CC | 1.72 [1.56, 1.88] |
| FineWiki | 1.64 [1.49, 1.79] |
| Common Corpus | 1.49 [1.36, 1.62] |
| FinePDFs | 1.34 [1.19, 1.49] |
| FineWeb | 1.21 [1.09, 1.33] |
| Model | Entities | Abstain (%) | S/(S+C), attempted (%) | |
|---|---|---|---|---|
| Baguettotron -MoE | in-seed | 500 | 19.8 [16.5, 23.5] | 81.8 [79.5, 84.0] |
| held-out | 600 | 66.8 [63.0, 70.5] | 61.7 [56.2, 66.8] | |
| no title mention | 302 | 74.2 [69.0, 78.8] | 55.2 [46.1, 63.9] | |
| OLMoE-1B-7B-Instruct | in-seed | 500 | 0.8 [0.3, 2.0] | 82.7 [81.0, 84.4] |
| held-out | 600 | 7.2 [5.4, 9.5] | 81.1 [79.6, 82.6] | |
| no title mention | 302 | 11.9 [8.7, 16.1] | 77.5 [75.0, 79.8] |
| Model | Responses | S/(S+C) (%) | ||||
|---|---|---|---|---|---|---|
| Baguettotron -MoE | attempt | 199 | 960 | 596 | 1,944 | 61.7 [56.2, 66.8] |
| abstain | 401 | 1,988 | 299 | 2,429 | 86.9 [85.0, 88.7] | |
| total | 600 | 2,948 | 895 | 4,373 | 76.7 [74.0, 79.3] | |
| OLMoE-1B-7B-Instruct | attempt | 557 | 10,581 | 2,459 | 11,763 | 81.1 [79.6, 82.6] |
| abstain | 43 | 401 | 87 | 559 | 82.2 [77.3, 86.5] | |
| total | 600 | 10,982 | 2,546 | 12,322 | 81.2 [79.7, 82.6] |
| Pre-training data | Post-training | MCQ (22) | MCQ w/o MMLU (21) | Open-ended (8) |
|---|---|---|---|---|
| FineWiki | SmolTalk + MMLU-aux | 26.1 | 26.2 | 10.2 |
| FinePDFs-Edu | SmolTalk + MMLU-aux | 25.1 | 25.1 | 13.5 |
| Synth | none | 42.2 | 42.2 | 24.3 |
| Group | Benchmark | With traces | Without | |
| Factual recall | NQ-Open | 17.7 | 16.6 | |
| PopQA | 10.0 | 9.8 | ||
| TriviaQA | 24.0 | 25.3 | ||
| WikiFact | 14.3 | 10.9 | ||
| SimpleQA | 1.9 | 2.7 | ||
| Truthfulness | TruthfulQA | 42.6 | 33.5 |
| Model | Training tokens | Micro | Macro |
|---|---|---|---|
| Baguettotron -600M + tool calling | 158B + 1.5B | 53.1 | 53.1 |
| Baguettotron -350M + tool calling | 200B + 1.5B | 47.7 | 46.9 |
| FunctionGemma-270M [ Google DeepMind, 2025 ] | 6T | 49.2 | 45.3 |
| Baguettotron -600M (base) | 158B | 28.0 | 11.1 |