Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a better starting point for subsequent learning of natural language. In this paper, we show that in small language models, such a warm-up improves specific capabilities that are not reflected in language-modeling perplexity. Our warm-up uses an abstract stack-manipulation task that requires compositional and state-tracking capabilities. Allocating as little as 1% of pretraining tokens to this data improves multi-hop question answering by up to 3.9 F1 points on MUSIQUE, with additional gains on HOTPOTQA and 2WIKIMULTIHOPQA despite comparable language-modeling perplexity. Controlled experiments show that the warm-up substantially accelerates the acquisition of deeper reasoning chains. We also explore what drives this transfer. First, the structure of the data matters: replacing the stack task with a queue fails to produce the same gains. Second, the gains are specific: performance improves on sequential reasoning chains, with no consistent benefit on tasks that combine or compare independent facts. Third, timing matters: mixing abstract data with natural language is far less effective than an initial dedicated phase, and exposure after pretraining completely removes the benefits. The early advantage persists through billions of subsequent language tokens. These results show that early abstract training can reliably shape the capabilities language models later acquire.
Figures & tables
Figure 1: We identify three practical principles for using abstract data to pretrain language models. (Left) It is most effective in an initial dedicated phase. (Middle) The benefits are not always apparent in language-modeling perplexity, but they show up more clearly in evaluations of specific capabilities such as multi-hop reasoning. (Right) The choice of abstract data matters: the benefits depend on the model learning specific structures rather than a generic headstart on the optimization.
Type of abstract data
Example sequences
Stack
push(A) p( push(B) p( pop() p( push(C) p( <SEP> p( A p( C p( <EOS> p(
Queue
push(A) p( push(B) p( pop() p( push(C) p( <SEP> p( B p( C p( <EOS> p(
StackRand
push(A) p( push(B) p( pop() p( push(C) p( <SEP> p( C p( A p( <EOS> p(
Sort
Z p( C p( B p( T p( A p( <SEP> p( A p( B p( C p( T p( Z p( <EOS> p(
Table 1: Examples of abstract data. Boxes represent tokens. The loss is computed on colored ones.
Figure 2: Stack improves multi-hop QA without improving language modeling. We compare language-only pretraining with Stack followed by language pretraining at two model sizes (SmolLM-135M and SmolLM-360M). (Top) Token-level F 1 after benchmark-specific fine-tuning. (Bottom) Validation loss during the pretraining phase on FineWeb-Edu .
Figure 3: Gains on QA benchmarks concentrate on sequentially composition. Mean token-level F 1 by question type. Models trained on Stack perform consistently better in every chain-dominated category, not in merge- and comparison-based categories. Annotations indicate relative differences with the baseline trained only on language from standard random initialization (no abstract data).
Figure 4: (Left) On the synthetic Depo task, the baseline models with standard pretraining on language struggle to perform multi-hop operations. (Right) In comparison, our models exposed to the Stack data learn faster and reach near-perfect accuracy across all tested hop counts.
Figure 5: Different types of abstract data are not equally effective. The test accuracy on Depo throughout training is best for Stack . See Figure 14 for additional options from prior work.
Figure 6: Similar loss during FineWeb - Edu pretraining without and with prior exposure to variants of Stack .
Figure 7: Close variants of the abstract data perform much worse. Replacing the simulated stack with a queue is worse than even a random initialization on QA benchmarks ( left , token-level F 1 after language pretraining) and on Depo ( right , accuracy after direct transfer). The StackRand variant is closer to the original data and performs accordingly better, but still worse than Stack .
Figure 8: Matching weight magnitudes does not reproduce Stack ’s benefits. The test accuracy on Depo improves slightly (center left/right) with random weights whose magnitude matches those trained on Stack , but they do not reach the near-perfect performance of the intact weights (right).
Figure 9: The initial training on Stack produces the largest gain. Mean change in F 1 over the three QA benchmarks relative to language-only pretraining.
Figure 10: Training order affects deep-chain reasoning acquisition. Four-hop Depo learning curves after pretraining on natural language (NL) only, Stack → language, or language → Stack .
Figure 11: Repeated exposure to abstract data has mixed effects.
Figure 12: The benefits of initial Stack training persist throughout language pretraining. Mean token-level F 1 after separately fine-tuning successive SmolLM-135M checkpoints on each natural language multihop QA benchmark.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Seq. len.
Vocab.
Optimization
Task-specific setting
Procedural tasks from Jiang et al. (2026)
Stack
128
102
10k steps, eff. bs 64, 5×10−4 const.
push/pop; emit stack top-first
Identity
128
102
10k steps, eff. bs 64, 5×10−4 const.
copy input
Reverse
128
102
10k steps, eff. bs 64, 5×10−4 const.
reverse input
Delete
128
102
10k steps, eff. bs 64, 5×10−4 const.
remove specified token
Set
128
102
10k steps, eff. bs 64, 5×10−4 const.
emit unique elements
Appendix
Table 2: Procedural pretraining configurations used in the controlled experiments. All models are 4-layer, 4-head, 512-dimensional GPT-2 transformers ( ∼ 12.7M parameters) trained with AdamW over three seeds. For tasks adopted from prior work, we follow the original protocols as closely as possible while adapting them to our controlled model and transfer setting. The FIFO and shuffled-output Stack controls use the same training configuration as Stack .
Model
Language tokens
Optimization steps
SmolLM-135M
2.69B
13,701
SmolLM-360M
7.20B
36,621
Appendix
Table 3: Language-pretraining budgets. All runs use 2,048-token sequences and an effective batch size of 96 sequences.
Muon LR
Weight decay
Auxiliary AdamW LR
Validation loss ↓
0.016
0.01
3×10−3
2.9144
0.020
0.01
3×10−3
2.9150
0.010
0.01
3×10−3
2.9225
0.008
0.01
3×10−3
2.9275
0.008
0.10
3×10−3
2.9361
0.010
0.10
3×10−3
2.9430
Appendix
Table 4: Muon hyperparameter search on SmolLM-135M. Each configuration is trained for 2.69B FineWeb-Edu tokens. The selected configuration has the lowest validation loss (in bold).
Benchmark
Training
Development
MuSiQue
19,937/19,938
2,417/2,417
HotpotQA
18,786/20,000
6,813/7,405
2WikiMultihopQA
19,674/20,000
12,075/12,576
Appendix
Table 5: Retained natural language QA benchmark examples after context-length filtering. Counts are retained/original examples using the SmolLM tokenizer.
Figure 13: Language-modeling validation loss on TinyStories and FineWeb-Edu . Stack -initialized models exhibit higher loss than randomly initialized models throughout language pretraining.
Figure 14: Full comparison of abstract data on DEPO . Accuracy by hop count during downstream training after initial training on abstract data, without intervening language pretraining. Stack enables the fastest learning to near-perfect accuracy across all four hop counts; other types of abstract data yield slower or less complete learning, particularly on longer chains.
Muon LR
Auxiliary LR
MuSiQue
HotpotQA
2WikiMH
4×10−3
10−4
23.57
25.07
44.96
1.3×10−3
3.3×10−5
35.97
38.66
49.26
4×10−4
10−5
37.93
40.57
50.01
Language-only baseline
38.10
41.34
50.37
Stack → FineWeb-Edu
41.99
43.54
51.24
Appendix
Table 6: Learning-rate sweep for post-language Stack training. Downstream F 1 for SmolLM-135M, using three pretraining seeds per configuration. Language-only and initial Stack exposure are included as references.