Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how much compute mid-training absorbs. We revisit how this compute should be allocated to a single run or multiple similar optimizations. We find that branches forked from a shared checkpoint under various controlled recipe reaches measurably different regions of parameter space, and establish a form of compatible diversity that extending one run cannot supply. Therefore, we introduce Trajectory Soup, which distributes a mid-training budget over several independent branches, and consolidates strongest checkpoints selected on validation through intra- and inter-trajectory averaging into a single model. A local bias and variance analysis separates the two averaging levels, showing that inter-trajectory averaging removes residual error beyond the reach of averaging within a trajectory, while checkpoint selection carries a bias that bounds how many checkpoints are worth merging. Across model scales, learning-rate schedules, token budgets, and trajectory counts, Trajectory Soup improves aggregate downstream performance over the strongest single-trajectory average under matched budgets and keeps improving as budgets expand, with the advantage preserved after an identical post-training pipeline. These results position trajectory allocation and merging as a practical way to extend the compute-scaling frontier of mid-training beyond serial saturation.
Figures & tables
Figure 1 : Motivation and compute-allocation overview of Trajectory Soup. (a) Performance saturation along serial mid-training trajectories. (b) Compute scaling of raw training and Trajectory Soup. Overall Average Accuracy is plotted against cumulative training FLOPs on a logarithmic scale.
Figure 2Figure 3
Method
Intra
Inter
Expected excess loss in the local quadratic model
Rank-1 checkpoint
No
No
Bn(1)+Vintra+Vinter
Intra-trajectory merging
Yes
No
Bn(K)+Vintra/Keff(K)+Vinter
Inter-trajectory merging
No
Yes
Bsoup(1)+Vintra/N+Vinter/N
Trajectory Soup
Yes
Yes
Bsoup(K)+Vintra/(NKeff(K))+Vinter/N
Table 1: Comparsive analysis of averaging operations under the local quadratic model and the homogeneous independent-anchor assumptions of Theorem 4.3 .
Figure 5 : Performance of different mid-training trajectories and merging strategies as the training-token budget increases. Panels report the overall average accuracy and the five capability categories. Branch identities, merge sizes, and the Limited and Extended budget conventions follow Sec . 5.1 .
Base Model
General Knowledge & Reasoning
Language Modeling
Professional Knowledge
Math
Code
Overall Average
Single-Trajectory Merge
62.80
85.86
63.70
72.33
65.36
68.55
Model Soup (Limited)
62.74
85.80
63.59
72.53
64.74
68.43
Model Soup (Extended)
62.78
85.95
63.68
72.49
65.78
68.67
Trajectory Soup (Limited)
62.99
86.19
64.11
72.66
65.07
68.72
Trajectory Soup (Extended)
63.16
86.06
64.17
72.55
66.19
68.96
Table 2 : Performance comparison of different checkpoint-merging strategies on the mid-training evaluation suite. Each row reports the best evaluated configuration of its strategy; the corresponding checkpoint counts are recorded in Appendix D.5 .
Instruct Model
Math
Code
Knowledge
Reasoning
Instruction Following
Function Call
Overall Average
Single-Trajectory Merge
68.30
46.58
70.51
62.96
59.46
50.67
61.11
Model Soup (Limited)
68.50
47.59
69.36
63.56
58.79
51.52
61.05
Model Soup (Extended)
68.63
46.76
69.65
64.41
59.29
51.63
61.23
Trajectory Soup (Limited)
69.44
48.33
69.81
63.60
58.88
51.49
61.39
Trajectory Soup (Extended)
68.53
47.50
69.26
64.48
60.38
52.85
61.52
Table 3 : Post-training performance after applying the same SFT procedure to mid-training checkpoints produced by different merging strategies. The merging protocols are those of Tab . 2 .
Model
General Knowledge & Reasoning
Language Modeling
Professional Knowledge
Math
Code
Overall Average
Single-Trajectory Merge
47.71
72.85
43.41
54.21
35.96
49.55
Trajectory Soup (Limited)
47.88
73.32
43.70
54.13
35.77
49.64
Trajectory Soup (Extended)
48.35
73.98
44.37
54.87
35.51
50.07
Table 4 : Comparison of single-trajectory merging and cross-trajectory Trajectory Soup during mid-training of the smaller 2B-parameter MoE model. The Limited and Extended budget conventions follow Sec . 5.1 .
Ablation
Coefficient
Selection
General Knowledge & Reasoning
Language Modeling
Professional Knowledge
Math
Code
Overall Average
Baseline
EQUAL
Top-K each
63.16
86.06
64.17
72.55
66.19
68.96
Coefficient
1SQRT
Top-K each
63.22
85.87
64.21
72.26
66.00
68.85
RANK
Top-K each
63.12
85.91
64.16
72.24
65.94
68.80
RSQRT
Top-K each
63.04
86.03
64.26
72.59
65.97
68.89
Selection
EQUAL
Global top-NK
62.99
86.24
64.18
72.46
65.43
68.76
EQUAL
Tail-K per branch
62.78
84.97
64.29
71.77
65.25
68.35
Table 5 : Ablation of merging coefficients and checkpoint-selection strategies. Every row follows the Trajectory Soup (Extended) setting of Tab . 2 , with Baseline the default configuration of Sec . 5.1 .
Figure 10Figure 11
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Supplied value
Model name
Ling-3.0-tiny
Total parameters
7.9B
Active parameters per token
1.3B
Layers
24
Hidden width
1,536
First feedforward layer
Dense
Appendix
Table 6 : Default model configuration. The architectural values are taken from the supplied model description and are not independently verified through a public model release.
Field
Value
Complete branch horizon
600B tokens
Optimizer
Muon
Warmup
1% of training tokens
Weight decay
0.1
Gradient clipping
1.0
Optimizer β1
0.9
Appendix
Table 7 : Training settings shared by all branches. Values recorded in the supplied recipe and held fixed across EXP1–EXP5. The perturbed fields are given in Tab . 8 .
Branch
Perturbed dimension
Data-shuffling seed
Peak learning rate
Global batch
Schedule after warmup
Muon momentum
EXP1
Baseline
1234
3.39×10−4
256
Constant
0.0
EXP2
Data order
1001
3.39×10−4
256
Constant
0.0
EXP3
Learning rate and batch size
1234
2.40×10−4
128
Constant
0.0
EXP4
Learning-rate schedule
1234
3.39×10−4
256
Decay to 3.39×10−5
0.0
EXP5
Optimizer momentum
1234
3.39×10−4
256
Constant
0.9
Appendix
Table 8 : Branch identity and recipe mapping. The five branches forked from the common pretrained checkpoint. Bold entries mark the fields a branch changes relative to the EXP1 baseline; all remaining settings follow Tab . 7 .
Base Model
General Knowledge & Reasoning
Language Modeling
Professional Knowledge
Math
Code
Overall Average
EXP1 Merge †
62.80
85.86
63.70
72.33
65.36
68.55
EXP2 Merge
62.79
85.44
63.97
72.39
64.99
68.46
EXP3 Merge
62.69
85.65
64.02
71.57
65.34
68.34
Appendix
Table 9 : Single-trajectory merges of every branch on the mid-training evaluation suite. Each row is the Single EXP Merge of one branch of Tab . 2 . † Reported as Single-Trajectory Merge in Tab . 2 . Bold marks the best value in each column.
Instruct Model
Math
Code
Knowledge
Reasoning
Instruction Following
Function Call
Overall Average
EXP1 Merge
68.22
47.69
68.98
64.88
57.54
51.75
60.67
EXP2 Merge †
68.30
46.58
70.51
62.96
59.46
50.67
61.11
EXP3 Merge
68.38
45.63
69.87
62.40
58.82
51.57
60.67
Appendix
Table 10 : Single-trajectory merges of every branch after post-training. Each row applies the SFT procedure of Tab . 3 to the Single EXP Merge of one branch. † Reported as Single-Trajectory Merge in Tab . 3 . Bold marks the best value in each column.
Model
General Knowledge & Reasoning
Language Modeling
Professional Knowledge
Math
Code
Overall Average
EXP1 Merge
47.97
73.21
43.92
54.09
34.88
49.49
EXP2 Merge †
47.71
72.85
43.41
54.21
35.96
49.55
Appendix
Table 11 : Single-trajectory merges of both branches of the smaller 2B-parameter MoE model. Each row is the Single EXP Merge of one branch of Tab . 4 . † Reported as Single-Trajectory Merge in Tab . 4 . Bold marks the best value in each column.
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, China · University of Chinese Academy of Sciences, China · Alibaba Group, China +2