Organizations: University of Wisconsin – Madison · Baidu Inc. · University of Arizona · University of Toronto · Wilfrid Laurier University · York University
Saved checkpoints record states along a training trajectory, but generally do not determine the updates at states that would be visited under a different schedule. We study how accurately these checkpoints can reconstruct the endpoint of a sequential reference with prescribed update strengths. Under a common local transition model, two checkpoint-index moment conditions characterize all convex merges that agree with this reference through second order. We then prove an information limit that for nondegenerate profiles, no algorithm using only a fixed-length gradient-descent (GD) history with step size h can achieve o(h3) endpoint error uniformly over a fixed class of smooth, strongly convex losses. The lower bound follows from two losses with identical GD checkpoint histories but sequential reference endpoints separated by Ω(h3). \textbf{Quadratic-Accurate Merging} (QAM) achieves a matching uniform O(h3) endpoint error bound. Its explicit coefficients also define the unique profile-dependent merge that exactly matches the sequential GD reference across all fixed quadratic objectives. Across two public Adam checkpoint trajectories (SmolLM3-3B and OpenEuroLLM-Prelude-9B), three windows and three profiles per model, and 15 tasks, QAM shows mixed results for short windows and broader advantages over \textbf{Warmup-Stable and Merge} (WSM) for longer windows. Matched-moment GSM8K diagnostics further show that local consistency alone does not fully determine downstream scores. These results characterize the reconstruction limits of saved histories, provide a coefficient rule that attains the optimal rate, and assess its practical utility.
Figures & tables
Figure 1: Illustrative example of QAM with a linearly decaying update profile Wj .
Model
Size
Checkpoints
Interval
Windows
SmolLM3
∼3.075 B
last 15 constant-LR checkpoints; iter 3.64M–4.20M
∼100 B tokens
Last-5/10/15
OpenEuroLLM Prelude
∼9.095 B
last 40 constant-LR checkpoints; iter 859200–952800
∼40 B tokens
Last-10/20/40
Table 1: Model details.
Figure 2: Paired score differences ( QAM minus WSM ) on SmolLM3 across six capability groups and profiles. Average weights tasks by evaluation-set size.
Figure 3: Same as Figure 2 , but for OpenEuroLLM Prelude.
Trajectory
All
Latest
Best
SmolLM3
1092/48/210
102/5/28
70/3/62
Prelude
2716/121/313
117/9/9
77/2/56
Total
3808/169/523
219/14/37
147/5/118
Table 2: QAM versus individual checkpoints across all nine configurations per trajectory. Entries are wins/ties/losses. “All” counts comparisons with every constituent checkpoint. “Latest” and “Best” each contain 135 task–configuration comparisons per trajectory.
Trajectory
QAM window
Linear QAM
Best WSM
Δ
SmolLM3
Last-10
43.71
41.74
+1.97
SmolLM3
Last-15
43.29
41.74
+1.55
Prelude
Last-20
44.47
42.91
+1.56
Prelude
Last-40
45.03
42.91
+2.12
Table 3: GSM8K with a fixed linear QAM profile versus the best WSM score over all nine tested window/profile combinations on each trajectory. Scores average flexible and strict extraction.
SmolLM3 Last-15
Prelude Last-40
Flex
Strict
Flex
Strict
Profile
λ
Mix
MaxEnt
Mix
MaxEnt
Mix
MaxEnt
Mix
MaxEnt
Linear
0.25
36.39
36.85
34.65
35.18
12.3
12.5
6.7
6.9
0.50
38.89
39.27
37.60
38.36
21.2
18.1
15.5
11.5
0.75
40.86
39.20
40.49
38.67
38.8
32.2
37.8
29.6
Cosine
0.25
39.50
40.33
38.44
39.73
19.2
18.4
13.2
13.6
Table 4: GSM8K intermediate-target controls (%). Mix stands for p(λ) . MaxEnt matches its two moments on the same grid.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Tasks
Metric
Knowledge
ARC-Easy, ARC-Challenge ( Clark et al., 2018 ) , OpenBookQA ( Mihaylov et al., 2018 ) , SciQ ( Welbl et al., 2017 ) , MMLU (5-shot) ( Hendrycks et al., 2021 )
acc.
Commonsense
HellaSwag ( Zellers et al., 2019 ) , PIQA ( Bisk et al., 2020 ) , COPA ( Gordon et al., 2012 ) , WinoGrande ( Sakaguchi et al., 2020 )
acc.
Reading
BoolQ ( Clark et al., 2019 )
acc.
Math
GSM8K flexible, GSM8K strict ( Cobbe et al., 2021 )
exact match
Semantics
SST-2 ( Socher et al., 2013 ) , WiC ( Pilehvar and Camacho-Collados, 2019 ) , WSC ( Levesque et al., 2012 )
acc.
Dialogue
MuTual ( Cui et al., 2020 )
MRR
Appendix
Table 5: Capability groups.
Model
Delimiter
No newline
Loop
Median length
λ=0
13.7
83.3
82.8
1027
λ=1/2
34.3
57.2
60.0
840
λ=3/4
79.1
11.0
14.8
374
λ=1
91.0
0.3
4.7
298
Appendix
Table 6: Output diagnostics along the linear Prelude Last-40 interpolation path on GSM8K. The delimiter column records the fraction of generations containing the standard GSM8K answer delimiter. The no-newline column records outputs without a newline. The loop column records repeated 10-gram behavior. Median length is measured in characters.
Profile
Method
Flexible
Strict
No newline
Loop
Median length
Linear
WSM
33.4
30.6
0.8
9.9
213
Linear
QAM
43.4
43.2
0.4
3.6
221
Cosine
WSM
37.1
35.3
0.2
5.9
233
Cosine
QAM
45.4
44.8
0.4
2.9
252
1−⋅
WSM
32.8
29.9
1.4
11.7
218
1−⋅
QAM
41.5
41.0
0.4
3.6
231
Appendix
Table 7: SmolLM3 Last-15 diagnostics on the 1,319-question GSM8K test set.
SmolLM3 Last-15
Prelude Last-40
Kernel
Flex
Strict
Flex
Strict
QAM
43.37
43.21
45.41
44.66
MaxEnt
43.75
43.52
46.10
45.34
Appendix
Table 8: Five-shot GSM8K scores for linear QAM and moment-matched MaxEnt controls.
Model
Window
Profile
Knowledge
Commonsense
Reading
Math
Semantics
Dialogue
SmolLM3
Last-5
Single-checkpoint average
58.27
58.64
76.71
34.33
63.48
69.68
Linear
60.90 (+2.63)
59.86 (+1.22)
78.75 (+2.04)
41.78 (+7.45)
68.09 (+4.61)
70.34 (+0.66)
Cosine
60.98 (+2.71)
59.85 (+1.21)
78.07 (+1.36)
41.21 (+6.88)
66.30 (+2.82)
70.79 (+1.11)
1−⋅
60.92 (+2.65)
59.85 (+1.21)
79.27 (+2.56)
40.30 (+5.97)
68.77 (+5.29)
69.90 (+0.22)
Last-10
Single-checkpoint average
58.34
58.59
76.20
34.09
62.98
69.68
Linear
61.24 (+2.90)
59.69 (+1.10)
79.63 (+3.43)
43.71 (+9.62)
68.40 (+5.42)
69.56 (-0.12)
Appendix
Table 9: SmolLM3 QAM compared with the average single checkpoint.
Model
Window
Profile
Knowledge
Commonsense
Reading
Math
Semantics
Dialogue
OpenEuroLLM
Last-10
Single-checkpoint average
61.16
61.94
79.52
33.16
57.56
71.65
Linear
64.55 (+3.39)
63.44 (+1.50)
81.87 (+2.35)
42.80 (+9.64)
64.19 (+6.63)
72.75 (+1.10)
Cosine
63.86 (+2.70)
63.22 (+1.28)
81.41 (+1.89)
41.43 (+8.27)
61.53 (+3.97)
72.04 (+0.39)
1−⋅
64.19 (+3.03)
63.42 (+1.48)
82.32 (+2.80)
42.65 (+9.49)
60.22 (+2.66)
72.42 (+0.77)
Last-20
Single-checkpoint average
61.09
61.94
79.66
33.37
58.65
71.42
Linear
64.06 (+2.97)
63.77 (+1.83)
82.97 (+3.31)
44.47 (+11.10)
65.74 (+7.09)
72.70 (+1.28)
Appendix
Table 10: OpenEuroLLM Prelude QAM compared with the average single checkpoint.
Linear
Cosine
1−⋅
Metric
WSM
QAM
WSM
QAM
WSM
QAM
GSM8K flexible
11.83
45.41
14.03
44.66
12.66
42.84
GSM8K strict
5.38
44.66
9.48
44.28
7.13
42.30
flexible − strict
6.44
0.76
4.55
0.38
5.53
0.53
Appendix
Table 11: GSM8K scores for Prelude Last-40 under flexible and strict answer extraction.
SmolLM3
Prelude
Task / metric
WSM
QAM
Δ
WSM
QAM
Δ
ARC-Challenge
47.61
48.12
+0.51
50.94
50.68
-0.26
ARC-Easy
79.46
79.84
+0.38
80.60
80.85
+0.25
BoolQ
79.17
79.97
+0.80
83.06
82.97
-0.09
COPA
87.00
88.00
+1.00
92.00
92.00
0.00
HellaSwag
54.89
55.20
+0.31
58.97
59.28
+0.31
Appendix
Table 12: Task-wise best observed scores. WSM uses all tested windows; QAM uses intermediate and long windows. Scores are percentages, except MuTual which is 100×MRR .
Trajectory
Window
Profile
All
Latest
Best
SmolLM3
Last-5
Linear
68/0/7
14/0/1
12/0/3
SmolLM3
Last-5
Cosine
66/2/7
14/0/1
11/1/3
SmolLM3
Last-5
1−⋅
58/5/12
10/2/3
8/1/6
SmolLM3
Last-10
Linear
116/5/29
10/1/4
7/0/8
SmolLM3
Last-10
Cosine
123/4/23
12/0/3
7/0/8
SmolLM3
Last-10
1−⋅
116/7/27
9/1/5
6/1/8
Appendix
Table 13: QAM versus individual checkpoints, broken down by window and profile.
Linear
Cosine
1−⋅
Task
Latest
Best
QAM
W/T/L
QAM
W/T/L
QAM
W/T/L
Last-5
ARC-Challenge
45.39
46.59
47.01
5/0/0
47.10
5/0/0
47.35
5/0/0
ARC-Easy
77.69
77.78
79.00
5/0/0
79.50
5/0/0
79.21
5/0/0
OpenBookQA
33.60
34.40
36.00
5/0/0
36.40
5/0/0
34.40
4/1/0
SciQ
95.00
95.10
95.40
5/0/0
95.50
5/0/0
95.00
3/1/1
Appendix
Table 14: SmolLM3 QAM versus individual checkpoints for all windows and profiles.
Linear
Cosine
1−⋅
Task
Latest
Best
QAM
W/T/L
QAM
W/T/L
QAM
W/T/L
Last-10
ARC-Challenge
47.61
48.46
50.77
10/0/0
49.23
10/0/0
49.91
10/0/0
ARC-Easy
79.08
79.50
80.68
10/0/0
80.30
10/0/0
80.26
10/0/0
OpenBookQA
31.80
33.20
33.80
10/0/0
35.00
10/0/0
34.80
10/0/0
SciQ
95.50
95.90
95.60
7/0/3
95.50
6/1/3
95.40
5/1/4
Appendix
Table 15: Prelude QAM versus individual checkpoints for all windows and profiles.