From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning
Organizations: School of Artificial Intelligence, Jilin University · Singapore University of Technology and Design · Key Laboratory of Symbolic Computation and Knowledge Engineering, Jilin University
Abstract
Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.
Figures & tables
| Rule | Budget gain | vs. matched NLL | vs. final |
|---|---|---|---|
| : Full 32 | : Full | : Full | |
| A. Original mathematical pools 60 trajectories / 19 configurations | |||
| GSM8K + MATH; full = 305/312/313 examples from 228–244 source groups | |||
| Generated accuracy | |||
| Checkpoint agreement | |||
| Matched NLL | (reference) | ||
| A. Independent validation selection | 81 trajectories, 26 configurations | ||
|---|---|---|---|
| Rule | Gain vs. matched NLL | Gain vs. final | |
| Generated accuracy | |||
| Checkpoint agreement | |||
| B. Matched-half selection | 75 trajectories, 24 configurations | ||
| Rule | Selection-half gain | Complementary-half gain | Paired difference |
| Empirical accuracy maximum | |||
| Sensitivity | Rule | Budget gain | vs. matched NLL | vs. final |
|---|---|---|---|---|
| One variant / source | Generated accuracy | |||
| Agreement | ||||
| Test-question resampling | Generated accuracy | |||
| Agreement |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Question | Traj. / config. | Selection and scoring | Direct contrast |
|---|---|---|---|
| Recovery over NLL vs. gain over final | 81 / 26 | Select on original validation; score on independent test items. | Gain over matched NLL and gain over the final checkpoint. |
| Same-item vs. unseen-item gain | 75 / 24 | Select once on half the evaluation items; score that choice on both halves. | Selection-half minus complementary-half gain; fixed selection size. |
| Validation-budget response | 60 / 19 | Select on shared nested validation subsets; score on a fixed independent test. | Full-minus-32 gain ( ); same-budget comparisons. |
| Frozen GSM8K new-pool extension | 12 / 4 | Select on 32–1,024 newly constructed source questions; reuse existing training trajectories and test predictions. | 1,024-minus-32 gain ( ); gains over the final checkpoint and matched NLL. |
| Cross-domain Commonsense replication | 12 / 4 | Train 12 new trajectories; select on 32 or 1,024 questions across eight Commonsense tasks and score frozen choices on 4,000 independent test questions. | 1,024-minus-32 gain ( ); gains over the final checkpoint and matched NLL. |
| Rule | GSM8K (48 traj.) | MATH (15) | Commonsense (18) |
|---|---|---|---|
| Original NLL | |||
| Generated accuracy | |||
| Agreement |
| Rule | |||
|---|---|---|---|
| Pooled (60 trajectories, 19 configurations) | |||
| Generated accuracy | |||
| Agreement | |||
| Matched NLL | |||
| GSM8K (45 trajectories, 15 configurations) | |||
| Generated accuracy | |||
| Rule | Budget | ||
|---|---|---|---|
| Generated accuracy | 32 | ||
| 64 | |||
| 128 | |||
| 256 | |||
| Full | |||
| Agreement | 32 |
| Rule | Reference | ||||
|---|---|---|---|---|---|
| Generated accuracy | -0.06 | +0.06 | +0.27 | +0.35 | |
| Generated accuracy | +0.05 | +0.17 | +0.38 | +0.45 | |
| Agreement | -0.06 | -0.02 | +0.02 | +0.03 | |
| Agreement | +0.04 | +0.08 | +0.13 | +0.14 | |
| Matched NLL | -0.10 | -0.10 | -0.11 | -0.10 | |
| Matched NLL | 0.00 | 0.00 | 0.00 | 0.00 |
| Rule | Contrast | Mean | Config. 95% interval | Test-question 95% interval |
|---|---|---|---|---|
| Generated accuracy | ||||
| Agreement | ||||
| Rule | Full 32 | Full vs. final | Full vs. matched NLL |
|---|---|---|---|
| Generated accuracy | |||
| Agreement | |||
| Matched NLL |
| Rule | Paired contrast | Mean | Configuration CI | Test-question CI |
|---|---|---|---|---|
| Generated accuracy | Full versus final checkpoint | |||
| Full versus matched NLL | ||||
| Full minus 32 | ||||
| Agreement | Full versus final checkpoint | |||
| Full versus matched NLL | ||||
| Full minus 32 |
| Task group | Model / adaptation | Steps | Seeds |
| GSM8K | Llama-3-8B | 252 | 3 |
| GSM8K | Mistral-7B | 252 | 3 |
| GSM8K | Qwen2.5-7B | 252 | 3 |
| GSM8K | Qwen2.5-7B-Instruct | 252 | 3 |
| GSM8K | Qwen3-1.7B | 252 | 3 |
| GSM8K | Qwen3-1.7B (full fine-tuning) | 252 | 3 |
| Source pool | Before split | Training | Validation | Intersection |
|---|---|---|---|---|
| GSM8K | 3,000 / 2,421 | 2,695 / 2,177 | 305 / 244 | 0 / 0 |
| MATH | 2,990 / 2,083 | 2,677 / 1,854 | 313 / 229 | 0 / 0 |
| Commonsense | 3,000 / 2,992 | 2,683 / 2,676 | 317 / 316 | 0 / 0 |