cs.LGSep 29, 2026

From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning

Authors: Yupeng Chang, Wenxuan Zhang, Yuan Wu

Organizations: School of Artificial Intelligence, Jilin University · Singapore University of Technology and Design · Key Laboratory of Symbolic Computation and Knowledge Engineering, Jilin University

Abstract

Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Instance Selection to Fixed-Pool Data Recipe Search for Supervised Fine-Tuning

    May 13, 2026Haodong Wu, Jiahao Zhang, Lijie Hu +1Supervised FinetuningModel Fine-Tuning

  2. Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning

    May 13, 2026Siyuan Liu, Tinghong Chen, Xinghan Li +2Large Language Model Fine-TuningModel Fine-Tuning

  3. Diversity in Large Language Models under Supervised Fine-Tuning

    Apr 30, 2026Roman Klypa, Oleksandr CherednichenkoDiversityTaco