Don't stop me now: How Validation Criteria Affect Checkpoint Selection and Early Stopping
Authors: Andrea Apicella, Francesco Isgrò, Andrea Pollastro, Roberto Prevete
Organizations: Department of Information Engineering, Electrical Engineering, and Applied Mathematics (DIEM), University of Salerno, Via Giovanni Paolo II, 132, Fisciano (Salerno), 84084, Italy. · Department of Electrical Engineering and Information Technology, University of Naples Federico II, Via Claudio 21, Naples, 80125, Italy.
Checkpoint selection is a standard component of neural network training, yet the validation criterion used to select a checkpoint is often chosen heuristically. Moreover, the same criterion may be used either only to rank checkpoints after completion of a predefined training run or also to determine when training should stop, thereby affecting both the selected checkpoint and the set of checkpoints available for selection. In this work, we systematically investigate the role of validation criteria under these two settings. We separately vary the training loss, the validation criterion, and the target evaluation metric, and compare post-hoc checkpoint selection, in which training proceeds for all predefined epochs, with patience-based early stopping, in which the validation criterion also controls training termination. We consider three Cross-Entropy, C-Loss, and PolyLoss as training losses, and accuracy, macro-F1, and Matthews correlation coefficient as target metrics. Selection quality is assessed through the relative gap between the test performance of the validation-selected checkpoint and the best-observed test performance for the same target metric over the complete predefined training run.
Figures & tables
Figure 1 : Illustrative validation-loss and validation-metric trajectories computed along the same training trajectory. In this figure, we adopt the conventional orientation in which the loss function ℓ(Val,θe) has to be minimized ( ↓ ), whereas the evaluation metric m(Val,θe) has to be maximized ( ↑ ). Accordingly, eℓ,Val⋆ denotes the epoch minimizing the validation loss, while em,Val⋆ denotes the epoch maximizing the validation metric. The vertical dashed lines indicate the corresponding selected epochs, whereas the horizontal dotted lines indicate the best-observed validation values ℓVal⋆ and mVal⋆ . Displaying the loss in its original orientation is equivalent to the common maximization convention adopted in the formal analysis, since maximizing −ℓ selects the same checkpoint as minimizing ℓ .
Figure 2 : Illustration of patience-based early stopping and post-hoc checkpoint selection on the same training trajectory. The figure displays the validation loss in its classical minimization form. In this example, the best validation loss observed up to t^PS is not improved during the following T epochs. Training therefore stops at tstop , and early stopping retains checkpoint at epoch t^ . Under fixed-horizon training, the trajectory continues until E , and post-hoc selection retains checkpoint θPH . The dashed portion of the trajectory would not be observed under the actual early-stopping protocol and is shown only to illustrate the difference between the two protocols.
Figure 3 : Illustration of post-hoc checkpoint selection using validation loss and the target validation metric. The raw validation loss is shown on the left axis, whereas validation and test metric values are shown on the right axis. The figure displays the validation loss in its original minimization form. Metric-based validation selection retains checkpoint θt^PH(m) , whereas loss-based selection retains checkpoint θt^PH(ℓ) . Their test performances are compared retrospectively with mTest⋆=t∈Emaxm(Test,θt) . Test-set information is shown only to quantify the resulting checkpoint-selection gaps and is not used to select either checkpoint.
Table 1 : Summary of the datasets used in the experimental evaluation retrieved from the UCI Machine Learning Repository ( Asuncion et al., 2007 ) . For each dataset, we report the total number of instances and the number of target classes.
Metric
Dataset Type
Training Loss
Val. Criterion
CE
CLoss
PolyLoss
Metric
MCC
Image
CE
0.083
0.005
0.069
0.005
CLoss
0.055
0.011
0.050
0.007
PolyLoss
0.096
0.005
0.095
0.006
UCI
CE
0.165
0.129
0.159
0.107
CLoss
0.272
0.095
0.217
0.090
Table 2 : Empirical 95th percentile of the relative checkpoint-selection gap under post-hoc checkpoint selection, for each combination of target metric, dataset type, training loss, and validation criterion. Lower values indicate smaller upper-tail gaps and are therefore preferable. In boldface the best values.
Metric
Dataset Type
Training Loss
Val. Criterion
CE
CLoss
PolyLoss
Metric
MCC
Image
CE
0.083
0.005
0.069
0.007
CLoss
0.346
0.025
0.346
0.018
PolyLoss
0.096
0.005
0.095
0.007
UCI
CE
0.185
0.279
0.187
0.324
CLoss
0.395
0.201
0.393
0.382
Table 3 : Empirical 95th percentile of the relative checkpoint-selection gap under patience-based early stopping, for each combination of target metric, dataset type, training loss, and validation criterion. Lower values indicate smaller upper-tail gaps and are therefore preferable. In boldface the best values.
Figure 4 : Relative checkpoint-selection gaps obtained on the UCI datasets using Cross-Entropy as the training loss and macro-F1 as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Figure 5 : Relative checkpoint-selection gaps obtained on the image datasets under post-hoc checkpoint selection. From top to bottom, the target metrics are accuracy, macro-F1, and MCC. Within each panel, columns correspond to the considered validation criteria, while marker shapes identify the loss used to train the model. A value of zero indicates that the selected checkpoint attains the best-observed test performance over the complete predefined training run; lower values are preferable.
Figure 6 : Relative checkpoint-selection gaps obtained on the image datasets under patience-based early stopping with T=20 . From top to bottom, the target metrics are accuracy, macro-F1, and MCC. Within each panel, columns correspond to the considered validation criteria, while marker shapes identify the loss used to train the model. A value of zero indicates no gap with respect to the best-observed test performance over the complete predefined training run; lower values are preferable. The reported gap may reflect both checkpoint ranking and termination before better-performing checkpoints are reached.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Relative checkpoint-selection gaps obtained on the UCI datasets using Cross-Entropy as the training loss and accuracy as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Figure 8 : Relative checkpoint-selection gaps obtained on the UCI datasets using Cross-Entropy as the training loss and MCC as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Figure 9 : Relative checkpoint-selection gaps obtained on the UCI datasets using C-Loss as the training loss and accuracy as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Figure 10 : Relative checkpoint-selection gaps obtained on the UCI datasets using C-Loss as the training loss and macro-F1 as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Figure 11 : Relative checkpoint-selection gaps obtained on the UCI datasets using C-Loss as the training loss and MCC as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Figure 12 : Relative checkpoint-selection gaps obtained on the UCI datasets using PolyLoss as the training loss and accuracy as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Figure 13 : Relative checkpoint-selection gaps obtained on the UCI datasets using PolyLoss as the training loss and macro-F1 as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Figure 14 : Relative checkpoint-selection gaps obtained on the UCI datasets using PolyLoss as the training loss and MCC as the target metric. Columns correspond to the considered validation criteria, while colours identify the parameter-to-sample ratio r . Datasets are ordered according to their GDV.
Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.
Yupeng Chang, Wenxuan Zhang, Yuan Wu
School of Artificial Intelligence, Jilin University · Singapore University of Technology and Design · Key Laboratory of Symbolic Computation and Knowledge Engineering, Jilin University
Gradient boosted decision trees require a stopping rule to avoid overfitting. The standard rule monitors a validation loss and stops if the loss fails to improve for a fixed patience period. However, the patience parameter has no interpretable scale and validation losses can be noisy or implicitly defined by a user-specified gradient. We propose ScoreStop, a gradient-based early-stopping rule that casts the stopping decision at each iteration as a test of the null hypothesis that the current predictor is the population risk minimizer. We use a functional score test, computed on validation data, with a statistic that is scale-invariant in the update direction, with a known asymptotic distribution under the null. Because our test uses gradients rather than loss values, the same construction applies to implicit losses such as LambdaRank, and data-dependent losses such as Cox regression via influence functions. In synthetic experiments and real-data benchmarks, we show that ScoreStop is competitive with loss-based methods.
Oliver J. Hines, Christian L. Hines
Columbia University, NY, USA · Alan Turing Institute, London, UK
Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We identify an empirical selection opportunity within fixed training trajectories: reselecting among checkpoints with near-optimal source accuracy can improve mean target probability quality with small observed changes in mean target accuracy. We study accuracy-constrained reliability selection (AC), which retains checkpoints within a tolerance of the best source-validation accuracy and ranks them by source reliability. Our reference rule aggregates within-set normalized negative log-likelihood (NLL) and class-wise calibration error (CwECE) using D∞. AC uses no target data and requires neither additional training nor weight averaging. We evaluate five domain generalization training algorithms on three benchmarks, using PACS to develop the objectives and a 0.5-percentage-point tolerance. In exploratory aggregation comparisons on 360 OfficeHome and TerraIncognita runs, the reference rule reduces mean target soft-bin squared-gap ECE and CwECE by 0.240% and 0.182%, respectively, and NLL by 0.030 relative to Source-Acc. Mean target accuracy changes by +0.213 percentage points. These results identify opportunities for reliability-aware reselection, while the additional benefit of joint over single-objective ranking remains unresolved.
Jinshi Liu, Jiahao Li, Pan Liu +5
Shenzhen University · Xiamen University · The Hong Kong University of Science and Technology (Guangzhou) +3