We study whether persistent out-of-distribution (OOD) degradation can be predicted before it is directly observed using only source-side training dynamics. In a controlled shortcut-learning setting, a simple logistic regression predictor develops a clear prospective signal, while training time alone does not. Temporal summaries of the source-side quantities are substantially more informative than their current values. When transferred without additional training from a CNN to an MLP, confidence and entropy dynamics retain substantial predictive information. These results provide a proof of principle that source-side training dynamics can contain an early warning signal for future OOD failure.
Figures & tables
Figure 1 : Architecture of the CNN used as the main model organism.
Training parameter
Value
Batch size B
50
Optimizer
Adam
Learning rate
10−3
(β1,β2)
(0.9,0.999)
Loss
Cross entropy
Maximum training steps Tmax
12000
Table 1: Training parameters used throughout the experiments.
Observable
Definition
Training loss
Cross-entropy loss on the current batch
Training accuracy
Accuracy on the current batch
Gradient norm
∥gt∥
Parameter norm
∥θt∥
Local parameter change
∥θt−θt−1∥
Global parameter change
∥θt−θ0∥
Table 2: Source-side quantities recorded during training. Here θt and gt denote the concatenated parameter and gradient vectors at training step t .
Figure 2 : Sample CNN training trajectories for different random seeds. Shown are the source and clean accuracies.
Figure 3 : Sample training trajectories for the MLP trained on the shortcut dataset. Different random seeds lead to different training behavior. Shown are the clean accuracy and the source accuracy.
Figure 4 : Performance of the time-only predictor on the balanced prediction dataset. Shown are the predictor-training trajectories and the held-out test trajectories. The upper panels show the predicted probability as a function of the number of training steps before degradation, and the lower panels show the corresponding fraction of flagged trajectories.
Figure 5 : Performance of the logistic regression predictor on the balanced prediction datasets. Shown are the predictor-training and held-out test trajectories from the training run, together with two independently generated CNN runs. The upper panels show the predicted probability as a function of the number of training steps before degradation, and the lower panels show the corresponding fraction of flagged trajectories.
Figure 6 : Performance of the same trained predictor over the last 800 training steps before degradation. Shown are all qualifying degrading trajectories from the training run, including those used to train the predictor, together with the same two independently generated CNN runs as in Figure 5 . The upper panels show the predicted probability as a function of the number of training steps before degradation, while the lower panels show the corresponding fraction of flagged trajectories.
Predictor inputs
Training-run test
Independent CNN 1
Independent CNN 2
All, four features
0.74
0.72
0.71
All, mean
0.71
0.70
0.70
All, current
0.63
0.61
0.63
No loss/accuracy
0.71
0.69
0.67
Output statistics
0.66
0.67
0.64
Parameter/gradient
0.65
0.63
0.67
Table 3 : Prediction accuracy on the balanced prediction datasets for different choices of source-side metrics and temporal features. The datasets contain 114 , 28 , and 88 trajectories for the Training-run test set, Independent CNN 1, and Independent CNN 2, respectively. “All” denotes all source-side metrics. “Four features” denotes the current value, mean, standard deviation, and slope over the preceding window. Unless stated otherwise, the four temporal features are used.
Features
Training run
Independent CNN 1
Independent CNN 2
Four features
0.03
0.03
0.03
Mean
0.09
0.12
0.13
Current
0.75
0.82
0.77
Table 4 : Mean trajectory-level false-positive rates on trajectories with no observed degradation for different temporal features. The training run, Independent CNN 1, and Independent CNN 2 contain 84 , 6 , and 10 such trajectories, respectively. For each trajectory, the false-positive rate is the fraction of eligible windows classified as positive, and the reported value is the mean across trajectories.
Figure 7 : Transfer of CNN-trained predictors to MLP trajectories without additional training. Left: predictor using all source-side quantities and the four temporal features. Center: predictor using only the parameter and gradient quantities. Right: predictor using only the output statistics. The complete predictor transfers only weakly, the parameter and gradient predictor transfers poorly, and the output statistics retain a substantially clearer prospective signal.
Predictor
CNN held-out
MLP
All metrics, four features
0.74
0.58
All metrics, mean
0.71
0.61
Output statistics, four features
0.66
0.65
Parameter/gradient, four features
0.65
0.38
Time only
0.50
0.50
Table 5 : Transfer of CNN-trained predictors to the MLP architecture. The predictors and their normalization are trained only on CNN trajectories and are applied to the MLP trajectories without additional training. The balanced evaluation datasets contain 114 held-out CNN trajectories and 49 MLP trajectories. Reported values are prediction accuracies on these datasets.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Representative training trajectories on the clean dataset for the CNN (left) and MLP (right). The CNN reaches at least 95% accuracy within the maximal training time Tmax=12000 steps, while the MLP reaches approximately 95% .
Figure 9 : Performance on the balanced prediction datasets after removing the training loss and training accuracy from the predictor inputs while retaining the four temporal features for the remaining source-side metrics.
Figure 10 : Comparison of different temporal features on the balanced prediction datasets. Left: predictor using only the current value of each source-side metric. Right: predictor using only the mean over the preceding window.
Figure 11 : Comparison of different groups of source-side metrics on the balanced prediction datasets, using all four temporal features. Left: parameter- and gradient-based quantities. Right: output statistics consisting of confidence, confidence on correctly classified examples, and entropy.
University of Wisconsin Madison Madison, WI, USA · IMC University of Applied Sciences Krems, Austria · Indian Institute of Technology Roorkee Roorkee, India
School of Life Science and Technology, Institute of Science Tokyo · Department of Computational Biology and Medical Sciences, Graduate School of Frontier Sciences, The University of Tokyo