Do Not Train Away Uncertainty: Early Uncertainty Anchored Calibration
Authors: Yutong Xie, Jiawei Tang, Zhenglin Hua, Yuxiang Ma, Si Qin, Yaxin Hou, Hui Liu, Junhui Hou, +1 more
Organizations: School of Software Engineering, Southeast University, Nanjing, China · School of Computer Science and Engineering, Southeast University, Nanjing, China · School of Computing Information Sciences, Saint Francis University, Hong Kong, China · Department of Computer Science, City University of Hong Kong, Hong Kong, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China
Deep neural networks, including large language models, have achieved remarkable performance across various tasks. However, they are prone to overconfidence during training or fine-tuning. In this work, we observe a consistent phenomenon across different models that the early model is better calibrated, while later training or fine-tuning yields marginal accuracy gains but substantially increases calibration errors. Our analysis suggests that the early model retains uncertainty awareness in both its predictions and features, which is gradually lost with continued training. To avoid training away this uncertainty awareness, we propose \textbf{EUA-Cal}, a novel method that exploits the \textbf{E}arly model as an \textbf{U}ncertainty \textbf{A}nchor for \textbf{Cal}ibration. EUA-Cal introduces early prediction regularization to preserve early predictive uncertainty and prototype structure regularization to exploit uncertainty reflected in the early feature space, jointly mitigating overconfidence. Extensive experiments on image classification and multiple-choice question answering across eight diverse models demonstrate that EUA-Cal outperforms state-of-the-art calibration methods.
Figures & tables
Figure 1: (a) Training and validation classification errors and validation ECE for ResNet-50 on CIFAR-10, ViT-B/32 on CIFAR-100, and Llama3-8B on ARC-Challenge. (b) : t-SNE visualizations of features from ResNet-50 on CIFAR-10 at epochs 40, 60, and 200. Lighter shades indicate features associated with lower confidence. More examples are provided in Appendix B .
Figure 2: An illustration of the proposed EUA-Cal approach for calibration. In the pre-learning stage, we freeze the early model as an uncertainty anchor at the selected training epoch. Its predictions are used for EPR, while confidence-weighted class prototypes constructed from early features provide structural targets for PSR. The overall objective guides training or fine-tuning to preserve uncertainty and mitigate overconfidence.
Figure 3: (a) Confidence of certain and uncertain samples under the vanilla CE model and our method. Samples are divided according to the confidence of the selected early model. (b) Classification errors and ECE of the prototype predictions on the validation set using prototypes constructed from different training epochs.
15 Method
CIFAR-10
CIFAR-100
Tiny-ImageNet
Avg. Rank
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
Vanilla
3.36 ± 0.17
3.31 ± 0.19
94.57 ± 0.15
9.19 ± 0.16
9.07 ± 0.26
76.75 ± 0.62
5.54 ± 0.44
5.52 ± 0.43
64.45 ± 0.45
10.11
LS
3.74 ± 0.24
3.93 ± 0.33
94.78 ± 0.13
3.96 ± 0.22
3.99 ± 0.24
76.77 ± 0.13
3.78 ± 0.27
3.70 ± 0.36
63.58 ± 0.67
8.78
FL
1.81 ± 0.31
1.58 ± 0.27
94.40 ± 0.28
1.77 ± 0.31
1.71 ± 0.33
75.88 ± 0.35
1.57 ± 0.11
1.48 ± 0.18
62.94 ± 0.33
6.67
CRL
0.86 ± 0.09
0.61 ± 0.03
93.94 ± 0.22
5.81 ± 0.30
5.74 ± 0.27
76.87 ± 0.39
2.99 ± 0.31
3.04 ± 0.33
63.43 ± 0.33
7.67
Distillation
3.22 ± 0.52
3.27 ± 0.49
94.10 ± 0.67
7.51 ± 0.20
7.40 ± 0.16
76.64 ± 0.12
6.04 ± 0.24
6.01 ± 0.19
65.21 ± 0.20
10.22
Table 1: Performance comparison between different methods with ResNet-50 on CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets. Results are reported as mean ± standard deviation (%). The best and second best results are shown in bold and underlined , respectively.
14 Method
ARC-C
ARC-E
SciQ
Avg. Rank
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
Vanilla
17.24 ± 0.25
17.12 ± 0.31
81.31 ± 0.64
6.54 ± 0.03
6.48 ± 0.03
92.81 ± 0.12
6.42 ± 0.38
6.10 ± 0.41
93.40 ± 0.52
8.56
LS
15.23 ± 0.96
15.54 ± 1.26
79.81 ± 0.96
3.83 ± 0.17
5.44 ± 0.78
91.91 ± 0.16
3.16 ± 0.35
5.61 ± 0.44
92.83 ± 0.21
8.11
FL
12.97 ± 1.31
12.85 ± 0.93
80.94 ± 0.81
4.34 ± 0.24
4.91 ± 0.24
91.93 ± 0.10
4.49 ± 0.12
4.78 ± 0.22
93.63 ± 0.06
5.89
CRL
18.17 ± 0.62
18.04 ± 0.55
80.46 ± 1.19
6.98 ± 0.84
6.88 ± 0.86
91.83 ± 0.87
6.21 ± 0.22
5.99 ± 0.27
93.27 ± 0.38
10.67
Distillation
11.99 ± 0.76
12.26 ± 0.69
81.77 ± 0.50
3.79 ± 0.27
4.69 ± 0.45
93.00 ± 0.45
3.90 ± 0.13
5.03 ± 0.22
93.57 ± 0.12
4.00
Table 2: Performance comparison between different methods with Llama3-8B on ARC-C, ARC-E, and SciQ datasets. Results are reported as mean ± standard deviation (%). The best and second best results are shown in bold and underlined , respectively.
14 Method
ARC-C
SciQ
CSQA
Avg. Rank
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
Vanilla
18.25 ± 0.24
18.03 ± 0.12
79.81 ± 0.05
6.72 ± 0.39
6.64 ± 0.42
92.43 ± 0.35
25.42 ± 0.94
25.39 ± 0.99
72.07 ± 0.98
7.56
LS
18.69 ± 0.26
18.80 ± 0.43
76.17 ± 0.32
4.98 ± 0.82
7.65 ± 0.20
90.90 ± 0.90
21.82 ± 1.55
21.66 ± 1.60
72.24 ± 1.44
8.11
FL
16.26 ± 1.39
16.29 ± 1.33
76.99 ± 1.45
5.47 ± 1.19
6.16 ± 1.10
90.97 ± 1.32
19.49 ± 1.33
19.38 ± 1.33
72.89 ± 0.57
5.67
CRL
21.47 ± 1.05
21.41 ± 0.99
75.81 ± 0.91
8.07 ± 1.00
7.77 ± 1.18
91.20 ± 1.13
25.01 ± 0.26
25.01 ± 0.26
72.48 ± 0.12
10.00
Distillation
16.55 ± 0.78
16.53 ± 0.65
77.87 ± 1.04
5.17 ± 0.62
6.03 ± 0.78
92.00 ± 0.95
21.84 ± 0.49
21.78 ± 0.56
72.67 ± 0.58
5.78
Table 3: Calibration performance of various methods under distribution shift. Llama3-8B is fine-tuned on ARC-E and evaluated on ARC-C, SciQ, and CSQA. Results are reported as mean ± standard deviation (%). The best and second best results are shown in bold and underlined , respectively.
Figure 4: Calibration performance of various methods under distribution shift on CIFAR-10-C and CIFAR-100-C. The horizontal axis denotes the shift level from weakest to strongest and the vertical axis reports ECE (lower is better).
Table 4: OOD detection results (%). The best and second best results are shown in bold and underlined , respectively.
Although deep neural networks (DNNs) achieve high predictive accuracy, their confidence estimates are often unreliable, potentially compromising user trust in their decisions. This has motivated research on calibrated models, where calibration measures how well a model's predicted confidence aligns with the empirical probability of correctness. However, calibration metrics can often be improved through post-processing techniques that merely mimic training-time uncertainty without genuinely improving the model's understanding. For this reason, statisticians recommend that models be not only calibrated but also refined. Intuitively, a model is considered more refined if it assigns significantly different confidence scores to correct and incorrect predictions, a property also referred to as sharpness. We observe that many existing calibration methods improve calibration at the cost of reduced refinement. To address this limitation, we propose: (1) a novel loss function that explicitly promotes refinement and can be optimized through supervised contrastive learning; and (2) a unified training framework, RefCal, that jointly optimizes calibration, refinement, and accuracy to improve DNN reliability. On the CIFAR-100-LT dataset with 10 percent class imbalance, RefCal achieves (accuracy, refinement, ECE) of (58.81, 95.67, 0.08), substantially outperforming the widely used Correctness Ranking Loss, which achieves (46.27, 93.7, 0.22).
Modern deep learning models remain notoriously prone to overconfidence, limiting their reliability in high-stakes applications. Bayesian methods aim to counter this by learning a distribution over model parameters, and recent advances now make this feasible for large-scale architectures at costs comparable to AdamW. However, a challenge remains at test time: predictions must be averaged across many forward passes with weights sampled from the posterior, which is prohibitively expensive. Variance propagation offers an efficient alternative, computing layer-wise analytical approximations of uncertainty in a single forward pass. While such techniques are effective for MLPs, their extension to modern architectures remains challenging, due to increased depth and diversity of layer types. To fill this gap, we propose Calibrated Variance Propagation (CVP), which introduces a new propagation method for normalization layers, combines it with recent techniques for handling activation functions, and absorbs residual error through a light calibration step. CVP yields comparably accurate uncertainty estimates to MC sampling across transformers and CNNs, at a fraction of the cost. Against prior variance propagation work, CVP improves coverage at 0.5% risk from 8.2% to 14.6% with BEiT-3 on Visual Reasoning (NLVR2) and from 2.6% to 10.8% with ViLT on VQAv2, with gains extending to convolutional architectures.
Tobias Jan Wieczorek, Leon de Andrade, Thomas Möllenhoff +1
TU Darmstadt & hessian.AI, Darmstadt, Germany · RIKEN Center for Advanced Intelligence Project, Tokyo, Japan
Modern neural networks can achieve high accuracy while remaining poorly calibrated, producing confidence estimates that do not match empirical correctness. Yet calibration is often treated as a post-hoc attribute. We take a different perspective: we study calibration as a training-time phenomenon on small vision tasks, and ask whether calibrated solutions can be obtained reliably by intervening on the training procedure. We identify a tight coupling between calibration, curvature, and margins during training of deep networks under multiple gradient-based methods. Empirically, Expected Calibration Error (ECE) closely tracks curvature-based sharpness throughout optimization. Mathematically, we show that both ECE and Gauss--Newton curvature are controlled, up to problem-specific constants, by the same margin-dependent exponential tail functional along the trajectory. Guided by this mechanism, we introduce a margin-aware training objective that explicitly targets robust-margin tails and local smoothness, yielding improved out-of-sample calibration across optimizers without sacrificing accuracy.