Do Not Train Away Uncertainty: Early Uncertainty Anchored Calibration
Authors: Yutong Xie, Jiawei Tang, Zhenglin Hua, Yuxiang Ma, Si Qin, Yaxin Hou, Hui Liu, Junhui Hou, +1 more
Organizations: School of Software Engineering, Southeast University, Nanjing, China · School of Computer Science and Engineering, Southeast University, Nanjing, China · School of Computing Information Sciences, Saint Francis University, Hong Kong, China · Department of Computer Science, City University of Hong Kong, Hong Kong, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China
Deep neural networks, including large language models, have achieved remarkable performance across various tasks. However, they are prone to overconfidence during training or fine-tuning. In this work, we observe a consistent phenomenon across different models that the early model is better calibrated, while later training or fine-tuning yields marginal accuracy gains but substantially increases calibration errors. Our analysis suggests that the early model retains uncertainty awareness in both its predictions and features, which is gradually lost with continued training. To avoid training away this uncertainty awareness, we propose \textbf{EUA-Cal}, a novel method that exploits the \textbf{E}arly model as an \textbf{U}ncertainty \textbf{A}nchor for \textbf{Cal}ibration. EUA-Cal introduces early prediction regularization to preserve early predictive uncertainty and prototype structure regularization to exploit uncertainty reflected in the early feature space, jointly mitigating overconfidence. Extensive experiments on image classification and multiple-choice question answering across eight diverse models demonstrate that EUA-Cal outperforms state-of-the-art calibration methods.
Figures & tables
Figure 1: (a) Training and validation classification errors and validation ECE for ResNet-50 on CIFAR-10, ViT-B/32 on CIFAR-100, and Llama3-8B on ARC-Challenge. (b) : t-SNE visualizations of features from ResNet-50 on CIFAR-10 at epochs 40, 60, and 200. Lighter shades indicate features associated with lower confidence. More examples are provided in Appendix B .
Figure 2: An illustration of the proposed EUA-Cal approach for calibration. In the pre-learning stage, we freeze the early model as an uncertainty anchor at the selected training epoch. Its predictions are used for EPR, while confidence-weighted class prototypes constructed from early features provide structural targets for PSR. The overall objective guides training or fine-tuning to preserve uncertainty and mitigate overconfidence.
Figure 3: (a) Confidence of certain and uncertain samples under the vanilla CE model and our method. Samples are divided according to the confidence of the selected early model. (b) Classification errors and ECE of the prototype predictions on the validation set using prototypes constructed from different training epochs.
15 Method
CIFAR-10
CIFAR-100
Tiny-ImageNet
Avg. Rank
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
Vanilla
3.36 ± 0.17
3.31 ± 0.19
94.57 ± 0.15
9.19 ± 0.16
9.07 ± 0.26
76.75 ± 0.62
5.54 ± 0.44
5.52 ± 0.43
64.45 ± 0.45
10.11
LS
3.74 ± 0.24
3.93 ± 0.33
94.78 ± 0.13
3.96 ± 0.22
3.99 ± 0.24
76.77 ± 0.13
3.78 ± 0.27
3.70 ± 0.36
63.58 ± 0.67
8.78
FL
1.81 ± 0.31
1.58 ± 0.27
94.40 ± 0.28
1.77 ± 0.31
1.71 ± 0.33
75.88 ± 0.35
1.57 ± 0.11
1.48 ± 0.18
62.94 ± 0.33
6.67
CRL
0.86 ± 0.09
0.61 ± 0.03
93.94 ± 0.22
5.81 ± 0.30
5.74 ± 0.27
76.87 ± 0.39
2.99 ± 0.31
3.04 ± 0.33
63.43 ± 0.33
7.67
Distillation
3.22 ± 0.52
3.27 ± 0.49
94.10 ± 0.67
7.51 ± 0.20
7.40 ± 0.16
76.64 ± 0.12
6.04 ± 0.24
6.01 ± 0.19
65.21 ± 0.20
10.22
Table 1: Performance comparison between different methods with ResNet-50 on CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets. Results are reported as mean ± standard deviation (%). The best and second best results are shown in bold and underlined , respectively.
14 Method
ARC-C
ARC-E
SciQ
Avg. Rank
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
Vanilla
17.24 ± 0.25
17.12 ± 0.31
81.31 ± 0.64
6.54 ± 0.03
6.48 ± 0.03
92.81 ± 0.12
6.42 ± 0.38
6.10 ± 0.41
93.40 ± 0.52
8.56
LS
15.23 ± 0.96
15.54 ± 1.26
79.81 ± 0.96
3.83 ± 0.17
5.44 ± 0.78
91.91 ± 0.16
3.16 ± 0.35
5.61 ± 0.44
92.83 ± 0.21
8.11
FL
12.97 ± 1.31
12.85 ± 0.93
80.94 ± 0.81
4.34 ± 0.24
4.91 ± 0.24
91.93 ± 0.10
4.49 ± 0.12
4.78 ± 0.22
93.63 ± 0.06
5.89
CRL
18.17 ± 0.62
18.04 ± 0.55
80.46 ± 1.19
6.98 ± 0.84
6.88 ± 0.86
91.83 ± 0.87
6.21 ± 0.22
5.99 ± 0.27
93.27 ± 0.38
10.67
Distillation
11.99 ± 0.76
12.26 ± 0.69
81.77 ± 0.50
3.79 ± 0.27
4.69 ± 0.45
93.00 ± 0.45
3.90 ± 0.13
5.03 ± 0.22
93.57 ± 0.12
4.00
Table 2: Performance comparison between different methods with Llama3-8B on ARC-C, ARC-E, and SciQ datasets. Results are reported as mean ± standard deviation (%). The best and second best results are shown in bold and underlined , respectively.
14 Method
ARC-C
SciQ
CSQA
Avg. Rank
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
ECE ↓
AECE ↓
ACC ↑
Vanilla
18.25 ± 0.24
18.03 ± 0.12
79.81 ± 0.05
6.72 ± 0.39
6.64 ± 0.42
92.43 ± 0.35
25.42 ± 0.94
25.39 ± 0.99
72.07 ± 0.98
7.56
LS
18.69 ± 0.26
18.80 ± 0.43
76.17 ± 0.32
4.98 ± 0.82
7.65 ± 0.20
90.90 ± 0.90
21.82 ± 1.55
21.66 ± 1.60
72.24 ± 1.44
8.11
FL
16.26 ± 1.39
16.29 ± 1.33
76.99 ± 1.45
5.47 ± 1.19
6.16 ± 1.10
90.97 ± 1.32
19.49 ± 1.33
19.38 ± 1.33
72.89 ± 0.57
5.67
CRL
21.47 ± 1.05
21.41 ± 0.99
75.81 ± 0.91
8.07 ± 1.00
7.77 ± 1.18
91.20 ± 1.13
25.01 ± 0.26
25.01 ± 0.26
72.48 ± 0.12
10.00
Distillation
16.55 ± 0.78
16.53 ± 0.65
77.87 ± 1.04
5.17 ± 0.62
6.03 ± 0.78
92.00 ± 0.95
21.84 ± 0.49
21.78 ± 0.56
72.67 ± 0.58
5.78
Table 3: Calibration performance of various methods under distribution shift. Llama3-8B is fine-tuned on ARC-E and evaluated on ARC-C, SciQ, and CSQA. Results are reported as mean ± standard deviation (%). The best and second best results are shown in bold and underlined , respectively.
Figure 4: Calibration performance of various methods under distribution shift on CIFAR-10-C and CIFAR-100-C. The horizontal axis denotes the shift level from weakest to strongest and the vertical axis reports ECE (lower is better).
Table 4: OOD detection results (%). The best and second best results are shown in bold and underlined , respectively.