In probabilistic classification, calibration error (CE) measures the average divergence of predicted probabilities f(X) from P(Y∣f(X)), the true class distribution for that predicted probability. While being a useful diagnostic tool, it is hard to estimate: popular binning-based estimators are often inconsistent and scale poorly beyond two classes. Recent work rewrites the CE as the excess risk of a model compared to the best recalibration of its own predictions, measured with a proper loss. However, this only works for Bregman-divergence-based calibration errors like the squared error, excluding the more popular L1-distance-based CE. We show that using prediction-dependent proper scores can alleviate this restriction, allowing us to estimate CEs with general convex divergences, including Lp distances with closed-form losses in the binary and multiclass settings. To estimate the excess risk, we introduce a more accurate recalibrator that fits a residual to temperature scaling with gradient boosting. The resulting variational estimator, V-ECE, needs no bins or clusters and lower-bounds the true calibration error in expectation. On a benchmark of semi-synthetic tasks built from real classifiers, with known true CE, V-ECE is among the most accurate binary estimators for every calibration error and significantly outperforms all multiclass estimators. Our results are accompanied by additional theory on Lp CE, estimator bias, and over- or under-confidence estimation.
Figures & tables
Figure 1: The L1 calibration error and how we estimate it , for an over-confident binary model. (a) The L1 CE is the orange area between E[Y∣f(X)] and the diagonal f(X) , weighted by the distribution of f(X) . (b) Re-expressing the L1 CE as a risk difference: writing Rℓf(g∘f)=E[ℓf(X)(g∘f(X),Y)] for the risk under the prediction-dependent score of Corollary 2 , the largest drop achievable over all recalibration functions is exactly the L1 CE. Any g^ therefore gives a lower bound.
Metric
Loss employed
CE estimated
Cross-entropy
−∑iYilog(zi)
E[KL(C∥f)]
Brier score
∥z−Y∥22
E[∥f−C∥22]
Ours ( L1 )
\mathds1z=f⟨\sign(z−f),f−Y⟩
E[∥f−C∥1]
Ours ( L2 )
\mathds1z=f⟨∥z−f∥2z−f,f−Y⟩
E[∥f−C∥2]
Ours (cvx Df )
−Df(z)−⟨Y−z,∇Df(z)⟩
E[Df(C)]
Table 1: Losses and the calibration error they estimate , writing f for f(X) . The estimate is the risk of f minus the risk of g^∘f under the loss of the second column. The first two rows are fixed proper losses and give Bregman divergences; the next two are prediction-dependent and give Lp distances, which no fixed proper loss can produce ( Proposition 5 ); and the last row works for any convex Df that satisfies the conditions of Theorem 1 .
Figure 2: Calibration error recovered by each recalibration function. For each classifier and calibration error, recalibration functions are ranked by their cross-validated estimate (rank 1 for the largest, i.e. the tightest lower bound), and ranks are averaged over classifiers. The vertical line is the average rank of a method no better than the others.
Figure 3: Accuracy of calibration error estimators on the semi-synthetic benchmark , where the true calibration error of each classifier is known. Normalized absolute error : on each classifier and label draw, the absolute error of an estimator divided by the largest absolute error among the estimators, averaged over draws and classifiers (lower is better; values depend on the pool of estimators, so they are not comparable between the binary (top) and multiclass (bottom) panels). % above truth : share of draws on which the estimate exceeds the true calibration error. Estimators are sorted by their normalized absolute error on the squared calibration error; Debiased-ECE only defines CEf,(⋅)2 , and Smooth-ECE, which only defines CEf,∣⋅∣ , is listed last. Exact values, NMAE and runtimes are in Tables F.2 and F.3 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1: The divergence of a proper loss as a gap under a supporting line. (a) For a fixed proper loss with entropy H , the divergence dℓ(f,C) is the vertical gap at z=C between the tangent to H at z=f and the graph of H . Where H is twice differentiable this gap closes quadratically, and for the Brier entropy H(z)=z(1−z) it is exactly (f−C)2 . (b) The prediction-dependent entropy Hf(z)=−∣z−f∣ has a kink at z=f . The supporting line with slope 0 is horizontal, and the gap at z=C is ∣f−C∣ . (c) The two divergences as functions of C : linear against quadratic near the prediction.
Figure A.2: What the estimator returns as a function of the value reported by g^ , at a fixed prediction f with C>f . (a) With the L1 score the value is a step: it equals the true ∣f−C∣ for every z on the same side of f as C , and −∣f−C∣ on the other side. (b) With the Brier score the value is a parabola peaked at z=C , so only the exact conditional probability attains (f−C)2 .
Algorithm 1 Computing CE with a prediction-dependent proper score.
Binary
Multiclass
Dataset
ncal
n
Dataset
ncal
n
k
2dplanes
36,691
4,077
artificial-characters
9,196
1,022
10
adult
43,957
4,885
connect-4
60,801
6,756
3
ailerons
12,375
1,375
Diabetes130US
91,589
10,177
3
Amazon_employee_access
29,492
3,277
dilbert
9,000
1,000
5
APSFailure
68,400
7,600
eye_movements
9,842
1,094
3
Appendix
Table E.1: TabRepo datasets used in the experiments , with the sizes of the calibration split ( ncal ) and of the test split ( n ), and the number of classes k of the multiclass datasets. Each dataset contributes eight classifiers, one per model family.
Estimator
Recalibration g^ / construction
Fitted
Functional
k=2
k>2
Uniform-ECE
15 bins, equal width
in-sample
binned distance
✓
Quantile-ECE
15 bins, equal mass
in-sample
binned distance
✓
Debiased-ECE
15 bins, equal width
in-sample
debiased binned, (⋅)2 only
✓
Smooth-ECE
Gaussian smoothing
in-sample
smoothed, ∣⋅∣ only
✓
Isotonic-ECE
isotonic regression
in-sample
plug-in distance
✓
KDE-ECE
Beta / Dirichlet kernel
leave-one-out
plug-in distance
✓
✓
Appendix
Table E.2: Estimators of the semi-synthetic benchmark , grouped by family: plug-in estimators (ECE), variational estimators with an in-sample g^ (IS), and variational estimators with an out-of-sample g^ , i.e. V-ECE with different recalibration functions. V-ECE (Default), in bold, is the configuration we recommend. Plug-in distance : n1∑id(f(Xi),g^(f(Xi))) . Binned distance : ∑bnnbd(fˉb,yˉb) , comparing the average prediction fˉb and the frequency yˉb of class 1 in each bin. Variational : the risk of f minus the risk of g^∘f , with the loss of Table 1 that corresponds to the calibration error. The last two columns indicate whether each estimator is run in the binary ( k=2 ) and multiclass ( k>2 ) benchmarks: the binned, isotonic, smoothing, quadratic-scaling and Beta-calibration estimators are only defined for binary predictions, and SMS is only used in the multiclass case.
Average rank ↓
Largest estimate (%) ↑
Time
Recalibrator
L1
L2
squared
L1
L2
squared
[ms/1k]
Binary
WS-CatBoost
3.15
–
2.82
19
–
31
1634
Beta
3.29
–
3.06
12
–
21
3.1
Quadratic
3.32
–
3.19
14
–
13
3.5
Isotonic
3.52
–
3.51
17
–
18
5.5
Uniform
3.67
–
4.02
23
–
11
0.44
Appendix
Table F.1: Recalibration functions inside the variational estimator ( Section 5.1 ), on 149 binary and 150 multiclass TabRepo classifiers with real labels. Average rank : rank of the estimate among the recalibration functions (1 for the largest), averaged over classifiers. Largest estimate : share of classifiers on which the recalibration function returns the largest estimate. L1 stands for CEf,∣⋅∣ (binary) and CEf,∥⋅∥1 (multiclass), L2 for CEf,∥⋅∥2 , and squared for CEf,(⋅)2 (binary) and CEf,∥⋅∥22 (multiclass). Time : wall-clock time of the 5-fold estimate in milliseconds per 1,000 samples. Best values in bold.
CE∣⋅∣
CE(⋅)2
Time
Estimator
NAE ↓
NMAE ↓
ratio
over
NAE ↓
NMAE ↓
ratio
over
[ms/1k]
Debiased-ECE
–
–
–
–
0.09 †
25
0.97
45
0.06
V-ECE (Beta)
0.20
22
0.84
24
0.12 †
38
0.48
16
3.4
V-ECE (Quadratic)
0.20
22
0.82
19
0.12 †
39
0.50
15
3.6
V-ECE (Default)
0.21
22
0.84
27
0.11
25
0.48
16
1619
V-ECE (Isotonic)
0.25
25
0.74
19
0.15
37
0.21
9
5.8
Appendix
Table F.2: Accuracy of the binary estimators on 176 semi-synthetic classifiers. NAE : normalized absolute error ( Figure 3 ). NMAE : ∑∣CE−CE∣/∑CE , in %. Ratio : median of CE/CE . Over : share of draws with CE>CE , in %. All statistics are computed on each of the three label draws of every classifier ( Section E.6 ). Time : milliseconds per 1,000 samples for one estimate. † : A dagger marks the estimators whose absolute errors are not significantly different from those of V-ECE (Default) (Wilcoxon signed-rank test, 5%); all other estimators are significantly less accurate than V-ECE (Default), except V-ECE (Beta), which is significantly more accurate for CEf,∣⋅∣ . Estimates are clipped at 0 after aggregating the folds, so a median ratio of 0 means that the estimate is clipped on most draws. Best values in bold.
CE∥⋅∥1
CE∥⋅∥2
CE∥⋅∥22
Time
Estimator
NAE ↓
NMAE ↓
ratio
over
NAE ↓
NMAE ↓
ratio
over
NAE ↓
NMAE ↓
ratio
over
[ms/1k]
V-ECE (Default)
0.27
26
0.70
2
0.26
26
0.70
2
0.15
44
0.34
0
2473
V-ECE (SMS)
0.32
37
0.64
2
0.31
37
0.65
2
0.17
58
0.25
2
102
V-ECE (TS)
0.36
44
0.63
2
0.36
44
0.61
2
0.20
63
0.12
1
12.8
V-ECE (KDE)
0.42
44
0.46
1
0.40
43
0.47
1
0.29
88
0.00
0
46.1
KDE-ECE
0.91
111
2.31
98
0.92
116
2.37
98
0.96
294
7.58
98
43.9
Appendix
Table F.3: Accuracy of the multiclass estimators on 168 semi-synthetic classifiers, with the columns of Table F.2 . V-ECE (Default) is significantly more accurate than every other estimator for the three calibration errors (Wilcoxon signed-rank test, 5%).
Figure F.1: Estimate divided by the true calibration error on each semi-synthetic classifier and label draw, against the true calibration error (log scale), for CEf,∣⋅∣ (binary, top) and CEf,∥⋅∥2 (multiclass, bottom). The horizontal line marks an exact estimate. Triangles are ratios outside the plotted range, drawn on its boundary.
Figure F.2: Calibration error estimators as tools for comparing models. Within each dataset, the eight models are ranked by their true and by their estimated calibration error. Spearman correlation : rank correlation between the two orderings, computed on each label draw and averaged over draws and datasets (error bars: one standard error over datasets). % pairs ordered : share of the 28 pairs of models ordered as by the truth, a tied pair counting as half (50% is chance).
CE∣⋅∣
CE(⋅)2
Estimator
Spearman ↑
pairs (%)
top-1 (%)
Spearman ↑
pairs (%)
top-1 (%)
Smooth-ECE
0.78 ± 0.06
84
45
–
–
–
Isotonic-ECE
0.79 ± 0.06
85
59
0.72 ± 0.05
80
50
Isotonic-IS
0.74 ± 0.07
81
27
0.70 ± 0.04
79
45
Uniform-ECE
0.79 ± 0.05
85
59
0.65 ± 0.05
77
32
V-ECE (Default)
0.66 ± 0.08
78
32
0.70 ± 0.06
79
38
Appendix
Table F.4: Ordering of the eight models of each binary dataset (22 datasets). Spearman : rank correlation between the estimated and the true ordering, computed on each label draw and averaged over draws and datasets, ± one standard error over datasets. Pairs : share of the 28 pairs of models ordered as by the truth, a tied pair counting as half. Top-1 : share of draws on which the model with the smallest estimate is the best-calibrated one, shared among tied models (chance: 12.5%). Best values in bold.
CE∥⋅∥1
CE∥⋅∥2
CE∥⋅∥22
Estimator
Spearman ↑
pairs (%)
top-1 (%)
Spearman ↑
pairs (%)
top-1 (%)
Spearman ↑
pairs (%)
top-1 (%)
V-ECE (Default)
0.89 ± 0.03
89
57
0.88 ± 0.02
89
52
0.83 ± 0.04
86
60
V-ECE (KDE)
0.82 ± 0.03
86
52
0.82 ± 0.04
86
45
0.56 ± 0.08
59
14
V-ECE (SMS)
0.76 ± 0.06
84
43
0.76 ± 0.06
83
48
0.66 ± 0.07
78
36
KDE-ECE
0.73 ± 0.05
81
43
0.68 ± 0.06
79
43
0.41 ± 0.10
67
14
V-ECE (TS)
0.65 ± 0.07
77
24
0.64 ± 0.07
77
19
0.51 ± 0.08
71
21
Appendix
Table F.5: Ordering of the eight models of each multiclass dataset (21 datasets), with the columns of Table F.4 .
Figure G.1: Separating over- and under-confidence. Top: the three simulated scenarios, showing C=E[Y∣f(X)] against the prediction f(X) (points) and the diagonal of perfect calibration (dashed). Predictions are over-confident (left), under-confident (middle), or a mix of both (right). Bottom: the corresponding V-ECE estimates, clipped at 0 and averaged over 10 independent runs.
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, United Arab Emirates · École pour l’informatique et les techniques avancées (EPITA), Paris, France