A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliability relation under covariate shifts that preserve the distribution of the confidence score, constraining the reweighting within each confidence level by a χ2 budget; the resulting worst case, as a function of the budget, is a fragility profile. On an interval of budgets that can be computed from the source distribution, the profile equals exactly the square root of the budget times the within-level variance of the correctness propensity -- the grouping-loss term of calibration-refinement decompositions. Beyond this interval the profile is governed by the tails of the propensity law, and the entire upward profile determines the centred within-level law; consequently, calibration residual and grouping variance do not determine fragility in general, though they do when labels and predictions are deterministic. Since the propensity is not observed, we restrict reweightings to a learned finite readout within confidence bins, bound the part the restriction misses by the grouping variance remaining inside readout cells, estimate the restricted profile with role-separated labels, and provide a separate split-sample lower confidence bound. On ImageNet this bound is positive in both splits for four of six primary classifiers and nine of twelve additional ones as released, and for three of eighteen after temperature scaling. Held-out drift under optimised reweightings fitted without evaluation labels tracks the estimated profile; an exploratory label-permutation diagnostic yields near-zero agreement for this statistic while largely reproducing the correlation observed for unsigned random reweightings.
Figures & tables
Figure 1: Reliability can move while the confidence distribution does not. (a) Within a confidence level set, mass moves between inputs that share a score but differ in their propensity to be correct; E[r∣S=s]=1 keeps the level’s mass fixed. (b) The score distribution, and so any confidence histogram, is unchanged. (c) The reliability curve moves by at most the fragility profile; a threshold policy keeps its coverage while its error rate rises. Movement is possible exactly where GP>0 (Proposition 2 ). The schematic is the population construction; the finite-sample one of Section 4 preserves the binned marginal.
Figure 2: The profile beyond its variance-controlled regime. Left: the three regimes of Theorem 1 for η∈{0.2,0.6,1.0} with probabilities (41,21,41) ( ρ+=0.5 , ρsat=3 ). Right: that law and a four-point law on {0,0.4,0.7,0.9} with the same mean and variance ( 0.6 , 0.08 ) agree for ρ≤0.22 and separate by 0.119 at ρ=3 . The two-point witnesses of Proposition 4 (iv) have no tail-controlled regime and separate by 0.100 once both saturate (Appendix M ).
T=1
T=T^
Model
Acc.
ECE
GR
FAUC
ECE
GR
FAUC
ResNet-50 V1
0.761
0.038
0.0004
0.042
0.024
0.0009
0.049
ResNet-50 V2
0.809
0.412
0.0133
0.135
0.032
0.0009
0.046
ConvNeXt-T
0.825
0.169
0.0081
0.112
0.031
0.0011
0.048
ViT-B/16
0.811
0.057
0.0011
0.048
0.042
0.0005
0.042
Swin-T
0.815
0.068
0.0015
0.056
0.032
0.0004
0.038
Table 1: Primary cohort. ECE: 20 -bin equal-mass out-of-fold estimate; GR with the binomial-noise correction of Proposition 9 ; FAUC of the upward plug-in profile. Neither GR nor FAUC is a confidence bound. Additional checkpoints: Table 12 .
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
bins used
target mass kept
mean Lcomp/Ltotal
Lcond>Lcomp
severity 1
severity 5
all
100%
41.8%
400/450
50.6%
33.0%
unflagged
77.9% ( 55.8% – 98.6% )
43.9%
401/450
52.7%
37.0%
Appendix
Table 2: Low-target-count sensitivity of the ImageNet-C summaries (post-hoc; original partition, fold-pooled records). “Unflagged bins” drops, within each (model, domain), the bins carrying the stored low-count flag of the pooled table (set by the preregistered per-fold rule of fewer than 200 target examples and taken from the first available fold; not a threshold on the fold-averaged count) and renormalises the target bin mass over the rest; every one of the 450 (model, domain) units keeps at least one bin. This changes the estimand to higher-count bins; it is not a revised-pipeline rerun and carries no interval or coverage statement.
Figure 3: Source calibration error against restricted fragility, exploratory. One point per checkpoint and calibration state, all post-audit and regenerated from the same source as Table 1 ; the six checkpoints appear four times each and are not independent observations. Marginally the two are associated (Pearson 0.91, Spearman 0.88); conditioning on GR removes it (partial 0.00, while GR retains 0.86 given ECE). Appendix J gives both parent sets, the rank versions and the leave-one-checkpoint-out ranges, and why none of this establishes non-determination.
Figure 4: The composition-to-total norm ratio decreases with corruption severity. Original protocol, fixed partition; not recomputed. ImageNet-C, macro-averaged over fifteen corruptions and six architectures. The conditional term tracks the total; the governed composition term grows far more slowly, and Lcomp/Ltotal falls from 50.6% at severity 1 to 33.0% at 5. Norms are not additive, so this is a ratio and not a share.
Target
Ltotal
Lcomp
Lcond
median U
ImageNet-C (75 domains)
0.103
0.034
0.102
0.63
ImageNetV2
0.058
0.006
0.055
0.37
ImageNet-Sketch
0.240
0.024
0.223
0.89
Appendix
Table 3: Original protocol, fixed partition; not recomputed under the revised data roles. Reliability drift under natural shift, averaged over six checkpoints and domains, with L and U as defined in Section 6 . The independent-rate check of Appendix P tests the qualitative dominance only; it does not regenerate these coefficients. The L are quadratic means of binwise drift, so Ltotal=Lcomp+Lcond despite the binwise identity being exact.
changed assignments / violations
max absolute difference
bins
frozen R
rank
ties
cells
GR
F±
FAUC +
FAUC -
sel. risk
ResNet-50 V1
0
0
0
0
0
0
0
0
0
0
ResNet-50 V2
0
0
0
0
0
0
0
0
0
0
ConvNeXt-T
0
0
0
0
0
0
0
0
0
0
ViT-B/16
0
0
0
0
0
0
0
0
0
0
Swin-T
0
0
0
0
0
0
0
0
0
0
Appendix
Table 4: Structural invariance under the fold-wise monotone map, all six architectures. Counts are over all 250,000 train and evaluation bin assignments per model. Every entry is exactly zero; the preregistered tolerance was 10−12 .
ECE
binary NLL
ECE 15
T=1
monotone
reduction
T=1
monotone
T=1
monotone
ResNet-50 V1
0.03614
0.01288
64.4%
0.39579
0.37929
0.03688
0.01349
ResNet-50 V2
0.41178
0.02162
94.8%
0.80867
0.38869
0.41178
0.02064
ConvNeXt-T
0.16933
0.01477
91.3%
0.44668
0.35481
0.16933
0.01312
ViT-B/16
0.05589
0.02577
53.9%
0.36461
0.34621
0.05605
0.02328
Swin-T
0.06794
0.01989
70.7%
0.36261
0.34134
0.06794
0.01740
Appendix
Table 5: Cross-fitted calibration under the fold-wise monotone map. ECE is the 20 -bin equal-mass estimator of Appendix N with the bin partition held fixed; ECE 15 is the equal-width secondary metric. All six architectures exceed the preregistered 20% relative-reduction bar.
af range
cf range
Spearman
Kendall
pooled ties
edge gap
before
after
ResNet-50 V1
0.710 – 0.719
−0.128 – −0.091
1.000
1.000
29
4
2.2×10−9
ResNet-50 V2
1.430 – 1.457
+2.554 – +2.570
1.000
1.000
29
4
2.2×10−9
ConvNeXt-T
1.342 – 1.352
+1.026 – +1.043
1.000
1.000
29
4
2.3×10−9
ViT-B/16
1.343 – 1.360
+0.124 – +0.155
1.000
1.000
29
4
0.7×10−9
Swin-T
1.251 – 1.264
+0.351 – +0.377
1.000
1.000
29
4
1.9×10−9
Appendix
Table 6: Fitted parameters and rank diagnostics. Spearman and Kendall are within-fold , where a single gf acts; no optimisation reached a numerical boundary in any of the 30 fold fits. The edge gap is maxj∣g(edgej)−edgejrefit∣ , the distance between the pushed-forward boundaries the primary analysis uses and independently refitted transformed-score quantiles.
gflat
gsharp
ECE
bins chg.
merged
max∣ΔF∣
ECE
bins chg.
merged
max∣ΔF∣
ResNet-50 V1
0.04713
0
0
0
0.09074
0
146
0
ResNet-50 V2
0.36709
0
0
0
0.46333
0
0
0
ConvNeXt-T
0.23640
0
0
0
0.10666
0
0
0
ViT-B/16
0.16995
0
0
0
0.05472
0
0
0
Swin-T
0.16788
0
0
0
0.05107
0
0
0
Appendix
Table 7: Deterministic stress maps. Large deliberate changes in numerical confidence leave every structural quantity exactly unchanged.
ARI (R)
exact R agreement
relative ∣ΔGR∣
relative ∣Δ FAUC +∣
ResNet-50 V1
0.136
0.379
108.7%
27.2%
ResNet-50 V2
0.615
0.795
1.9%
0.5%
ConvNeXt-T
0.584
0.775
9.9%
4.5%
ViT-B/16
0.611
0.796
25.4%
11.3%
Swin-T
0.575
0.770
27.5%
13.4%
EffNet-B0
0.672
0.832
13.6%
7.2%
Appendix
Table 8: Secondary end-to-end analysis, in which the bin quantiles and the low-capacity logistic readout are refitted on the transformed confidence coordinate. Diagnostic only.
Figure 5: Random score-preserving shifts, post-audit. (a) One point per bin–radius cell: two-sided plug-in profile against the median held-out movement over 100 random directions (Spearman 0.901 , most of which a within-bin label-permutation null reproduces; Appendix T ); dashed line y=x , equal axes. (b) Per-shift ratio ∣δrandom∣/F by quartile of the bin’s GR : median and 90th percentile (bars, left axis) and the percentage of shifts with ratio above one (line, right axis). Descriptive; the shifts are not independent.
stratum
shifts
median ratio
90th pct.
ratio >1 (%)
all
720,000
0.264
0.741
3.86
GR Q1 [−0.00148,0.00012]
180,000
0.352
0.994
9.84
GR Q2 [0.00012,0.00066]
180,000
0.270
0.742
3.43
GR Q3 [0.00066,0.00235]
180,000
0.232
0.638
1.54
GR Q4 [0.00235,0.03797]
180,000
0.226
0.595
0.63
T=1
360,000
0.257
0.708
3.08
Appendix
Table 9: Random score-preserving shifts: the per-shift ratio ∣δrandom∣/FB,R± , with the branch chosen by the sign of δrandom . Shifts within a bin–radius cell share data and are not independent; the percentages are descriptive and are not coverage rates. Post-audit, not preregistered.
human proxy
cross-target (pre-specified)
same proxy (post-hoc)
model
GH×
vote noise
GR
ratio
GRH
ratio
≤GH×
ResNet-56
9.29
9.3%
0.021
0.33%
0.521
6.3% [3.8, 10.1]
50/50
VGG16-BN
8.33
9.9%
0.012
0.17%
0.093
1.7% [0.2, 3.5]
50/50
MobileNetV2
9.23
8.3%
0.035
0.29%
0.251
2.8% [1.3, 4.3]
50/50
pooled
0.25% [0.0, 2.1]
3.3% [1.4, 6.3]
150/150
Appendix
Table 10: CIFAR-10H case study, J=10 bins and K=4 groups in each of 5 folds, 50 fold–bin cells per model. GH× is the cross-half covariance of the two annotator-half proxies; GH× , GR and GRH are medians over cells, in units of 10−3 . Columns 4–5: the pre-specified cross-target diagnostic, whose restricted term uses the fixed CIFAR-10 label and so a different propensity from GH× ; Proposition 5 does not order the two. Columns 6–8: the post-hoc same-proxy diagnostic, in which both terms use ηH . Ratios are medians over cells with interquartile ranges in brackets. Top-1 accuracy on the CIFAR-10 test set: ResNet-56 94.37%, VGG16-BN 94.15%, MobileNetV2 94.05%.
Figure 6: CIFAR-10H case study, post-audit. One point per (model, fold, bin) cell, 150 in total; dashed line y=x . (a) Human-proxy heterogeneity GH from annotator half A against half B. (b) Post-hoc same-proxy diagnostic: the vote-noise-corrected cross-half GH× against the debiased restricted GRH computed from the same proxy.
benchmark
partition, state
units
Lcomp/Ltotal
range
Lcond>Lcomp
median U
ImageNetV2
original, T=1
30
12.2%
[2.1, 29.7]
100%
0.399
revised, T=1
30
12.0%
[2.4, 31.5]
100%
0.357
revised, T=T^
30
6.3%
[3.0, 12.4]
100%
0.331
ImageNet-Sketch
original, T=1
30
9.8%
[3.7, 24.7]
100%
0.891
revised, T=1
30
10.0%
[2.8, 30.0]
100%
0.768
revised, T=T^
30
7.8%
[1.6, 22.1]
100%
0.661
Appendix
Table 11: Natural shift under the original and the revised partition, with matched aggregation: per (checkpoint, source fold) unit, then the mean over the 30 units. The primary comparison is between the first two rows of each block, both in the as-released state; the temperature-scaled row has no original-partition counterpart and is an additional result, not a partition sensitivity. Units share target images and source data and are not independent. Post-audit, not preregistered.
T=1
T=T^
Model
family
Acc.
ECE
GR
FAUC
ECE
GR
FAUC
DenseNet-201
DenseNet
0.773
0.032
0.0006
0.043
0.021
0.0009
0.050
ResNeXt-50 32x4d
ResNeXt
0.776
0.065
0.0009
0.049
0.028
0.0015
0.057
RegNetX-4.0GF
RegNet
0.785
0.052
0.0008
0.047
0.025
0.0017
0.061
MobileNetV3-L
MobileNetV3
0.758
0.067
0.0029
0.076
0.021
0.0007
0.047
DeiT-S/16
DeiT
0.799
0.081
0.0010
0.049
0.041
0.0002
0.037
Appendix
Table 12: Post-audit architecture-breadth cohort: twelve additional frozen ImageNet-1K checkpoints, one per family, under the revised data roles of Appendix N (same estimators as Table 1 ). Acc. is the measured top-1 on the 50,000 validation images; every checkpoint reproduces its public reference within 0.04 pp. Not preregistered; the six checkpoints of Table 1 remain the frozen primary cohort.
cohort
(ECE, GR )
(ECE, FAUC)
( GR , FAUC)
primary six
+1.00 / -0.37
+1.00 / -0.83
+1.00 / +0.60
additional twelve
+0.61 / -0.28
+0.81 / -0.36
+0.92 / +0.99
all eighteen
+0.75 / -0.26
+0.85 / -0.36
+0.95 / +0.95
Appendix
Table 13: Cohort summaries. Top: Spearman rank correlations among the three source metrics across checkpoints (one point per checkpoint), as released / temperature-scaled. Bottom: held-out controlled-shift rank agreement over (checkpoint, state) combinations with the aggregation of Appendix N , and random score-preserving shifts (Appendix R.1 ; “corr.” is the Spearman correlation over bin–radius cells between the two-sided plug-in profile and the median, or secondarily the maximum, of 100 held-out movements; most of it is reproduced by the permutation null of Appendix T ). Primary-six values in this table use the revised estimator; Appendix P reports the original-protocol values. All rank correlations are computed from the unrounded source values. Descriptive; checkpoints and cells are not independent samples and no interval is attached.
as released T=1
split-fitted T=T^
monotone map
Model
A fits/B cert.
B fits/A cert.
A fits/B cert.
B fits/A cert.
ECE reduction
lower NLL
DenseNet-201
0
0
0
0
57.4%
yes
ResNeXt-50 32x4d
0.00068
0
0.00018
0.00058
not fitted
–
RegNetX-4.0GF
0
0
0.00053
0.00037
not fitted
–
MobileNetV3-L
0.0011
0.0011
0
0.00034
85.2%
yes
DeiT-S/16
0.00011
0.00064
0
0
70.0%
yes
Appendix
Table 14: Additional twelve checkpoints: split-sample lower confidence bound LCBmodel+(0.3) of Proposition 7 , one column per split arm (construction and split seed of Appendix N.1 unchanged; each entry carries its own per-analysis 95% statement, never joint), and the fold-wise monotone reparameterisation of Appendix Q replicated post-audit (relative reduction of cross-fitted ECE; structural quantities unchanged exactly). The original 4-of-6 outcome and Outcome A are historical and are not redefined.
Figure 7: Source metrics across eighteen checkpoints, as released (top) and temperature-scaled (bottom). One point per checkpoint (blue: primary six; red: additional twelve; marker = family), log axes. Panel titles give the Spearman rank correlation on the six / the twelve / all eighteen, computed from unrounded values. Descriptive; the checkpoints are not independent samples.
Figure 8: Held-out rank agreement against restricted heterogeneity. All eighteen checkpoints in four calibration states (blue: primary six; red: breadth twelve; marker = family). The weakest agreement occurs where GR is smallest.
Figure 9: Random score-preserving shifts on the breadth and combined cohorts. One point per bin–radius cell: two-sided plug-in profile against the median held-out movement over 100 random directions; dashed line y=x , equal axes. Descriptive; most of the association is reproduced by the permutation null of Appendix T .
(ECE, FAUC)
( GR , FAUC)
(ECE, GR )
primary six
+1.00 / − 0.83
+1.00 / +0.60
+1.00 / − 0.37
additional twelve
+0.81 / − 0.36
+0.92 / +0.99
+0.61 / − 0.28
all eighteen
+0.85 / − 0.36
+0.95 / +0.95
+0.75 / − 0.26
leave one family out
[+0.79, +0.90] / [ − 0.50, − 0.28]
[+0.93, +0.98] / [+0.94, +0.96]
[+0.65, +0.90] / [ − 0.42, − 0.16]
one per family
+0.82 / − 0.35
+0.94 / +0.96
+0.71 / − 0.28
Appendix
Table 15: Rank correlations among source metrics across checkpoints (Spearman, unrounded inputs), as released / temperature-scaled. Leave-one-family-out gives the range over the 17 deletions; one-per-family keeps the first checkpoint of each family in the frozen cohort order. Post-hoc and descriptive; checkpoints are not independent samples.
state
pair
∣ΔECE∣
FAUC
FAUC ratio
GR
T=T^
ResNet-50 V2 / PoolFormer-S24
1.41×10−4
0.0461 / 0.0770
1.67
0.00086 / 0.00459
T=T^
ResNet-50 V2 / MViTv2-T
4.54×10−4
0.0461 / 0.0407
1.13
0.00086 / 0.00061
T=T^
Swin-T / MViTv2-T
0.77×10−4
0.0384 / 0.0407
1.06
0.00038 / 0.00061
T=T^
DenseNet-201 / MobileNetV3-L
3.72×10−4
0.0495 / 0.0471
1.05
0.00094 / 0.00072
T=T^
RegNetX-4.0GF / ResMLP-12
1.09×10−4
0.0610 / 0.0384
1.59
0.00166 / 0.00022
T=1
ResNet-50 V1 / PoolFormer-S24
4.98×10−4
0.0423 / 0.0723
1.71
0.00041 / 0.00417
Appendix
Table 16: All checkpoint pairs among the eighteen whose estimated ECE agrees within the original tolerance 5×10−4 (all 153 pairs per state enumerated; none within the primary six in either state). Post-hoc enumeration on an enlarged cohort; similar estimated ECE is not equal population calibration error, and no significance is claimed.
cohort
null variant
pooled
fixed radius
within unit
primary six
both roles
0.901 / 0.876 [0.886]
0.828 / 0.747 [0.771]
0.776 / 0.748 [0.777]
Drate only
0.901 / 0.835 [0.843]
0.828 / 0.692 [0.714]
0.776 / 0.755 [0.776]
Deval only
0.901 / 0.846 [0.854]
0.828 / 0.659 [0.681]
0.776 / 0.743 [0.765]
additional twelve
both roles
0.901 / 0.877 [0.881]
0.838 / 0.755 [0.768]
0.803 / 0.755 [0.774]
Drate only
0.901 / 0.851 [0.856]
0.838 / 0.726 [0.741]
0.803 / 0.759 [0.775]
Deval only
0.901 / 0.866 [0.872]
0.838 / 0.711 [0.722]
0.803 / 0.764 [0.776]
Appendix
Table 17: Random-shift ordering statistic against the within-bin label-permutation null ( M=100 replicates per variant; correctness permuted within each bin in Drate and/or Deval , with partitions, cell counts, proportions and directions fixed). Entries: observed / null mean [null maximum]. “Fixed radius”: for the observed data and for each replicate, the mean over the six radii of the Spearman correlation over cells at that radius; the null mean and maximum are taken over replicates of that mean. (The r7 version of this table printed, as the bracket, the maximum over radii and replicates of the radius-specific correlations, a looser quantity; corrected here.) “Within unit”: median over (checkpoint, state, fold, radius) units of the Spearman correlation across the 20 bins. Permutations are fold-conditional, so the null does not reproduce cross-fold dependence; no p -value is reported. Post-hoc.
ρ
0.01
0.03
0.1
0.3
1
3
outside variance regime
0%
0%
0%
45%
99%
100%
excess of ρG^
0%
0%
0%
0%
10%
44%
excess of capped variance
0%
0%
0%
0%
9%
10%
matched-variance difference
0%
0%
0%
1%
10%
19%
pairs differing by >10%
0%
0%
0%
6%
48%
72%
Appendix
Table 18: The plug-in profile against its variance summaries on identical cells (all eighteen checkpoints, as released and T=T^ , 3,600 bin units). Regime: share of units outside the variance-controlled regime (upward profile). Excess: median relative overstatement of F+ by ρG^ and by min{ρG^,z^+} . Matched variance: median relative profile difference within disjoint consecutive pairs of units whose G^ agree within 1% ( 1,742 pairs), and the share of pairs differing by more than 10% . Post-hoc; plug-in quantities, no coverage.
n=250
n=5000
law of η
F+(3)
cap slack
full
capped
full
capped
four-point (B)
0.299
0.101
0.022
0.148
0.006
0.102
mixture λ=21
0.362
0.038
0.029
0.052
0.006
0.038
three-point (A)
0.400
0.000
0.030
0.030
0.006
0.006
B, variance /16
0.075
0.025
0.042
0.052
0.008
0.027
Appendix
Table 19: Known-propensity check, primary arm: one bin, upward, ρ=3 . Exact profile, population slack of the capped variance, and mean absolute error over 2000 replicates of the full plug-in and the capped variance from the same rates.
n=250
n=1000
n=5000
law, direction
ρ
F
MAE f/c
Δ
MAE f/c
Δ
MAE f/c
Δ
B
0.3
0.155
0.011/0.011
+0.000
0.006/0.006
+0.000
0.003/0.003
+0.000
1
0.283
0.020/0.020
+0.000
0.010/0.010
+0.000
0.005/0.005
+0.000
3
0.299
0.022/0.148
−0.126
0.014/0.110
−0.096
0.006/0.102
−0.095
λ=41
0.3
0.155
0.011/0.011
+0.000
0.005/0.005
+0.000
0.002/0.002
+0.000
1
0.279
0.020/0.021
+0.000
0.010/0.011
−0.001
0.004/0.005
−0.001
Appendix
Table 20: Known-propensity magnitude, primary arm (proportional stratified labels, 2000 replicates). F : exact profile. MAE f/c: mean absolute error of the full plug-in and of the capped variance; Δ : mean paired difference ∣full−F∣−∣capped−F∣ (negative favours the full plug-in; Monte-Carlo standard errors are at most 0.002 ). Upward unless stated.
condition
sampling
bias f/c, n=250
MAE f/c, n=250
n=1000
n=5000
B, up, ρ=3
proportional
+0.010 / +0.141
0.022/0.148
0.014/0.110
0.006/0.102
equal
+0.000 / +0.096
0.025/0.097
0.013/0.101
0.006/0.101
iid, drop empty
+0.006 / +0.120
0.023/0.130
0.014/0.106
0.006/0.102
iid, drop <30
−0.017 / −0.017
0.025/0.025
0.020/0.020
0.006/0.102
asym., down, ρ=1
proportional
+0.003 / +0.004
0.020/0.020
0.010/0.010
0.004/0.004
equal
+0.004 / +0.004
0.012/0.013
0.006/0.006
0.003/0.003
Appendix
Table 21: Sampling arms at two conditions: bias and MAE of the full plug-in (f) and the capped variance (c).
ρ=1
ρ=3
law, ordering
K
GR
HR
gap
B −FR
A −FR
gap
B −FR
A −FR
smooth, feature
1
0.0000
0.0469
0.212
0.217
0.217
0.296
0.375
0.375
2
0.0385
0.0085
0.016
0.020
0.092
0.088
0.126
0.159
4
0.0417
0.0053
0.010
0.014
0.072
0.030
0.057
0.126
8
0.0440
0.0029
0.006
0.008
0.054
0.016
0.036
0.094
16
0.0458
0.0012
0.003
0.003
0.034
0.003
0.013
0.059
Appendix
Table 22: Nested readouts, exact population values (upward). Gap: FB−FB,R ; B and A: the middle and right-hand bounds of Proposition 6 , reported as excess over FB,R .
law
n
K=1
K=2
K=4
K=8
K=16
smooth
250
0.000/0.296
0.030/0.094
0.039/0.048
0.040/0.037
0.045/0.043
1000
0.000/0.296
0.015/0.089
0.020/0.036
0.020/0.023
0.020/0.019
5000
0.000/0.296
0.007/0.088
0.009/0.031
0.009/0.018
0.008/0.008
rare
250
0.000/0.264
0.038/0.063
0.040/0.046
0.044/0.041
0.044/0.044
1000
0.000/0.264
0.019/0.053
0.021/0.037
0.021/0.020
0.021/0.021
5000
0.000/0.264
0.008/0.051
0.010/0.033
0.010/0.012
0.010/0.010
Appendix
Table 23: Nested readouts at a fixed label budget, feature ordering, upward, ρ=3 : RMSE of the plug-in restricted profile against its own target / against the binned oracle ( 1000 replicates).
state
FAUC 4 /FAUC 8
FAUC 2 /FAUC 8
GRdeb : K=4 / K=8
ΔGdeb/ΔGplug , 4→8
as released
0.89 [0.76, 0.94]
0.65 [0.56, 0.75]
0.89 [0.85, 1.05]
0.35 [-0.05, 0.79]
temperature-scaled
0.80 [0.73, 0.89]
0.62 [0.48, 0.72]
0.89 [0.83, 1.08]
0.14 [-0.04, 0.61]
Appendix
Table 24: Nested merges on ImageNet, eighteen checkpoints, median [range]. The last column is the share of the plug-in increase in GR from four to eight groups that remains after the binomial-noise correction.
Confidence calibration for classification models is vital in safety-critical decision-making scenarios and has received extensive attention. General confidence calibration methods assume training and test data are independent and identically distributed, limiting their effectiveness under covariate shifts. Previous calibration methods under covariate shift struggle with class-wise or canonical calibrations and often rely on unstable importance weighting when density ratios are large or unbounded. Given the above limitations, this paper rethinks confidence calibration under covariate shifts. First, we derive a necessary and sufficient condition for confidence calibration under covariate shifts, named Expectation consistency condition, which reveals covariate shifts do not necessarily lead to uncalibrated confidence and provides a weaker condition for confidence calibration than global covariate distribution alignment. Then, utilizing Expectation consistency condition, this paper proposes an unsupervised domain adaptation loss to calibrate confidence of the target domain, named Expectation consistency loss (ECL), which is compatible with canonical calibration, class-wise calibration, and top-label calibration. Third, we prove that computing ECL loss has the same sample complexity as Expected Calibration Error (ECE) and provide a theoretically grounded mini-batch trainable scheme for ECL loss. Finally, we validate the effectiveness of our method on both simulated and real-world covariate shift datasets.
Jinzong Dong, Zhaohui Jiang, Bo Yang
School of Automation, Central South University, Changsha, China.
Calibrated probability outputs of trained classifiers are increasingly used as inputs to downstream regression estimands such as effects, prevalences, or disparities for a latent group observed only on a small labelled subset. A standard practice is to threshold the calibrated score at a confidence cutoff and treat the hard label as the truth. Building on a recent identification result for the underlying moment equation, we develop a calibration-aware diagnostic apparatus for pseudo-labelling pipelines. We derive a closed-form expression for the attenuation bias that confidence thresholding induces in the downstream regression coefficient, and show that the bias can be predicted, before any inference is run, from the residual score variance V∗=E[Var(p∣X)] on the unlabelled set after partialling out the downstream controls X. We further obtain a sharp sensitivity bound under bounded calibration drift, and identify the boundary V∗=0, which holds iff p is a deterministic function of X; this motivates a structural separation between classifier features W and downstream controls X⊊W. Five controlled simulations and a UCI Adult illustration trace the predictions. The contribution is operational: a (V∗,κ) decision rule that practitioners can compute from any classifier output to decide whether confidence thresholding is safe.
Marcell T. Kurbucz
Institute for Global Prosperity, The Bartlett, University College London, 9–11 Endsleigh Gardens, London, WC1H 0EH, United Kingdom
Prediction sets can make deployed classifiers safer by returning several plausible labels when a single prediction is uncertain. Their value depends on classwise reliability: average coverage can meet its target while rare or difficult classes fail repeatedly. This concern is sharper after distribution shift, when calibration labels come from a source environment but reliability is needed on the target. We ask what labeled source data and unlabeled target inputs reveal about class-conditional prediction sets, and when target labels are necessary. Under unrestricted joint shift, two target laws can produce the same observable data while requiring different classwise thresholds; any label-free rule covering both must enlarge its sets on one law. We give a labeled target audit that estimates the missing quantiles with a simultaneous guarantee. Probability-scale error is invariant to increasing score transformations, and threshold recovery follows under local regularity. At fixed confidence, achieving threshold tolerance ε with fixed, nonadaptive class-stratified labeled pairs has total complexity Θ(Kε−2logK), or Θ(ε−2logK) labels per class under equal allocation. Class imbalance creates a separate acquisition cost; for foreground class probabilities of order 1/K, the mixed-stream label complexity is also Θ(Kε−2logK) at fixed confidence. Experiments on action-recognition and image shifts show that marginal coverage can conceal severe class failures and that source classwise calibration depends on the shift. The results connect the information available at deployment to the target labels needed for useful class-conditional prediction.