cs.AISep 30, 2026

How Much Can Reliability Drift Under a Fixed Confidence Distribution?

Authors: Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen

Organizations: Adelaide University

Abstract

A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliability relation under covariate shifts that preserve the distribution of the confidence score, constraining the reweighting within each confidence level by a χ2χ^2 budget; the resulting worst case, as a function of the budget, is a fragility profile. On an interval of budgets that can be computed from the source distribution, the profile equals exactly the square root of the budget times the within-level variance of the correctness propensity -- the grouping-loss term of calibration-refinement decompositions. Beyond this interval the profile is governed by the tails of the propensity law, and the entire upward profile determines the centred within-level law; consequently, calibration residual and grouping variance do not determine fragility in general, though they do when labels and predictions are deterministic. Since the propensity is not observed, we restrict reweightings to a learned finite readout within confidence bins, bound the part the restriction misses by the grouping variance remaining inside readout cells, estimate the restricted profile with role-separated labels, and provide a separate split-sample lower confidence bound. On ImageNet this bound is positive in both splits for four of six primary classifiers and nine of twelve additional ones as released, and for three of eighteen after temperature scaling. Held-out drift under optimised reweightings fitted without evaluation labels tracks the estimated profile; an exploratory label-permutation diagnostic yields near-zero agreement for this statistic while largely reproducing the correlation observed for unsigned random reweightings.

Figures & tables

Appendix figures & tables30 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Expectation Consistency Loss: Rethink Confidence Calibration under Covariate Shift

    May 20, 2026Jinzong Dong, Zhaohui Jiang, Bo YangCovariate ShiftConfidence Estimation

  2. When to Trust Confidence Thresholding: Calibration Diagnostics for Pseudo-Labelled Regression

    May 12, 2026Marcell T. KurbuczDecision ThresholdsLogistic Regression