cs.AISep 30, 2026

How Much Can Reliability Drift Under a Fixed Confidence Distribution?

Authors: Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen

Organizations: Adelaide University

Abstract

A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliability relation under covariate shifts that preserve the distribution of the confidence score, constraining the reweighting within each confidence level by a χ2χ^2 budget; the resulting worst case, as a function of the budget, is a fragility profile. On an interval of budgets that can be computed from the source distribution, the profile equals exactly the square root of the budget times the within-level variance of the correctness propensity -- the grouping-loss term of calibration-refinement decompositions. Beyond this interval the profile is governed by the tails of the propensity law, and the entire upward profile determines the centred within-level law; consequently, calibration residual and grouping variance do not determine fragility in general, though they do when labels and predictions are deterministic. Since the propensity is not observed, we restrict reweightings to a learned finite readout within confidence bins, bound the part the restriction misses by the grouping variance remaining inside readout cells, estimate the restricted profile with role-separated labels, and provide a separate split-sample lower confidence bound. On ImageNet this bound is positive in both splits for four of six primary classifiers and nine of twelve additional ones as released, and for three of eighteen after temperature scaling. Held-out drift under optimised reweightings fitted without evaluation labels tracks the estimated profile; an exploratory label-permutation diagnostic yields near-zero agreement for this statistic while largely reproducing the correlation observed for unsigned random reweightings.

Figures & tables

Appendix figures & tables30 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 20, 2026cs.LG

Expectation Consistency Loss: Rethink Confidence Calibration under Covariate Shift

Confidence calibration for classification models is vital in safety-critical decision-making scenarios and has received extensive attention. General confidence calibration methods assume training and test data are independent and identically distributed, limiting their effectiveness under covariate shifts. Previous calibration methods under covariate shift struggle with class-wise or canonical calibrations and often rely on unstable importance weighting when density ratios are large or unbounded. Given the above limitations, this paper rethinks confidence calibration under covariate shifts. First, we derive a necessary and sufficient condition for confidence calibration under covariate shifts, named Expectation consistency condition, which reveals covariate shifts do not necessarily lead to uncalibrated confidence and provides a weaker condition for confidence calibration than global covariate distribution alignment. Then, utilizing Expectation consistency condition, this paper proposes an unsupervised domain adaptation loss to calibrate confidence of the target domain, named Expectation consistency loss (ECL), which is compatible with canonical calibration, class-wise calibration, and top-label calibration. Third, we prove that computing ECL loss has the same sample complexity as Expected Calibration Error (ECE) and provide a theoretically grounded mini-batch trainable scheme for ECL loss. Finally, we validate the effectiveness of our method on both simulated and real-world covariate shift datasets.
May 12, 2026stat.ME

When to Trust Confidence Thresholding: Calibration Diagnostics for Pseudo-Labelled Regression

Calibrated probability outputs of trained classifiers are increasingly used as inputs to downstream regression estimands such as effects, prevalences, or disparities for a latent group observed only on a small labelled subset. A standard practice is to threshold the calibrated score at a confidence cutoff and treat the hard label as the truth. Building on a recent identification result for the underlying moment equation, we develop a calibration-aware diagnostic apparatus for pseudo-labelling pipelines. We derive a closed-form expression for the attenuation bias that confidence thresholding induces in the downstream regression coefficient, and show that the bias can be predicted, before any inference is run, from the residual score variance V∗=E[Var⁡(p∣X)]V^{*}=\mathbb{E}[\operatorname{Var}(p\mid X)] on the unlabelled set after partialling out the downstream controls XX. We further obtain a sharp sensitivity bound under bounded calibration drift, and identify the boundary V∗=0V^{*}=0, which holds iff pp is a deterministic function of XX; this motivates a structural separation between classifier features WW and downstream controls X⊊WX\subsetneq W. Five controlled simulations and a UCI Adult illustration trace the predictions. The contribution is operational: a (V∗,κ)(V^{*}, κ) decision rule that practitioners can compute from any classifier output to decide whether confidence thresholding is safe.
Jul 20, 2026cs.LG

The Label Complexity of Useful Class-Conditional Prediction Sets under Distribution Shift

Prediction sets can make deployed classifiers safer by returning several plausible labels when a single prediction is uncertain. Their value depends on classwise reliability: average coverage can meet its target while rare or difficult classes fail repeatedly. This concern is sharper after distribution shift, when calibration labels come from a source environment but reliability is needed on the target. We ask what labeled source data and unlabeled target inputs reveal about class-conditional prediction sets, and when target labels are necessary. Under unrestricted joint shift, two target laws can produce the same observable data while requiring different classwise thresholds; any label-free rule covering both must enlarge its sets on one law. We give a labeled target audit that estimates the missing quantiles with a simultaneous guarantee. Probability-scale error is invariant to increasing score transformations, and threshold recovery follows under local regularity. At fixed confidence, achieving threshold tolerance ε\varepsilon with fixed, nonadaptive class-stratified labeled pairs has total complexity Θ(Kε−2log⁡K)Θ(K\varepsilon^{-2}\log K), or Θ(ε−2log⁡K)Θ(\varepsilon^{-2}\log K) labels per class under equal allocation. Class imbalance creates a separate acquisition cost; for foreground class probabilities of order 1/K1/K, the mixed-stream label complexity is also Θ(Kε−2log⁡K)Θ(K\varepsilon^{-2}\log K) at fixed confidence. Experiments on action-recognition and image shifts show that marginal coverage can conceal severe class failures and that source classwise calibration depends on the shift. The results connect the information available at deployment to the target labels needed for useful class-conditional prediction.