stat.MEMay 12, 2026

When to Trust Confidence Thresholding: Calibration Diagnostics for Pseudo-Labelled Regression

Authors: Marcell T. Kurbucz

Abstract

Calibrated probability outputs of trained classifiers are increasingly used as inputs to downstream regression estimands such as effects, prevalences, or disparities for a latent group observed only on a small labelled subset. A standard practice is to threshold the calibrated score at a confidence cutoff and treat the hard label as the truth. Building on a recent identification result for the underlying moment equation, we develop a calibration-aware diagnostic apparatus for pseudo-labelling pipelines. We derive a closed-form expression for the attenuation bias that confidence thresholding induces in the downstream regression coefficient, and show that the bias can be predicted, before any inference is run, from the residual score variance V=E[Var(pX)]V^{*}=\mathbb{E}[\operatorname{Var}(p\mid X)] on the unlabelled set after partialling out the downstream controls XX. We further obtain a sharp sensitivity bound under bounded calibration drift, and identify the boundary V=0V^{*}=0, which holds iff pp is a deterministic function of XX; this motivates a structural separation between classifier features WW and downstream controls XWX\subsetneq W. Five controlled simulations and a UCI Adult illustration trace the predictions. The contribution is operational: a (V,κ)(V^{*}, κ) decision rule that practitioners can compute from any classifier output to decide whether confidence thresholding is safe.

Explore similar work

Sep 18, 2025cs.LG

Efficient Conformal Prediction for Regression Models under Label Noise

In high-stakes scenarios, such as medical imaging applications, it is critical to equip the predictions of a regression model with reliable confidence intervals. Recently, Conformal Prediction (CP) has emerged as a powerful statistical framework that, based on a labeled calibration set, generates intervals that include the true labels with a pre-specified probability. In this paper, we address the problem of applying CP for regression models when the calibration set contains noisy labels. We begin by establishing a mathematically grounded procedure for estimating the noise-free CP threshold. Then, we turn it into a practical algorithm that overcomes the challenges arising from the continuous nature of the regression problem. We evaluate the proposed method on two medical imaging regression datasets with Gaussian label noise. Our method significantly outperforms the existing alternative, achieving performance close to the clean-label setting.
Yahav Cohen, Jacob Goldberger, Tom Tirer
Jul 20, 2026cs.LG

The Label Complexity of Class-Conditional Coverage under Distribution Shift

Conformal prediction certifies that a classifier's prediction sets cover the truth, and that certificate is marginal. Many recognition benchmarks build distribution shift into evaluation, placing disjoint conditions in the training and test splits. Under that shift the certificate stays reassuring while per class coverage fails silently: on a real cross subject skeleton benchmark marginal coverage holds near ninety percent while the worst class is covered about seventy percent and ten of sixty classes fall below eighty percent. This class specific undercoverage stays hidden behind a single reassuring marginal number. Once the shift acts jointly on covariates and labels, the target class conditional score law is unidentified, so no label free method is at once per class valid and efficient uniformly over target laws consistent with the observed source joint distribution and target covariate marginal. The per class labels needed to recover every class threshold to a given tolerance grow as the inverse square of that tolerance and the logarithm of the class count, with matching bounds for classwise threshold procedures. Pseudo labels do not shortcut it: the best prediction powered estimator gains at most a small constant factor where coverage collapses. Across three real shifts and an image corruption benchmark, source label calibration recovers much of the gap while marginal coverage holds, and stops once it breaks.
Weijia Han, Lisha Qu
Jun 10, 2026stat.ML

Conformal Bayes under Label Shift: Post-Hoc Calibration vs. In-Training Adaptation

Conformal Bayes combines Bayesian posterior predictives with conformal calibration to produce prediction sets that are both statistically valid and geometrically efficient. We study conformal Bayes under label shift from a unified perspective, identifying two complementary approaches that restore nominal target-domain coverage through importance-weighted conformal calibration but operate through independent mechanisms. \emph{Post-hoc calibration} tilts the posterior predictive toward the target domain and corrects the conformal threshold via an importance-weighted quantile, leaving the parameter posterior unchanged. \emph{In-training adaptation} tilts the parameter posterior itself to the target domain, producing a corrected predictive whose highest predictive density region serves as the highest predictive density (HPD)-based prediction set under the fitted target predictive; efficiency is model-dependent and does not imply finite-sample conditional optimality. Two controlled experiments isolate the regime-dependence of each strategy: in the low-dimensional, well-estimated regime StrategyA produces the narrowest valid intervals, while in the high-dimensional, underdetermined regime StrategyB achieves up to 43%43\% width reduction at unchanged coverage, under the stated source-sampling and label-shift assumptions.
Seungjin Choi