Uncertainty estimation in machine learning typically decomposes uncertainty into aleatoric uncertainty (AU) and epistemic uncertainty (EU) using the standard information-theoretic framework. However, in practice, two critical issues arise: entanglement (AU and EU are highly correlated) and epistemic collapse (EU magnitude shrinks with increasing model capacity). We analyze this framework on a functional level and discover that significant portions of the assumed AU, EU range are infeasible in finite settings, and cannot be attained with any class probabilities. We characterize how this infeasible region scales with the number of classes and Monte Carlo samples N (e.g., from ensembles with N members), revealing it is bounded by AU≤log(2)/N. Crucially, the infeasible region's boundary helps explain epistemic collapse: when model confidence is high, AU>EU is guaranteed by this fundamental structural limitation. Our findings show that increasing ensemble size mitigates epistemic collapse by reducing the infeasible area. Lastly, we caution against interpreting AU and EU as independent quantities in low AU regimes, since we show they are coupled when AU≤log(2)/N.
Figures & tables
Figure 1: Visualizations of the uncertainty space (Aleatoric Uncertainty vs Epistemic Uncertainty) along with sample uncertainties from probabilities in different settings. The probabilities are simulated in the first two panels, while the last panel shows real probabilities from an ensemble of five models ( N=5 ) trained on C=5 classes of CIFAR-10.
Figure 2: Simulated (AU,TU) values for different posterior sample counts N and class counts C , with the bounds from Theorem 1 .
Figure 3: Simulated (AU,TU) values for different posterior sample counts N and class counts C . The points with AU=0 are highlighted and labeled by their class counts Nq .
Figure 4: Binary entropy at the original and perturbed prediction probabilities.
Figure 5: Simulated (AU,TU) values for different posterior sample counts N and class counts C . Curves from D are labeled by deterministic class counts. The underlined entries mark the two classes supported by the non-deterministic posterior sample. For example, [2,0,0] connects the count vectors [3,0,0] and [2,1,0] .
Figure 6: Infeasible-boundary shapes, shown as lower envelopes of interpolating curves, for different posterior sample counts N and class counts C .
Figure 7: Visualizations of the ( AU , EU ) space along with points from an ensemble of 10 EfficientNet-B0 networks trained on CIFAR-10.
Figure 8: OOD detection performance and epistemic collapse as the number of Enet-B0 networks (N) is varied.
Figure 9: Visualizations of the ( AU , EU ) space along with points from an Enet-B4 ensemble trained on CIFAR as the number of networks (N) is varied. The infeasible boundary is shown in red, and the dashed line is AU=EU .
Figure 10: OOD detection performance and epistemic collapse as the network capacity is varied.
Figure 11: Visualizations of the ( AU , EU ) space along with points from an ensemble/dropout model as the network capacity is varied. The model is trained on CIFAR. The infeasible boundary is shown in red, and the dashed line is AU=EU .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12: Simulated (AU,TU) values for different settings with C=N=5 .
Figure 13: Comparison of jackknife-corrected EU estimates for different posterior sample counts N on CIFAR.
Figure 14: Comparison of jackknife-corrected EU estimates for different posterior sample counts N on MNIST.
The standard taxonomy of predictive uncertainty defines epistemic uncertainty as the part removable by collecting more data, while the standard measure identifies it with a mutual-information term. We prove the definition and the measure are extensionally inconsistent. On an explicit construction, the measure assigns all uncertainty to the epistemic class, yet no quantity of training data reduces it. Reducibility is instead a property of the pair (uncertainty, acquisition class), and the dichotomy resolves into three parts: aleatoric, sample-reducible epistemic, and mechanism-reducible epistemic uncertainty. An exact identity for the value of an observation shows that in-distribution data never reduces mechanism-irreducible uncertainty and generically increases it. Ensemble disagreement, the deployed epistemic estimate, tracks the training procedure rather than the epistemic term. It collapses to zero beneath a positive truth under consistent training, and equals hyperparameter-scaled initialization noise under interpolation. A finite-sample falsification test and seed-swept experiments confirm the theory.
Uncertainty estimation is critical for deploying machine learning models in high-stakes settings. However, classical calibration only assesses the reliability of predicted probabilities and does not evaluate whether epistemic uncertainty estimates are themselves trustworthy. This limitation is particularly relevant for second-order classification models. We introduce epistemic calibration, a principled criterion that measures whether reported epistemic uncertainty faithfully reflects the dispersion of model predictions around the ground truth. We show that epistemic calibration is a strictly stronger notion than classical calibration and captures failure modes invisible to standard metrics. We relate this work to the existing literature through an impossibility theorem that holds under the epistemic calibration hypothesis. To operationalize this concept, we propose the Expected Epistemic Calibration Error (EECE), which we prove to be a consistent estimator of a True Epistemic Calibration Error (TECE). Experiments across a broad range of uncertainty quantification methods show that epistemic calibration is a coherent and meaningful criterion and reveal substantial differences across methods, despite similar predictive performance.
Arthur Hoarau
Université de Lorraine, CentraleSupélec · Loria, CNRS, Metz, France
Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.
Jakub Paplhám, Willem Waegeman, Eyke Hüllermeier +1
Czech Technical University in Prague · Ghent University · LMU Munich, MCML, DFKI