Uncertainty quantification (UQ) is crucial in safety-critical applications such as medical image segmentation. Total uncertainty is typically decomposed into data-related aleatoric uncertainty (AU) and model-related epistemic uncertainty (EU). Many methods exist for modeling AU (such as Probabilistic UNet, Diffusion) and EU (such as ensembles, MC Dropout), but it is unclear how they interact when combined. Additionally, recent work has revealed substantial entanglement between AU and EU, undermining the interpretability and practical usefulness of the decomposition. We present a comprehensive empirical study covering a broad range of AU-EU model combinations, propose an entanglement proxy based on the relative performance of uncertainty measures, and evaluate model combinations across downstream uncertainty quantification tasks. Ensembles consistently show more favorable proxy values and superior performance. Softmax models usually beat other AU methods, except in calibration where the results are dataset-dependent. A softmax ensemble performs remarkably well on all tasks. Finally, we analyze potential sources of uncertainty entanglement and outline directions for mitigating this effect.
Figures & tables
Figure 1 : A visualization of aleatoric and epistemic modeling components and their aggregation into uncertainty measures for downstream tasks.
Dataset
Modality
nc
#Im
Ann/Im
Train
Val
Test (ID/OOD)
LIDC-IDRI
CT
2
15096
4
9355
2689
3052 / 3052×3
MMIS NPC
MRI
2
2260
4
1483
302
475 / 475×3
Chákṣu IMAGE
Fundus imaging
3
1345
5
648
162
264 / 271
Table 1 : Summary of the datasets used in our experiments ( nc:= number of classes).
Figure 2 : Example images from the three datasets, showing both ID and OOD images. The four binary-mask delineations are shown on LIDC and NPC as red lines, while the five cup and disc segmentations are shown as blue and green lines on Chaksu.
Figure 3 : A visualization of the entanglement proxy Δ . (a) Point 1 has a more favorable Δ value than point 2. The proxy is proportional to the signed angles ( ϕ ). (b) Angles are converted to Δ based on the shown number line.
Figure 4 : Mean model predictions ( Eθ[Ey[p]] ) on NPC data (ID).
Figure 5 : Aleatoric (AU), epistemic (EU) and total uncertainty (TU) maps for ID NPC images, with red numbers indicating the maximum value.
Figure 6 : Scatter plots showing the performance of model combinations on the different tasks (OOD detection, AMB, CAL). Each point corresponds to a model combination, with y-values based on the theoretically consistent uncertainty measure and x-values using the inconsistent one.
Figure 7 : A scatter plot of the mean performance (aggregated over datasets) and entanglement-proxy ( Δ ) rankings for different tasks. Lines indicate confidence intervals based on Student’s t -distribution ( N=5 ).
Table 2 : Average ranking (1-19) of performance and the entanglement proxy ( Δ ) when comparing the 19 AU-EU model combinations. The ranks were calculated separately per task and averaged across tasks.
Figure 8 : Top : Scatter plots comparing mean AU and EU across models and datasets. Bottom : Mean ratio of EU/AU across models and datasets on OOD images (using image-wise mean). No EU models are left out of axis limits for better resolution.
Figure 9 : Scatter plots showing OOD detection performance as different aggregation strategies are used. Each small point is an AU-EU model combination, and the large crosses are the mean performances across all combinations.
Table 3 : Our recommendations for AU-EU model combinations across tasks, based on whether performance, favorable entanglement-proxy values or both are important. AU-EU model combinations are assigned colors for an easy overview.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Abbreviation
Meaning
AU
Aleatoric uncertainty
EU
Epistemic uncertainty
TU
Total uncertainty
ID
In-distribution
OOD
Out-of-distribution
AMB
Ambiguity modeling
Appendix
Table 4 : Abbreviations used throughout the paper.
Figure 10 : Mean predictions ( Eθ[Ey[p]] ) for the different model combinations across datasets and distribution settings.
Figure 11 : Aleatoric (AU), epistemic (EU) and total uncertainty (TU) maps for in-distribution (ID) images from the LIDC and Chaksu datasets. Red numbers indicate the maximum value.
Figure 12 : Aleatoric (AU), epistemic (EU) and total uncertainty (TU) maps for out-of-distribution (OOD) images from the LIDC and Chaksu datasets. Red numbers indicate the maximum value.
Figure 13 : Mean performance and entanglement-proxy ( Δ ) ranking for different datasets (aggregated over tasks). Lines indicate confidence intervals based on Student’s t -distribution ( N=5 ).
Figure 14 : Identical to Fig. 6 but with OOD splits for AMB and CAL.
Figure 15 : Same as the left column of Fig. 6 but with each OOD split for LIDC (blur, noise, contrast) and NPC (Gibbs, noise, hist) shown separately.
Figure 16 : The performance metrics as the number of AU predictions are varied. We used 10 samples in the paper as a performance vs compute trade-off.
Table 5 : Average ranking (1-19) of performance and the entanglement proxy ( Δ ) when comparing the 19 AU-EU model combinations. The ranks were calculated separately per task and averaged across tasks.
Figure 17 : Scatter plots showing the OOD detection performance of model combinations with confidence intervals (otherwise the same as in Fig. 6 ). Each point corresponds to a model combination, with y-values based on the theoretically consistent uncertainty measure and x-values using the inconsistent one.
Figure 18 : Scatter plots showing the AMB performance of model combinations with confidence intervals (otherwise the same as in Fig. 6 ). Each point corresponds to a model combination, with y-values based on the theoretically consistent uncertainty measure and x-values using the inconsistent one.
Figure 19 : Scatter plots showing the CAL performance of model combinations with confidence intervals (otherwise the same as in Fig. 6 ). Each point corresponds to a model combination, with y-values based on the theoretically consistent uncertainty measure and x-values using the inconsistent one.
School of Data Science, Fudan University, Shanghai, China · Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University, Shanghai, China · National Heart and Lung Institute, Imperial College London, London, United Kingdom