Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification
Authors: Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep Kaushik, Marion Smits, Stefan Klein
Organizations: Department of Radiology and Nuclear Medicine, Erasmus MC, University Medical Center Rotterdam, Rotterdam, the Netherlands · Department of Medical Informatics, Amsterdam UMC, University of Amsterdam, Amsterdam, the Netherlands · GE HealthCare, Munich, Germany · GE HealthCare, USA · Brain tumor Centre, Erasmus MC Cancer Institute, Rotterdam, the Netherlands · Medical Delta, Delft, the Netherlands
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.
Figures & tables
Figure 1: Overview of the multi-task glioma subtyping framework with uncertainty quantification. The input data consist of four structural MRI sequences per patient, namely T1-weighted (T1w), contrast-enhanced T1-weighted (T1wCE), T2-weighted (T2w), and fluid-attenuated inversion recovery (FLAIR). The sequences are preprocessed using registration, bias field correction, and skull stripping. Subsequently, the data are used to train a Deep Learning model that jointly performs tumor segmentation and prediction of IDH mutation status, 1p/19q co-deletion status, and tumor grade. For each task, predictive, aleatoric, and epistemic uncertainty are computed. For tumor segmentation, voxel-wise uncertainty maps are converted into case-level uncertainty scores using four aggregation strategies: brain-level aggregation, predicted-tumor aggregation, dilated-tumor aggregation, and distance-weighted aggregation around the tumor region. In the tumor features panel, the colors of the icons vary from tumors associated with poorer prognosis (red) to those with better prognosis (green).
Set
Dataset
Number of cases
Total
Train
BraTS
156
1466
Brain tumor Progression
20
CPTAC-GBM
45
EGD
775
IvyGAP
39
In-house
322
Table 1: Train and test data distribution
Metric
Purpose
Possible Values
Interpretation
U-AUC
Measures the probability that a randomly chosen error case has higher uncertainty than a correct case.
0 – 1
Values close to 1 indicate uncertainty correctly ranks errors above successes.
AP
Evaluates precision-recall trade-off for predicting errors, suitable for imbalanced tasks.
Quantifies error-identification performance relative to baseline prevalence.
0−ϵ1
Values >1 indicate that uncertainty successfully prioritizes error cases. A value of 1 indicates random selection.
AURC
Evaluates the mean residual risk among retained cases as coverage varies from most certain to all cases.
0 – 1
Lower values indicate that uncertainty successfully retains lower-risk cases at reduced coverage.
Table 2: Summary of the metrics used to assess operational utility of uncertainty estimates, namely, Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC). For each metric, the purpose, possible values, and interpretation are presented.
Figure 2: First three columns from left to right: predictive, aleatoric and epistemic uncertainty value heatmaps across all tasks as a function of the number of MCD samples T and the dropout rate δ . The value at each cell of the heatmap represents the mean uncertainty across all cases of the test set. The last column presents the average value for each type of uncertainty over all experiments with the blue, green and red bars representing predictive, aleatoric and epistemic uncertainty, respectively. Uncertainty values for classification tasks were normalized at the same scale to ease the analysis of the patterns. For the segmentation task the heatmaps were kept in the original scale.
Figure 3: Expected Calibration Error (ECE), Negative Log-Likelihood (NLL) and Receiver Operating Characteristic Area Under the Curve (AUC) values as a function of dropout rate δ . For each point in every graph, a 95% confidence interval is presented. This interval was obtained by 1000× bootstrap resampling of the test set.
Figure 4: Tumor segmentation uncertainty quality analysis. The top panel presents the Pearson correlation coefficient between case-level predictive uncertainty and Dice Similarity Coefficient (DSC) as a function of the dropout rate δ used for the MCD ensemble. Correlations are shown for four spatial aggregation strategies: brain-mask aggregation, predicted-tumor aggregation, 5-voxel dilated predicted-tumor aggregation, and boundary-weighted aggregation using σ=5 voxels. More negative correlations indicate that higher case-level uncertainty is associated with lower segmentation accuracy. The bottom panel presents the mean DSC between the predicted tumor segmentation and the ground truth in the test set as a function of the dropout rate δ .
Task
Class distribution
Accuracy ↑
Balanced accuracy ↑
Macro F1 score ↑
AUC ↑
IDH
Wildtype: 129 Mutated: 85
0.82 [0.77,0.87]
0.79 [0.73,0.84]
0.80 [0.74,0.85]
0.90 [0.85,0.94]
1p/19q
Intact: 205 Co-deleted: 25
0.90 [0.85,0.93]
0.56 [0.50,0.63]
0.57 [0.47,0.67]
0.85 [0.77,0.91]
Grade
Grade 2: 45 Grade 3: 58 Grade 4: 132
0.71 [0.64,0.77]
0.60 [0.57,0.64]
0.51 [0.47,0.54]
0.85 [0.81,0.88]
Table 3: Predictive performance at the selected MCD configuration ( T=30 , δ=0.25 ) for the three classification tasks. For each task, the class distribution is also reported. Values are shown with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. AUC is reported at task level: for tumor grade, it corresponds to the macro-averaged one-vs-rest AUC across grade classes. Total test set size differs across tasks because molecular and histopathological annotations were not available for all patients. Arrows in metric column headers indicate that higher values are preferable ( ↑ ).
Figure 5: Violin plots of the predictive uncertainty across different tasks, comparing the distributions of high quality and low quality predictions, obtained with the optimal configuration of the MCD model at T=30 and δ=0.25 . For the classification tasks, high and low quality predictions correspond to correctly and incorrectly classified cases, respectively. For the tumor segmentation task, high and low quality predictions represent cases with DSC>0.7 and DSC≤0.7 , respectively. For all tasks, the difference between the distributions for both groups was statistically significant (Mann–Whitney U test).
Task
Error rate ↓
Uncertainty
Predictive
Aleatoric
Epistemic
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation status prediction
0.18
0.73 [0.62,0.82]
0.42 [0.28,0.57]
2.37 [1.72,3.34]
0.10 [0.05,0.15]
0.71 [0.61,0.80]
0.32 [0.22,0.46]
1.79 [1.39,2.54]
0.10 [0.05,0.15]
0.71 [0.61,0.80]
0.34 [0.23,0.50]
1.89 [1.45,2.80]
0.10 [0.05,0.16]
1p/19q co-deletion status prediction
0.10
0.84 [0.76,0.91]
0.37 [0.22,0.58]
3.55 [2.37,6.07]
0.03 [0.01,0.05]
0.84 [0.77,0.91]
0.39 [0.22,0.58]
3.70 [2.51,6.19]
0.03 [0.01,0.04]
0.68 [0.58,0.77]
0.16 [0.10,0.29]
1.54 [1.22,2.59]
0.05 [0.03,0.08]
Tumor grade prediction
0.29
0.78 [0.72,0.84]
0.59 [0.47,0.71]
2.00 [1.66,2.47]
0.12 [0.09,0.17]
0.79 [0.73,0.85]
0.59 [0.47,0.72]
2.00 [1.68,2.50]
0.12 [0.09,0.17]
0.70 [0.62,0.77]
0.45 [0.35,0.58]
1.52 [1.28,1.94]
0.17 [0.12,0.24]
Tumor segmentation
0.17
0.95 [0.92,0.98]
0.82 [0.72,0.91]
4.95 [3.89,6.67]
0.13 [0.12,0.14]
0.96 [0.94,0.98]
0.83 [0.72,0.91]
5.01 [3.93,6.80]
0.13 [0.12,0.14]
0.95 [0.92,0.98]
0.81 [0.71,0.90]
4.90 [3.85,6.64]
0.13 [0.12,0.14]
Table 4: Operational utility of uncertainty estimates per task. Assessment is made using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC), derived from using uncertainty as a predictor of errors. Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. The numerically highest values for U-AUC , AP, and Lift, and the numerically lowest values for AURC , are indicated in bold for each task. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Task
Method
Error ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation
DE
0.17
0.72 [0.62,0.82]
0.37 [0.24,0.56]
2.15 [1.57,3.23]
0.10 [0.05,0.15]
MCD
0.18
0.73 [0.62,0.82]
0.42 [0.28,0.57]
2.37 [1.72,3.34]
0.10 [0.05,0.15]
MCDE
0.17
0.73 [0.63,0.83]
0.34 [0.23,0.50]
2.03 [1.53,2.98]
0.09 [0.05,0.14]
1p/19q co-deletion
DE
0.10
0.84 [0.76,0.91]
0.39 [0.22,0.59]
3.90 [2.53,6.62]
0.03 [0.01,0.04]
MCD
0.10
0.84 [0.76,0.91]
0.37 [0.22,0.58]
3.55 [2.37,6.07]
0.03 [0.01,0.05]
MCDE
0.09
0.81 [0.74,0.89]
0.26 [0.16,0.45]
2.84 [2.06,4.91]
0.03 [0.01,0.04]
Table 5: Comparison of uncertainty quantification methods using predictive uncertainty. Assessment is made using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC), derived from using predictive uncertainty as a predictor of errors. Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. For tumor segmentation, predictive uncertainty was aggregated within the predicted tumor region. The numerically highest values for U-AUC , AP, and Lift, and the numerically lowest values for AURC , are indicated in bold for each task. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Figure 6: Distribution of tumor segmentation Dice Similarity Coefficient (DSC) scores for correctly and incorrectly classified cases for the IDH mutation status, 1p/19q co-deletion status and tumor grade. Cases are stratified according to whether the corresponding prediction was correct or incorrect. For each task, higher DSC values indicate better agreement between predicted and reference tumor segmentations. Statistical significance between groups was found for the IDH mutation status and tumor grade prediction (Mann–Whitney U test).
Task
Uncertainty
Predictive
Aleatoric
Epistemic
IDH mutation status prediction
0.48 [0.38, 0.57]
0.51 [0.42, 0.58]
0.09 [-0.04, 0.21]
1p/19q co-deletion status prediction
0.22 [0.12, 0.33]
0.24 [0.15, 0.33]
−0.15 [-0.25, -0.05]
Tumor grade prediction
0.53 [0.45, 0.61]
0.50 [0.42, 0.57]
0.10 [-0.03, 0.26]
Table 6: Pearson correlation ( ρ ) between tumor segmentation uncertainty and classification uncertainty for each task. Correlations were computed separately for predictive, aleatoric, and epistemic uncertainty components. Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. The numerically highest correlation coefficient per task is indicated in bold.
Task
Error rate ↓
Tumor segmentation uncertainty as a predictor of classification errors
Predictive
Aleatoric
Epistemic
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation status prediction
0.18
0.62 [0.51,0.72]
0.30 [0.21,0.46]
1.68 [1.27,2.60]
0.15 [0.08,0.22]
0.65 [0.55,0.75]
0.28 [0.20,0.40]
1.54 [1.22,2.23]
0.13 [0.07,0.20]
0.61 [0.50,0.72]
0.32 [0.22,0.46]
1.79 [1.30,2.72]
0.15 [0.09,0.23]
1p/19q co-deletion status prediction
0.10
0.68 [0.57,0.78]
0.23 [0.12,0.41]
2.16 [1.38,4.15]
0.06 [0.03,0.11]
0.69 [0.57,0.79]
0.27 [0.14,0.43]
2.55 [1.50,4.56]
0.06 [0.03,0.11]
0.67 [0.56,0.78]
0.24 [0.12,0.40]
2.26 [1.36,4.19]
0.06 [0.03,0.11]
Tumor grade prediction
0.29
0.65 [0.57,0.73]
0.44 [0.35,0.58]
1.52 [1.25,1.94]
0.22 [0.15,0.30]
0.67 [0.59,0.74]
0.42 [0.34,0.55]
1.45 [1.23,1.82]
0.19 [0.14,0.26]
0.65 [0.56,0.73]
0.45 [0.35,0.58]
1.54 [1.27,1.95]
0.22 [0.16,0.30]
Table 7: Ability of tumor segmentation uncertainty to predict classification errors. Assessment is performed using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC). Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. The numerically highest values for U-AUC , AP, and Lift, and the numerically lowest values for AURC , are indicated in bold for each task. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Task
Error rate ↓
Trust score
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation status prediction
0.18
0.73 [0.64,0.81]
0.38 [0.26,0.53]
2.12 [1.58,3.10]
0.09 [0.05,0.14]
1p/19q co-deletion status prediction
0.10
0.85 [0.77,0.91]
0.42 [0.25,0.58]
3.98 [2.71,6.39]
0.02 [0.01,0.04]
Tumor grade prediction
0.29
0.77 [0.71,0.83]
0.51 [0.41,0.65]
1.77 [1.48,2.23]
0.13 [0.09,0.17]
Table 8: Performance of the proposed trust score for predicting classification errors. Assessment is performed using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC). Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Overview of the multi-task DL network used for the tumor segmentation and prediction of IDH mutation status, 1p/19q co-deletion status and tumor grade. The architecture processes full 3D volumes. The labels in the figure define the layer type, and the numbers indicate either the number of filters, dense units or features. Figure based on the one presented in the work of van der Voort et al. (2023) .
τDSC
Error rate ↓
Uncertainty
U-AUC ↑
AP ↑
Lift ↑
Predictive
0.96 [0.93,0.98]
0.78 [0.64,0.90]
7.96 [5.73,12.41]
0.60
0.10
Aleatoric
0.97 [0.95,0.99]
0.82 [0.67,0.92]
8.33 [5.99,13.09]
Epistemic
0.96 [0.93,0.98]
0.76 [0.61,0.88]
7.77 [5.61,12.26]
Predictive
0.95 [0.92,0.98]
0.82 [0.72,0.91]
4.95 [3.89,6.67]
0.70
0.17
Aleatoric
0.96 [0.94,0.98]
0.83 [0.72,0.91]
5.01 [3.93,6.80]
Epistemic
0.95 [0.92,0.98]
0.81 [0.71,0.90]
4.90 [3.85,6.64]
Appendix
Table 9: Sensitivity of segmentation uncertainty error detection to the DSC threshold used to define segmentation error. For each threshold, segmentation error was defined as DSC≤τDSC . Error rate is reported as the proportion of cases classified as segmentation errors. U-AUC , AP, and Lift are reported for predictive, aleatoric, and epistemic uncertainty, with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Figure 8: Case-wise association between predictive segmentation uncertainty and tumor segmentation DSC at the selected MCD configuration ( T=30 , δ=0.25 ). Results are shown for brain mask, predicted tumor, 5-voxel dilated tumor, and boundary-weighted aggregation ( σ=5 voxels). Each point represents one test case; Pearson ρ and Spearman ρs are reported for each strategy.
Dataset
Subset
IDH mutation status
1p/19q co-deletion status
Tumor grade
Wildtype
Mutated
N/A
Intact
Co-deleted
N/A
2
3
4
N/A
Train
BraTS
0
0
156
0
0
156
0
0
0
156
Brain tumor Progression
0
0
20
0
0
20
0
0
0
20
CPTAC-GBM
0
0
45
0
0
45
0
0
0
45
EGD
312
155
308
186
73
516
135
80
502
58
IvyGAP
32
6
1
27
3
9
0
1
36
2
Appendix
Table 10: Detailed distribution of the data used for training and testing the Deep Learning (DL) architecture for each one of the classification tasks. N/A (not available) represents missing data. Tumor segmentation ground truth was available for all cases.
Task
Method
Error ↓
Aleatoric uncertainty
Epistemic uncertainty
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation
DE
0.17
0.73 [0.62,0.82]
0.39 [0.26,0.56]
2.26 [1.66,3.31]
0.09 [0.05,0.15]
0.64 [0.54,0.74]
0.28 [0.17,0.43]
1.60 [1.20,2.45]
0.12 [0.07,0.18]
MCD
0.18
0.71 [0.61,0.80]
0.32 [0.22,0.46]
1.79 [1.39,2.54]
0.10 [0.05,0.15]
0.71 [0.61,0.80]
0.34 [0.23,0.50]
1.89 [1.45,2.80]
0.10 [0.05,0.16]
MCDE
0.17
0.73 [0.62,0.82]
0.36 [0.24,0.53]
2.12 [1.55,3.13]
0.09 [0.05,0.14]
0.64 [0.54,0.73]
0.30 [0.20,0.47]
1.81 [1.31,2.83]
0.12 [0.07,0.19]
1p/19q co-deletion
DE
0.10
0.83 [0.75,0.90]
0.29 [0.18,0.49]
2.95 [2.12,5.25]
0.03 [0.01,0.04]
0.66 [0.55,0.76]
0.18 [0.10,0.34]
1.79 [1.19,3.65]
0.05 [0.03,0.08]
MCD
0.10
0.84 [0.77,0.91]
0.39 [0.22,0.58]
3.70 [2.51,6.19]
0.03 [0.01,0.04]
0.68 [0.58,0.77]
0.16 [0.10,0.29]
1.54 [1.22,2.59]
0.05 [0.03,0.08]
Appendix
Table 11: Comparison of Uncertainty Quantification methods using aleatoric and epistemic uncertainty. Assessment is made using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve ( AURC ), derived from using each uncertainty component as a predictor of errors. Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. The error column denotes the method-specific classification error rate for classification tasks and the proportion of cases with DSC ≤0.70 for tumor segmentation. For the latter, uncertainty was aggregated within the predicted tumor region. The highest displayed values for U-AUC, AP, and Lift, and the lowest displayed values for AURC , are indicated in bold for each task and uncertainty component. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Glioma segmentation in multiparametric MRI is a critical component of treatment planning. A segmentation model that fails silently on treatment-critical sub-regions represents a patient safety risk that overlap-based metrics such as Dice scores cannot expose. We ask whether voxel-level uncertainty estimation via Monte Carlo (MC) Dropout can reliably identify segmentation errors in clinically critical sub-regions, and whether calibration failure modes are detectable from standard reporting metrics alone. In an empirical two-model case study on 126 BraTS21 patients, we evaluate a high-performance pretrained SegResNet and a locally trained UNet with residual units (UNet-Res). MC dropout preserved segmentation accuracy (∣ΔDice∣<0.01) while achieving strong uncertainty-error alignment (AUROC for entropy (H) ≈0.97), indicating uncertainty correctly ranks erroneous voxels above correct ones. Entropy-based patient stratification identified a high-uncertainty subgroup with substantially lower segmentation performance (median whole-tumour Dice 0.835 vs. 0.925), supporting uncertainty as a practical triage signal. However, global alignment can mask important region-specific differences. Despite similar AUROC, UNet-Res exhibited near-zero enhancing tumour entropy (0.054) and Expected Calibration Error (ECE) of 0.915, with a Dice of only 0.714, indicating severely miscalibrated confidence on the most clinically critical sub-region, a failure mode invisible to standard Dice and AUROC reporting. These findings demonstrate that strong uncertainty-error alignment is necessary but insufficient for clinical safety: sub-region-specific calibration assessment must accompany AUROC evaluation when selecting models for clinical deployment.
Xin Ci Wong, Duygu Sarikaya, Kieran Zucker +2
Centre for Doctoral Training in AI for Medical Diagnosis and Care, School of Computing, University of Leeds · School of Computer Science, University of Leeds · Leeds Cancer Centre, St James’s University Hospital, Leeds, UK
Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning. We present a reproducible framework for evaluating uncertainty-aware segmentation under con- trolled clinical degradation. Our experiments use a synthetic multimodal brain tumor MRI cohort generated with a biophysical phantom simulator that follows the BraTS protocol. We train U-Net and Attention U-Net baselines for multi-class tumor sub-region segmentation and augment both models with Monte Carlo dropout to estimate per-voxel uncertainty. Across eight clinically motivated corruption types at five severity levels, we measure segmentation accuracy, calibration, failure detection, and selective prediction coverage. On clean data, Attention U-Net achieves a whole-tumor Dice of 0.990; under severe Gaussian noise, its performance falls to 0.089. Predictive uncertainty rises with degradation and tracks segmentation error (Pearson r = 0.53 under severity-3 Gaussian noise), allowing us to flag failures with an AUROC of 0.843. These results argue for uncertainty-aware inference as a practical safety layer in physician-in-the-loop radiology workflows. We release the code, trained models, and evaluation protocol to support direct reproduction.
Pranav Kaliaperumal, Manisha Kaliaperumal
Department of Computer Science University of Colorado Denver Aurora, CO, USA · Department of Biology (Pre-Medical) Creighton University Omaha, NE, USA
Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases. Here we show that Monte Carlo dropout, applied to a multi-task chest-radiograph classifier (eight thoracic findings, 137,593 training images), provides an epistemic uncertainty signal that tracks generalisation across training-set scales and flags confident yet error-prone predictions. Adding this signal to the point prediction raised error-detection AUROC from 0.74 to 0.77 (ΔAUROC +0.023, 95% CI [+0.014, +0.033]). In a controlled 2x2 factorial experiment, a clinical-decision-support agent exploited this uncertainty only when it was delivered as a binary error-risk flag rather than as raw scores, cutting confident misdiagnoses on unreliable findings from 8.5% to 2.7%. Epistemic uncertainty estimation thus carries decision-relevant information beyond point predictions, but its value for downstream agents depends on how it is communicated.
Frederik Hauke, Patrick Wienholt, Christiane Kuhl +4
Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany. · Department of Medical Oncology, National Center for Tumor Diseases (NCT), Heidelberg University Hospital, Heidelberg, Germany. · Else Kröner Fresenius Center for Digital Health, TU Dresden, Dresden, Germany. +1