Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification
Authors: Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep Kaushik, Marion Smits, Stefan Klein
Organizations: Department of Radiology and Nuclear Medicine, Erasmus MC, University Medical Center Rotterdam, Rotterdam, the Netherlands · Department of Medical Informatics, Amsterdam UMC, University of Amsterdam, Amsterdam, the Netherlands · GE HealthCare, Munich, Germany · GE HealthCare, USA · Brain tumor Centre, Erasmus MC Cancer Institute, Rotterdam, the Netherlands · Medical Delta, Delft, the Netherlands
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.
Figures & tables
Figure 1: Overview of the multi-task glioma subtyping framework with uncertainty quantification. The input data consist of four structural MRI sequences per patient, namely T1-weighted (T1w), contrast-enhanced T1-weighted (T1wCE), T2-weighted (T2w), and fluid-attenuated inversion recovery (FLAIR). The sequences are preprocessed using registration, bias field correction, and skull stripping. Subsequently, the data are used to train a Deep Learning model that jointly performs tumor segmentation and prediction of IDH mutation status, 1p/19q co-deletion status, and tumor grade. For each task, predictive, aleatoric, and epistemic uncertainty are computed. For tumor segmentation, voxel-wise uncertainty maps are converted into case-level uncertainty scores using four aggregation strategies: brain-level aggregation, predicted-tumor aggregation, dilated-tumor aggregation, and distance-weighted aggregation around the tumor region. In the tumor features panel, the colors of the icons vary from tumors associated with poorer prognosis (red) to those with better prognosis (green).
Set
Dataset
Number of cases
Total
Train
BraTS
156
1466
Brain tumor Progression
20
CPTAC-GBM
45
EGD
775
IvyGAP
39
In-house
322
Table 1: Train and test data distribution
Metric
Purpose
Possible Values
Interpretation
U-AUC
Measures the probability that a randomly chosen error case has higher uncertainty than a correct case.
0 – 1
Values close to 1 indicate uncertainty correctly ranks errors above successes.
AP
Evaluates precision-recall trade-off for predicting errors, suitable for imbalanced tasks.
Quantifies error-identification performance relative to baseline prevalence.
0−ϵ1
Values >1 indicate that uncertainty successfully prioritizes error cases. A value of 1 indicates random selection.
AURC
Evaluates the mean residual risk among retained cases as coverage varies from most certain to all cases.
0 – 1
Lower values indicate that uncertainty successfully retains lower-risk cases at reduced coverage.
Table 2: Summary of the metrics used to assess operational utility of uncertainty estimates, namely, Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC). For each metric, the purpose, possible values, and interpretation are presented.
Figure 2: First three columns from left to right: predictive, aleatoric and epistemic uncertainty value heatmaps across all tasks as a function of the number of MCD samples T and the dropout rate δ . The value at each cell of the heatmap represents the mean uncertainty across all cases of the test set. The last column presents the average value for each type of uncertainty over all experiments with the blue, green and red bars representing predictive, aleatoric and epistemic uncertainty, respectively. Uncertainty values for classification tasks were normalized at the same scale to ease the analysis of the patterns. For the segmentation task the heatmaps were kept in the original scale.
Figure 3: Expected Calibration Error (ECE), Negative Log-Likelihood (NLL) and Receiver Operating Characteristic Area Under the Curve (AUC) values as a function of dropout rate δ . For each point in every graph, a 95% confidence interval is presented. This interval was obtained by 1000× bootstrap resampling of the test set.
Figure 4: Tumor segmentation uncertainty quality analysis. The top panel presents the Pearson correlation coefficient between case-level predictive uncertainty and Dice Similarity Coefficient (DSC) as a function of the dropout rate δ used for the MCD ensemble. Correlations are shown for four spatial aggregation strategies: brain-mask aggregation, predicted-tumor aggregation, 5-voxel dilated predicted-tumor aggregation, and boundary-weighted aggregation using σ=5 voxels. More negative correlations indicate that higher case-level uncertainty is associated with lower segmentation accuracy. The bottom panel presents the mean DSC between the predicted tumor segmentation and the ground truth in the test set as a function of the dropout rate δ .
Task
Class distribution
Accuracy ↑
Balanced accuracy ↑
Macro F1 score ↑
AUC ↑
IDH
Wildtype: 129 Mutated: 85
0.82 [0.77,0.87]
0.79 [0.73,0.84]
0.80 [0.74,0.85]
0.90 [0.85,0.94]
1p/19q
Intact: 205 Co-deleted: 25
0.90 [0.85,0.93]
0.56 [0.50,0.63]
0.57 [0.47,0.67]
0.85 [0.77,0.91]
Grade
Grade 2: 45 Grade 3: 58 Grade 4: 132
0.71 [0.64,0.77]
0.60 [0.57,0.64]
0.51 [0.47,0.54]
0.85 [0.81,0.88]
Table 3: Predictive performance at the selected MCD configuration ( T=30 , δ=0.25 ) for the three classification tasks. For each task, the class distribution is also reported. Values are shown with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. AUC is reported at task level: for tumor grade, it corresponds to the macro-averaged one-vs-rest AUC across grade classes. Total test set size differs across tasks because molecular and histopathological annotations were not available for all patients. Arrows in metric column headers indicate that higher values are preferable ( ↑ ).
Figure 5: Violin plots of the predictive uncertainty across different tasks, comparing the distributions of high quality and low quality predictions, obtained with the optimal configuration of the MCD model at T=30 and δ=0.25 . For the classification tasks, high and low quality predictions correspond to correctly and incorrectly classified cases, respectively. For the tumor segmentation task, high and low quality predictions represent cases with DSC>0.7 and DSC≤0.7 , respectively. For all tasks, the difference between the distributions for both groups was statistically significant (Mann–Whitney U test).
Task
Error rate ↓
Uncertainty
Predictive
Aleatoric
Epistemic
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation status prediction
0.18
0.73 [0.62,0.82]
0.42 [0.28,0.57]
2.37 [1.72,3.34]
0.10 [0.05,0.15]
0.71 [0.61,0.80]
0.32 [0.22,0.46]
1.79 [1.39,2.54]
0.10 [0.05,0.15]
0.71 [0.61,0.80]
0.34 [0.23,0.50]
1.89 [1.45,2.80]
0.10 [0.05,0.16]
1p/19q co-deletion status prediction
0.10
0.84 [0.76,0.91]
0.37 [0.22,0.58]
3.55 [2.37,6.07]
0.03 [0.01,0.05]
0.84 [0.77,0.91]
0.39 [0.22,0.58]
3.70 [2.51,6.19]
0.03 [0.01,0.04]
0.68 [0.58,0.77]
0.16 [0.10,0.29]
1.54 [1.22,2.59]
0.05 [0.03,0.08]
Tumor grade prediction
0.29
0.78 [0.72,0.84]
0.59 [0.47,0.71]
2.00 [1.66,2.47]
0.12 [0.09,0.17]
0.79 [0.73,0.85]
0.59 [0.47,0.72]
2.00 [1.68,2.50]
0.12 [0.09,0.17]
0.70 [0.62,0.77]
0.45 [0.35,0.58]
1.52 [1.28,1.94]
0.17 [0.12,0.24]
Tumor segmentation
0.17
0.95 [0.92,0.98]
0.82 [0.72,0.91]
4.95 [3.89,6.67]
0.13 [0.12,0.14]
0.96 [0.94,0.98]
0.83 [0.72,0.91]
5.01 [3.93,6.80]
0.13 [0.12,0.14]
0.95 [0.92,0.98]
0.81 [0.71,0.90]
4.90 [3.85,6.64]
0.13 [0.12,0.14]
Table 4: Operational utility of uncertainty estimates per task. Assessment is made using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC), derived from using uncertainty as a predictor of errors. Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. The numerically highest values for U-AUC , AP, and Lift, and the numerically lowest values for AURC , are indicated in bold for each task. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Task
Method
Error ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation
DE
0.17
0.72 [0.62,0.82]
0.37 [0.24,0.56]
2.15 [1.57,3.23]
0.10 [0.05,0.15]
MCD
0.18
0.73 [0.62,0.82]
0.42 [0.28,0.57]
2.37 [1.72,3.34]
0.10 [0.05,0.15]
MCDE
0.17
0.73 [0.63,0.83]
0.34 [0.23,0.50]
2.03 [1.53,2.98]
0.09 [0.05,0.14]
1p/19q co-deletion
DE
0.10
0.84 [0.76,0.91]
0.39 [0.22,0.59]
3.90 [2.53,6.62]
0.03 [0.01,0.04]
MCD
0.10
0.84 [0.76,0.91]
0.37 [0.22,0.58]
3.55 [2.37,6.07]
0.03 [0.01,0.05]
MCDE
0.09
0.81 [0.74,0.89]
0.26 [0.16,0.45]
2.84 [2.06,4.91]
0.03 [0.01,0.04]
Table 5: Comparison of uncertainty quantification methods using predictive uncertainty. Assessment is made using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC), derived from using predictive uncertainty as a predictor of errors. Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. For tumor segmentation, predictive uncertainty was aggregated within the predicted tumor region. The numerically highest values for U-AUC , AP, and Lift, and the numerically lowest values for AURC , are indicated in bold for each task. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Figure 6: Distribution of tumor segmentation Dice Similarity Coefficient (DSC) scores for correctly and incorrectly classified cases for the IDH mutation status, 1p/19q co-deletion status and tumor grade. Cases are stratified according to whether the corresponding prediction was correct or incorrect. For each task, higher DSC values indicate better agreement between predicted and reference tumor segmentations. Statistical significance between groups was found for the IDH mutation status and tumor grade prediction (Mann–Whitney U test).
Task
Uncertainty
Predictive
Aleatoric
Epistemic
IDH mutation status prediction
0.48 [0.38, 0.57]
0.51 [0.42, 0.58]
0.09 [-0.04, 0.21]
1p/19q co-deletion status prediction
0.22 [0.12, 0.33]
0.24 [0.15, 0.33]
−0.15 [-0.25, -0.05]
Tumor grade prediction
0.53 [0.45, 0.61]
0.50 [0.42, 0.57]
0.10 [-0.03, 0.26]
Table 6: Pearson correlation ( ρ ) between tumor segmentation uncertainty and classification uncertainty for each task. Correlations were computed separately for predictive, aleatoric, and epistemic uncertainty components. Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. The numerically highest correlation coefficient per task is indicated in bold.
Task
Error rate ↓
Tumor segmentation uncertainty as a predictor of classification errors
Predictive
Aleatoric
Epistemic
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation status prediction
0.18
0.62 [0.51,0.72]
0.30 [0.21,0.46]
1.68 [1.27,2.60]
0.15 [0.08,0.22]
0.65 [0.55,0.75]
0.28 [0.20,0.40]
1.54 [1.22,2.23]
0.13 [0.07,0.20]
0.61 [0.50,0.72]
0.32 [0.22,0.46]
1.79 [1.30,2.72]
0.15 [0.09,0.23]
1p/19q co-deletion status prediction
0.10
0.68 [0.57,0.78]
0.23 [0.12,0.41]
2.16 [1.38,4.15]
0.06 [0.03,0.11]
0.69 [0.57,0.79]
0.27 [0.14,0.43]
2.55 [1.50,4.56]
0.06 [0.03,0.11]
0.67 [0.56,0.78]
0.24 [0.12,0.40]
2.26 [1.36,4.19]
0.06 [0.03,0.11]
Tumor grade prediction
0.29
0.65 [0.57,0.73]
0.44 [0.35,0.58]
1.52 [1.25,1.94]
0.22 [0.15,0.30]
0.67 [0.59,0.74]
0.42 [0.34,0.55]
1.45 [1.23,1.82]
0.19 [0.14,0.26]
0.65 [0.56,0.73]
0.45 [0.35,0.58]
1.54 [1.27,1.95]
0.22 [0.16,0.30]
Table 7: Ability of tumor segmentation uncertainty to predict classification errors. Assessment is performed using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC). Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. The numerically highest values for U-AUC , AP, and Lift, and the numerically lowest values for AURC , are indicated in bold for each task. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Task
Error rate ↓
Trust score
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation status prediction
0.18
0.73 [0.64,0.81]
0.38 [0.26,0.53]
2.12 [1.58,3.10]
0.09 [0.05,0.14]
1p/19q co-deletion status prediction
0.10
0.85 [0.77,0.91]
0.42 [0.25,0.58]
3.98 [2.71,6.39]
0.02 [0.01,0.04]
Tumor grade prediction
0.29
0.77 [0.71,0.83]
0.51 [0.41,0.65]
1.77 [1.48,2.23]
0.13 [0.09,0.17]
Table 8: Performance of the proposed trust score for predicting classification errors. Assessment is performed using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve (AURC). Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Overview of the multi-task DL network used for the tumor segmentation and prediction of IDH mutation status, 1p/19q co-deletion status and tumor grade. The architecture processes full 3D volumes. The labels in the figure define the layer type, and the numbers indicate either the number of filters, dense units or features. Figure based on the one presented in the work of van der Voort et al. (2023) .
τDSC
Error rate ↓
Uncertainty
U-AUC ↑
AP ↑
Lift ↑
Predictive
0.96 [0.93,0.98]
0.78 [0.64,0.90]
7.96 [5.73,12.41]
0.60
0.10
Aleatoric
0.97 [0.95,0.99]
0.82 [0.67,0.92]
8.33 [5.99,13.09]
Epistemic
0.96 [0.93,0.98]
0.76 [0.61,0.88]
7.77 [5.61,12.26]
Predictive
0.95 [0.92,0.98]
0.82 [0.72,0.91]
4.95 [3.89,6.67]
0.70
0.17
Aleatoric
0.96 [0.94,0.98]
0.83 [0.72,0.91]
5.01 [3.93,6.80]
Epistemic
0.95 [0.92,0.98]
0.81 [0.71,0.90]
4.90 [3.85,6.64]
Appendix
Table 9: Sensitivity of segmentation uncertainty error detection to the DSC threshold used to define segmentation error. For each threshold, segmentation error was defined as DSC≤τDSC . Error rate is reported as the proportion of cases classified as segmentation errors. U-AUC , AP, and Lift are reported for predictive, aleatoric, and epistemic uncertainty, with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Figure 8: Case-wise association between predictive segmentation uncertainty and tumor segmentation DSC at the selected MCD configuration ( T=30 , δ=0.25 ). Results are shown for brain mask, predicted tumor, 5-voxel dilated tumor, and boundary-weighted aggregation ( σ=5 voxels). Each point represents one test case; Pearson ρ and Spearman ρs are reported for each strategy.
Dataset
Subset
IDH mutation status
1p/19q co-deletion status
Tumor grade
Wildtype
Mutated
N/A
Intact
Co-deleted
N/A
2
3
4
N/A
Train
BraTS
0
0
156
0
0
156
0
0
0
156
Brain tumor Progression
0
0
20
0
0
20
0
0
0
20
CPTAC-GBM
0
0
45
0
0
45
0
0
0
45
EGD
312
155
308
186
73
516
135
80
502
58
IvyGAP
32
6
1
27
3
9
0
1
36
2
Appendix
Table 10: Detailed distribution of the data used for training and testing the Deep Learning (DL) architecture for each one of the classification tasks. N/A (not available) represents missing data. Tumor segmentation ground truth was available for all cases.
Task
Method
Error ↓
Aleatoric uncertainty
Epistemic uncertainty
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
U-AUC ↑
AP ↑
Lift ↑
AURC ↓
IDH mutation
DE
0.17
0.73 [0.62,0.82]
0.39 [0.26,0.56]
2.26 [1.66,3.31]
0.09 [0.05,0.15]
0.64 [0.54,0.74]
0.28 [0.17,0.43]
1.60 [1.20,2.45]
0.12 [0.07,0.18]
MCD
0.18
0.71 [0.61,0.80]
0.32 [0.22,0.46]
1.79 [1.39,2.54]
0.10 [0.05,0.15]
0.71 [0.61,0.80]
0.34 [0.23,0.50]
1.89 [1.45,2.80]
0.10 [0.05,0.16]
MCDE
0.17
0.73 [0.62,0.82]
0.36 [0.24,0.53]
2.12 [1.55,3.13]
0.09 [0.05,0.14]
0.64 [0.54,0.73]
0.30 [0.20,0.47]
1.81 [1.31,2.83]
0.12 [0.07,0.19]
1p/19q co-deletion
DE
0.10
0.83 [0.75,0.90]
0.29 [0.18,0.49]
2.95 [2.12,5.25]
0.03 [0.01,0.04]
0.66 [0.55,0.76]
0.18 [0.10,0.34]
1.79 [1.19,3.65]
0.05 [0.03,0.08]
MCD
0.10
0.84 [0.77,0.91]
0.39 [0.22,0.58]
3.70 [2.51,6.19]
0.03 [0.01,0.04]
0.68 [0.58,0.77]
0.16 [0.10,0.29]
1.54 [1.22,2.59]
0.05 [0.03,0.08]
Appendix
Table 11: Comparison of Uncertainty Quantification methods using aleatoric and epistemic uncertainty. Assessment is made using Uncertainty-based Receiver Operating Characteristic Area Under the Curve (U-AUC), Average Precision (AP), Lift, and Area Under the Risk–Coverage Curve ( AURC ), derived from using each uncertainty component as a predictor of errors. Values are reported with 95% confidence intervals obtained by 1000× bootstrap resampling of the test set. The error column denotes the method-specific classification error rate for classification tasks and the proportion of cases with DSC ≤0.70 for tumor segmentation. For the latter, uncertainty was aggregated within the predicted tumor region. The highest displayed values for U-AUC, AP, and Lift, and the lowest displayed values for AURC , are indicated in bold for each task and uncertainty component. Arrows in metric column headers indicate whether higher ( ↑ ) or lower ( ↓ ) values are preferable.
Centre for Doctoral Training in AI for Medical Diagnosis and Care, School of Computing, University of Leeds · School of Computer Science, University of Leeds · Leeds Cancer Centre, St James’s University Hospital, Leeds, UK
Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany. · Department of Medical Oncology, National Center for Tumor Diseases (NCT), Heidelberg University Hospital, Heidelberg, Germany. · Else Kröner Fresenius Center for Digital Health, TU Dresden, Dresden, Germany. +1