While graph neural networks (GNNs) have shown substantial promise in connectome-based diagnostic classification, deterministic models inevitably suppress pipeline-induced noise and model ambiguities, yielding overconfident predictions. Although uncertainty quantification (UQ) is widely adopted in voxel-level segmentation, its role in connectomic graph learning remains largely unaddressed. This paper presents a comprehensive narrative review of UQ frameworks tailored to connectome graph learning alongside an empirical case study demonstrating the perils of uncalibrated predictions. We delineate sources of aleatoric and epistemic uncertainty across neuroimaging pipelines and review prominent UQ paradigms, from Bayesian approximations and ensemble methods to evidential learning and conformal prediction. In our case study, a temporal Graph Attention Network (GAT) trained on dynamic functional connectivity (dFC) matrices from the SUDMEX CONN dataset achieves 80.0% diagnostic accuracy (F1 = 0.794) for Cocaine Use Disorder. However, a post-hoc uncertainty audit via Monte Carlo dropout reveals severe overconfidence (ECE = 0.127), with misclassified subjects assigned prediction confidences up to 95%. This empirical divergence between discrimination and calibration underscores the confidence paradox in deep connectomics. Our findings establish that rigorous UQ, calibration, and selective prediction mechanisms are indispensable for deploying trustworthy graph-based biomarkers in clinical neuroscience.
Figures & tables
Figure 1: Conceptual simulation of aleatoric and epistemic uncertainty. Persistent noise around the latent process represents aleatoric uncertainty. The widening predictive interval beyond the training boundary ( x=7 ) illustrates epistemic uncertainty due to extrapolation.
Figure 2: Bayesian view of predictive uncertainty: predictions are averaged over a posterior of plausible parameter configurations p(θ∣D) . Model disagreement signifies epistemic uncertainty.
Figure 3: MC dropout estimation: stochastic masks generate an ensemble of sub-networks, where variance quantifies model uncertainty.
Figure 4: Calibration regimes in predictive modeling: (a) overconfidence, where assigned probabilities outpace true accuracy; (b) ideal calibration, adhering to the diagonal identity; and (c) underconfidence, where the network is more accurate than its output probabilities indicate.
Figure 5: Comprehensive calibration assessment and mitigation framework: (a) Reliability diagram depicting the gap between binned confidence and empirical accuracy (ECE); (b) Brier score formulation for penalty-based probabilistic alignment; and (c) Temperature scaling mechanism for post-hoc logit entropy adjustment.
Characteristic
CUD ( n=74 )
HC ( n=64 )
Age, years; mean ± SD
30.60±8.26
30.99±7.25
Sex, male/female
65/9
51/11
AMAI score; mean ± SD
133.45±50.63
107.96±50.64
Table 1: Demographic and clinical characteristics of the SUDMEX CONN cohort.
Hyperparameter / Component
Value / Specification
Evaluation scheme
5-Fold Stratified Cross-Validation
GNN layer type
Graph Attention Network (GATConv)
Attention heads ( h )
8
Input node features ( F )
2 ( d(v) and dnn(v) )
Hidden / output dimensions
64 / 128
Dropout rate ( p )
0.70
Table 2: Specification of network architecture and optimization hyperparameters.
Fold
Best Epoch
Accuracy
Precision
Recall
F1 Score
Fold 1
35
0.7917
0.8111
0.7917
0.7822
Fold 2
24
0.7917
0.8529
0.7917
0.7822
Fold 3
26
0.9167
0.9286
0.9167
0.9161
Fold 4
30
0.7500
0.7600
0.7500
0.7450
Fold 5
28
0.7500
0.7600
0.7500
0.7450
Mean ± SD
—
0.800±0.068
0.835±0.065
0.800±0.068
0.794±0.072
Table 3: Per-fold and aggregate classification performance of the temporal GAT model under 5-fold stratified cross-validation. Metrics are computed on the held-out validation split of each fold; bold marks the best-performing fold. Mean and standard deviation are computed over the five folds.
Figure 6: Five-fold cross-validation learning curves of the temporal GAT model trained on the SUDMEX CONN connectome dataset. Left: cross-entropy loss trajectories for training (solid) and validation (dashed) sets across all five folds, together with fold-averaged curves (black: mean train; dark red: mean validation ± one standard deviation shaded). All folds exhibit rapid loss decrease within the first five epochs, with training loss approaching zero by epoch 35, while validation loss stabilises between 0.45 and 0.65 , consistent with the applied dropout regularization ( p=0.70 ) and early stopping. Right: corresponding accuracy curves. Mean validation accuracy (dark red) converges to approximately 0.80 , with Fold 3 (teal dashed) reaching the highest validation accuracy ( 0.917 ) at epoch 26. The gap between mean training accuracy (black, ≈1.0 ) and mean validation accuracy reflects the high dropout rate used to support Monte-Carlo Dropout uncertainty quantification at inference time.
Figure 7: Uncertainty quantification and calibration audit of the temporal GNN on SUDMEX CONN connectomes. (a) Reliability diagram showing systematic overconfidence in dominant high-confidence bins, resulting in an expected calibration error of ECE=0.127 . (b) Posterior confidence density distributions across correct predictions (teal) versus misclassifications (red), highlighting that errors are sharply concentrated in the high-confidence regime ( pmax>0.80 ). (c) Decomposition of normalized uncertainty into total predictive entropy, expected entropy (aleatoric), and mutual information (epistemic) for correct versus incorrect classifications. (d) Subject-level confidence paradox map (predictive confidence vs. predictive entropy), identifying high-certainty diagnostic failures (Subjects #10 and #11) residing within the critical risk region.
Uncertainty quantification (UQ) in graph neural networks (GNNs) is crucial in high-stakes domains but remains a significant challenge. In graph settings, message passing often relies on strong assumptions such as exchangeability, which are rarely satisfied in practice, and achieving reliable UQ typically requires costly resampling or post-hoc calibration. To address these issues, we introduce Quantile-free Prediction Interval GNN (QpiGNN), a framework that builds on quantile regression (QR) to enable GNN-based UQ by directly optimizing coverage and interval width without requiring quantile inputs or post-processing. QpiGNN employs a dual-head architecture that decouples prediction and uncertainty, and is trained with label-only supervision through a quantile-free joint loss. This design allows efficient training and yields robust prediction intervals, with theoretical guarantees of asymptotic coverage and near-optimal width under mild assumptions. Experiments on 19 synthetic and real-world benchmarks show QpiGNN achieves average 22% higher coverage and 50% narrower intervals than baselines, while ensuring efficiency and robustness to noise and structural shifts.
Soyoung park, Hwanjun Song, Sungsu Lim
Department of Computer Science and Engineering, Chungnam National University, Daejeon, Korea · Department of Industrial and Systems Engineering, KAIST, Daejeon, Korea
While deep ensembles are widely considered to be the default method for uncertainty quantification in deep learning, their effectiveness for graph-structured data is often simply assumed based on successes in domains like computer vision. We investigate standard deep ensembles specifically for message-passing graph neural networks. Benchmarking across seven datasets representing varied tasks and complexities, we reveal that ensembles provide surprisingly little improvement over a single model. Instead, the observed marginal gains stem primarily from stabilizing optimization noise in point predictions rather than yielding meaningfully better uncertainty estimates. Through an aleatoric-epistemic decomposition, we identify epistemic collapse: independently trained networks consistently converge to overly similar predictions. Because disagreement is the fundamental mechanism through which ensembles capture epistemic uncertainty, this lack of diversity neutralizes their key advantage. Analyzing this phenomenon further, we suggest this collapse is driven by functional rather than weight-space convexity, where distinct parameter solutions induce almost identical behavior. Our results suggest that deep ensemble success does not seamlessly transfer to graph machine learning.
Pedro C. Vieira, Pedro Ribeiro, Viacheslav Borovitskiy
University of Edinburgh · DCC/FCUP, University of Porto
Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications. In a series of experiments, we contrast Deep Ensembles, Bayesian Neural Networks, and Monte Carlo-dropout methods. We assess their ability to quantify uncertainty in different scenarios, accounting for common dataset characteristics in two genomic application areas and modalities: sequence-to-activity models, and single-cell expression analysis. Our systematic comparison framework provides guidelines for the applicability and reliability of UQ methods in genomics, highlighting their strengths and limitations in different scenarios. We show that Bayesian Neural Networks are better at capturing uncertainty caused by strong class imbalance and out-of-distribution data in genomics, despite their computational disadvantages. Moreover, we show how uncertainty scores can be used to select high-quality predictions in protein-RNA interactions.
Sepideh Saran, Mahsa Ghanbari, Uwe Ohler
The Berlin Institute for Medical Systems Biology, Max Delbrück Center for Molecular Medicine, Germany · Department of Electrical Engineering and Computer Science, Technical University of Berlin, Germany · Department of Biology, Humboldt University of Berlin, Germany +1