Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as random graph signals: graph Fourier filters capture structural variation, and a scalar orthogonal-polynomial chaos coordinate captures latent stochastic variation. The resulting doubly-spectral stochastic (DSS) expansion supplies task-matched readouts from one representation: the mean coefficient encodes class evidence for the energy-based OOD score, the higher-order coefficients encode structured logit variation, and quadrature averaging over the chaos coordinate defines the single predictive distribution used for prediction and calibration. A capacity theorem shows that, under a full-rank feature assumption, a restricted subfamily matches the chaos coefficients of any Gaussian-latent random graph signal, with exponentially decaying truncation error under a growth condition; the task-level claims are established empirically. DSS-GNN has two deployment modes: standalone, or as a residual branch beside a deterministic encoder (DSS-Hybrid). Standalone DSS-GNN achieves the lowest Brier score among the compared uncertainty-aware baselines on all 14 node classification benchmarks without post-hoc correction; DSS-Hybrid achieves the best AUROC on most node-OOD settings, competitive cross-graph OOD detection, and the strongest shifted accuracy on all 7 GOOD concept-shift benchmarks under standard empirical risk minimization (ERM). Cross-evaluating both modes on all three tasks shows that each remains effective on the other's tasks, with documented exceptions, and yields explicit deployment guidance.
Figures & tables
Figure 1: (a) Second-order correction ci (Proposition 2 ) on Minesweeper. Each node is positioned by loss curvature ( tr(Hi) ) and chaos energy ( Ei ); color shows correction magnitude. (b) Nodes binned by correction quartile: the correction concentrates on hard nodes (low accuracy, high Brier).
Accuracy (%)
Brier Score
Dataset
DSS-GNN
TFE-GNN
G- Δ UQ
DSS-GNN
TFE-GNN
G- Δ UQ
Cora
86.63 ± 1.26
71.25 ± 1.90
84.35 ± 1.97
0.207 ± 0.019
0.449 ± 0.024
0.226 ± 0.023
Citeseer
80.20 ± 1.28
68.87 ± 1.87
70.85 ± 2.22
0.308 ± 0.010
0.463 ± 0.020
0.433 ± 0.026
PubMed
89.72 ± 0.31
82.84 ± 0.76
88.18 ± 0.66
0.156 ± 0.005
0.265 ± 0.011
0.179 ± 0.008
Texas
91.31 ± 3.44
84.92 ± 4.26
9.02 ± 1.83
0.240 ± 0.111
0.255 ± 0.083
0.779 ± 0.029
Cornell
85.11 ± 5.12
81.28 ± 5.45
21.06 ± 3.49
0.231 ± 0.079
0.290 ± 0.069
0.793 ± 0.023
Table 1: Node classification: test accuracy (%, ↑ ) and Brier score ( ↓ ). Mean ± standard deviation over 10 runs. Best result per dataset in bold . DSS-GNN uses the per-dataset configuration of Table 12 and the mean-logit readout.
Model
Cora
Amazon-Photo
Coauthor-CS
Cross-graph
Structure
Feature
Label
Structure
Feature
Label
Structure
Feature
Label
Twitch
Arxiv
GNNSafe
87.52
93.44
92.80
99.58
98.55
97.35
99.60
99.64
97.23
66.82
71.06
GNNSafe++
90.62
95.56
92.75
99.82
99.64
97.51
99.99
99.97
97.89
95.36
74.77
Graph-EBM
61.14
72.42
92.69
75.22
86.49
97.24
72.75
89.28
97.91
44.43
52.80
MC-dropout
87.50
93.12
93.03
98.67
98.50
96.80
99.51
99.53
97.25
68.58
66.34
Deep Ensemble
87.87
93.63
93.75
98.58
98.43
97.35
98.18
98.45
94.97
72.25
67.72
Table 2: OOD detection (AUROC, %) on the nine node-OOD settings and the two cross-graph settings (AUPR, FPR95, and ID accuracy for the latter are in Table 3 ). GNNSafe/GNNSafe++ and GPN [ 30 ] values as reported by Wu et al. [36] on this benchmark (their Tables 1–2; their appendix lists 81.77/93.24 for GPN on Coauthor-CS feature/label; GPN out-of-memory on Arxiv); Graph-EBM [ 9 ] , MC-dropout ( M=20 ), and Deep Ensembles ( M=5 ) run in our pipeline on the same GCN backbone (3 seeds; the latter two with the same score propagation). Bold: best.
Twitch
Arxiv
Model
AUROC
AUPR
FPR95
ID Acc.
AUROC
AUPR
FPR95
ID Acc.
GNNSafe
66.82
70.97
76.24
70.40
71.06
80.44
87.01
53.39
GNNSafe++
95.36
97.12
33.57
70.18
74.77
83.21
77.43
53.50
Graph-EBM
44.43
57.91
94.84
63.04
52.80
65.88
97.40
53.45
MC-dropout
68.58
81.25
94.84
64.10
66.34
74.88
89.69
53.56
Deep Ensemble
72.25
83.19
94.29
66.94
67.72
75.98
87.32
47.58
Table 3: Cross-graph OOD (AUROC, AUPR: ↑ ; FPR95, false positive rate at 95% true positive rate: ↓ ; ID: in-distribution; bold: best). GNNSafe/GNNSafe++ and GPN values as reported by Wu et al. [36] (OOM: out of memory on a 24 GB GPU); Graph-EBM [ 9 ] , MC-dropout, and Deep Ensembles run in our pipeline (3 seeds).
Model
GOOD-CBAS
GOOD-WebKB
GOOD-Twitch
GOOD-Cora
GOOD-Cora
GOOD-Arxiv
GOOD-Arxiv
color
university
language
word
degree
time
degree
ERM
82.43
27.16
51.59
64.03
60.30
65.64
54.81
IRM
82.00
26.06
49.78
63.93
60.26
65.54
56.72
VREx
82.86
26.61
55.75
64.03
60.53
65.92
56.68
Coral
81.57
28.07
51.80
64.04
60.30
65.79
55.14
DANN
83.57
29.36
51.67
63.96
60.23
65.67
55.34
Table 4: Test accuracy (%) on the concept-shift GOOD node benchmark for DSS-Hybrid (standard ERM training) and 13 baselines; ERM through TAR as reported by Zheng et al. [38] , G- Δ UQ run in our pipeline. Best in bold, second-best underlined; OOM: out of memory.
Task
Standalone DSS-GNN
DSS-Hybrid
Calibration (14 datasets)
Best Brier on all 14 and best accuracy on 13 (Table 1 ); stronger than the hybrid on 13 of 14
Never loses beyond noise to its own GCN base; Roman-Empire +27.7 accuracy, −0.33 Brier
OOD detection (11 settings)
Effective on 10 of 11: Cora 79.8/88.2/94.0 , Photo 96.5 – 98.1 , CS 94.3 – 97.8 , Arxiv 73.7 (margin objective); the Twitch energy score does not separate
Best published AUROC on 7 of 11 (Table 2 ); stronger than the standalone on 10 of 11
GOOD concept shift (7 settings)
Above every published baseline on 5 of 7, at ERM level on the other 2
Strongest on all 7 (Table 4 )
Table 5: Both deployment modes on all three tasks. Each cell summarizes the full per-dataset comparison in Appendix G (Tables 24 , 25 , and 26 ); the standalone OOD and GOOD cells and the hybrid calibration cell are the cross-evaluations; the other cells restate Tables 1 , 2 – 3 , and 4 .
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Calibration
OOD Detection
Robust Classif.
MC-dropout [ 10 ]
(✓)
✗
✗
Deep Ensembles [ 17 ]
✓
(✓)
✗
G- Δ UQ [ 33 ]
✓
✓ ∗
✓
TFE-GNN [ 8 ]
(✓)
✗
✗
GNNSafe [ 36 ]
✗
✓
✗
GNNSafe++ [ 36 ]
✗
✓
✗
Appendix
Table 6: Tasks evaluated by each method. ✓= evaluated in the original paper; (✓) = evaluated as a secondary metric or in the appendix; ✓ ∗ = evaluated in the original paper but OOM on our benchmark; ✗= not evaluated in the original paper or not designed for that task.
Table 8: Node classification datasets for calibration experiments (Section 4.2 ). Results use 10 splits per dataset (Appendix D ); edges are counted in both directions. Cornell uses the corrected files of the geom-gcn repository, whose initial release duplicated the Texas features and labels. Datasets below the mid-rule are from the heterophilous suite of Platonov et al. [26] .
Dataset
OOD type
Nodes
Edges
Classes
Features
Cora
structure / feature / label
2,708
10,556
7
1,433
Amazon-Photo
structure / feature / label
7,650
238,162
8
745
Coauthor-CS
structure / feature / label
18,333
163,788
15
6,805
Twitch (DE → ES/FR/RU)
cross-graph
9,498
306,276
2
3,170
Arxiv ( ≤ 2015 → 2018–20)
cross-graph (temporal)
169,343
2,315,598
40
128
Appendix
Table 9: OOD detection datasets (Section 4.3 ). For node-OOD (Cora, Amazon-Photo, Coauthor-CS), OOD nodes are generated via structure (stochastic block model), feature (random interpolation of node features), or label (held-out classes) perturbations. For cross-graph OOD, the ID and OOD graphs are from different domains. Statistics are for the in-distribution graph.
Setting
Domain
Nodes
Edges
Classes
Features
GOOD-Cora / degree
node degree
19,793
126,842
70
8,710
GOOD-Cora / word
word diversity
19,793
126,842
70
8,710
GOOD-Arxiv / degree
node degree
169,343
2,315,598
40
128
GOOD-Arxiv / time
publication year
169,343
2,315,598
40
128
GOOD-CBAS / color
node color
700
3,962
4
4
GOOD-WebKB / university
university
617
1,138
5
1,703
Appendix
Table 10: GOOD benchmark settings (Section 4.4 ). All use concept shift. The domain column indicates the covariate that defines the distribution shift. Node counts are for the full graph before splitting; GOOD-Cora is built on the full Cora citation graph (CoraFull), not on the 7-class Planetoid subset of Table 8 .
Dataset
Architecture
Hidden
L
Optimizer
Overrides
Cora
Prop-first (rw)
64
2
Adam
dropout 0.8
Citeseer
Cheb
64
2
Adam
none
PubMed
Cheb
128
2
Adam
none
Texas
Prop-first
64
2
RMSprop
none
Cornell
Cheb
128
2
Adam
dropout 0.8
Wisconsin
Cheb
64
2
Adam
dropout 0.6
Appendix
Table 11: Per-dataset non- P DSS-GNN configuration for the calibration experiments (Table 1 ). “Prop-first” = propagate-once structure (spectral filtering followed by MLP); “TFE-adj” = TFE-style adjacency propagation; “Cheb” = Chebyshev filters on rescaled Laplacian. Unlisted hyperparameters use the shared defaults above; P , λreg , and S per dataset are listed in Table 12 .
Dataset
P
λreg
S
Accuracy
Brier
Val. Brier
Cora
1
0.01
4
86.63 ± 1.26
0.207 ± 0.019
0.203
Citeseer
2
0
4
80.20 ± 1.28
0.308 ± 0.010
0.315
PubMed
5
0.01
4
89.72 ± 0.31
0.156 ± 0.005
0.152
Texas
2
0.01
4
91.31 ± 3.44
0.240 ± 0.111
0.190
Cornell
1
0.01
4
85.11 ± 5.12
0.231 ± 0.079
0.166
Wisconsin
2
0
4
93.25 ± 3.39
0.104 ± 0.047
0.088
Appendix
Table 12: Selected configuration per dataset for the calibration experiments (Table 1 ): chaos order P , chaos energy regularization λreg , quadrature size S , and that single configuration’s test accuracy (%) and Brier score (mean ± SD over the same 10 splits). The last column reports the validation Brier of the selected configuration, recorded in a separate run on the same 10 splits (for Cornell, in the same run).
Dataset
P=0
P=1
P=2
P=3
P=4
P=5
P=6
P=8
Cora
85.12 ± 1.06
86.63 ± 1.26
86.33 ± 1.12
86.33 ± 0.82
86.63 ± 1.14
86.32 ± 1.13
86.61 ± 0.84
86.38 ± 0.96
Citeseer
65.39 ± 1.02
79.21 ± 0.79
80.20 ± 1.28
78.47 ± 0.80
71.50 ± 4.58
78.92 ± 1.13
76.77 ± 2.19
78.49 ± 1.08
PubMed
89.12 ± 0.57
89.33 ± 0.49
89.33 ± 0.66
89.38 ± 0.67
89.33 ± 0.54
89.72 ± 0.31
89.40 ± 0.57
89.39 ± 0.51
Texas
75.08 ± 23.70
88.36 ± 4.57
91.31 ± 3.44
89.18 ± 4.92
91.31 ± 3.30
87.54 ± 5.59
89.01 ± 2.82
87.37 ± 5.43
Cornell
83.62 ± 6.80
85.11 ± 5.12
85.74 ± 6.73
83.62 ± 5.39
84.26 ± 6.87
83.83 ± 7.38
83.62 ± 6.53
82.98 ± 7.61
Wisconsin
93.25 ± 3.46
92.87 ± 2.92
93.25 ± 3.39
88.50 ± 3.35
93.00 ± 3.46
90.75 ± 2.84
89.63 ± 4.82
90.50 ± 2.92
Appendix
Table 13: Accuracy (%) of DSS-GNN for varying chaos order P across all datasets.
Dataset
P=0
P=1
P=2
P=3
P=4
P=5
P=6
P=8
Cora
0.228 ± 0.013
0.207 ± 0.019
0.212 ± 0.016
0.209 ± 0.016
0.208 ± 0.019
0.215 ± 0.015
0.211 ± 0.011
0.211 ± 0.015
Citeseer
0.427 ± 0.008
0.346 ± 0.008
0.308 ± 0.010
0.326 ± 0.012
0.422 ± 0.027
0.327 ± 0.012
0.370 ± 0.023
0.331 ± 0.013
PubMed
0.162 ± 0.007
0.158 ± 0.007
0.160 ± 0.008
0.159 ± 0.008
0.161 ± 0.007
0.156 ± 0.005
0.159 ± 0.008
0.159 ± 0.007
Texas
0.269 ± 0.082
0.213 ± 0.111
0.240 ± 0.111
0.233 ± 0.096
0.253 ± 0.096
0.253 ± 0.086
0.249 ± 0.114
0.242 ± 0.082
Cornell
0.241 ± 0.087
0.231 ± 0.079
0.228 ± 0.088
0.250 ± 0.090
0.263 ± 0.129
0.247 ± 0.107
0.252 ± 0.079
0.273 ± 0.111
Wisconsin
0.103 ± 0.048
0.107 ± 0.041
0.104 ± 0.047
0.161 ± 0.060
0.103 ± 0.048
0.129 ± 0.037
0.153 ± 0.052
0.135 ± 0.040
Appendix
Table 14: Brier Score ( ↓ ) of DSS-GNN for varying chaos order P across all datasets.
Setting
P=0
P=1
P=2
GOOD-Cora / word
63.24 ± 0.40
64.89 ± 0.21
64.01 ± 0.25
GOOD-Cora / degree
61.38 ± 0.81
62.60 ± 0.59
61.22 ± 0.07
GOOD-Arxiv / time
64.81 ± 0.27
65.43 ± 1.07
64.48 ± 0.27
Appendix
Table 15: Shifted test accuracy (%) of standalone DSS-GNN on GOOD concept shift as the chaos order varies at the validation-selected configuration (mean ± SD over 3 seeds).
Dataset
Structure
Feature
Label
Cora
+14.5
+5.4
−6.7
Amazon-Photo
−0.4
+1.7
+3.9
Coauthor-CS
+1.9
+1.1
+1.3
Twitch (cross-graph)
+18.2
Arxiv (cross-graph)
+7.3
Appendix
Table 16: Effect of energy-score propagation on DSS-Hybrid: change in AUROC (points) from the raw to the propagated energy score at identical checkpoints (3 runs, fixed configuration).
Dataset
λreg=0
10−3
10−2
10−1
Cora
0.219
0.219
0.207
0.220
Citeseer
0.308
0.310
0.318
0.499
PubMed
0.159
0.156
0.156
0.157
Texas
0.366
0.246
0.213
0.251
Cornell
0.230
0.237
0.228
0.342
Wisconsin
0.103
0.107
0.108
0.174
Appendix
Table 17: Effect of chaos energy regularization λreg on Brier score ( ↓ ), with P=2 fixed. Bold indicates the best λreg per dataset.
Dataset
S=2
S=3
S=4
S=6
S=8
Cora
0.216
0.207
0.217
0.209
0.212
Citeseer
0.310
0.308
0.312
0.313
0.311
PubMed
0.157
0.156
0.156
0.158
0.156
Texas
0.216
0.221
0.213
0.217
0.218
Cornell
0.235
0.286
0.228
0.218
0.242
Wisconsin
0.142
0.109
0.115
0.103
0.108
Appendix
Table 18: Brier score ( ↓ ) vs. number of quadrature nodes S , with P=2 fixed. The theoretical minimum is S=3 .
Dataset
ChebNet K=4
P=0
Best P>0
Δ Chaos
Cora
0.281
0.228
0.207
−9.2%
Citeseer
0.470
0.427
0.308
−27.9%
PubMed
0.171
0.162
0.156
−3.7%
Texas
0.208
0.269
0.213
−20.8%
Cornell
0.324
0.241
0.228
−5.4%
Wisconsin
0.073
0.103
0.103
−0%
Appendix
Table 19: Contribution decomposition: Brier score ( ↓ ). ChebNet K=4 is a vanilla single-filter diagnostic; P=0 and Best P>0 are the P=0 column and the best P>0 column of Table 14 (same per-dataset architecture). Boldface compares the controlled P=0 versus Best P>0 pair, and the rightmost column isolates the chaos contribution. Table 1 reports the selected order of Table 12 , which can differ from the best swept order.
Dataset
P=0
P=1
P=2
P=3
Ratio P=2 / P=0
Cora (2.7K nodes)
35
67
102
131
2.9×
Amazon-Rat. (25K nodes)
464
526
576
646
1.2×
ogbn-arxiv (169K nodes)
2002
2088
2174
2273
1.1×
ogbn-arxiv peak memory (MiB)
1721
2813
3976
5141
2.3×
Appendix
Table 20: Wall-clock time per epoch (ms) vs. chaos order P (Cora and Amazon-Ratings on an A100; ogbn-arxiv on one H100, 3 runs, with peak GPU memory in MiB).
Component
Comparison
Effect
Chaos expansion ( P )
P=0 vs best P>0 , same architecture (Tables 13 , 14 , 19 )
Calibration: Citeseer Brier −27.9% and accuracy +14.8 points (the endpoint also beats plain GCN, 0.308 vs 0.318 Brier, Tables 1 / 23 ); Texas −20.8% ; Minesweeper −11.1%
Chaos expansion ( P )
OOD ablation varying only P (Appendix E.2 )
Energy AUROC flat across P (Cora-structure 91.0/91.2/90.9/90.8 at one fixed configuration); detection reads the mean logit, so P can be selected for calibration without changing detection
Chaos expansion ( P )
GOOD ablation at the selected configurations (Table 15 )
Shifted accuracy: P=1 beats P=0 on GOOD-Cora word ( +1.7 ), degree ( +1.2 ), and GOOD-Arxiv time ( +0.6 )
Spectral backbone (chaos off)
P=0 vs GCN (Tables 13 , 23 )
Heterophilous accuracy comes from the spectral backbone: Roman-Empire 79.0 at P=0 vs GCN 51.3
Filtering design vs single filter
Per-dataset architecture at P=0 vs single-branch ChebNet (Table 19 )
Lower or equal Brier for the full design on 12 of 14 datasets (equal on Questions; Chameleon 0.362 vs 0.776 , Squirrel 0.416 vs 0.798 , Roman-Empire 0.298 vs 0.368 ); exceptions are Texas and Wisconsin
Table 21: Component-wise evidence. Each row names one component, the comparison that isolates it, and the measured effect.
Dataset
DSS-GNN
MC-dropout
Deep Ensemble
Plain GCN
Cora
0.207 (86.63)
0.2188 (87.36)
0.2134 (87.45)
0.2125 (87.47)
Citeseer
0.308 (80.20)
0.3230 (80.03)
0.3170 (79.77)
0.3176 (80.00)
PubMed
0.156 (89.72)
0.2176 (86.20)
0.2153 (86.16)
0.2155 (86.10)
Texas
0.240 (91.31)
0.5862 (64.43)
0.5686 (62.79)
0.5758 (64.43)
Cornell
0.231 (85.11)
0.6829 (52.77)
0.6826 (52.98)
0.6839 (52.34)
Wisconsin
0.104 (93.25)
0.6370 (50.38)
0.6373 (51.25)
0.6435 (50.25)
Appendix
Table 22: Brier score ( ↓ ) with test accuracy (%) in parentheses for DSS-GNN (Table 1 ) and stochastic GCN baselines under the same 10-split protocol (means over 10 splits).
Accuracy (%)
Brier Score
Dataset
DSS-GNN
GCN
GAT
DSS-GNN
GCN
GAT
Cora
86.63 ± 1.26
87.45 ± 1.61
87.75 ± 1.71
0.207 ± 0.019
0.213 ± 0.014
0.222 ± 0.015
Citeseer
80.20 ± 1.28
79.84 ± 0.96
80.57 ± 1.12
0.308 ± 0.010
0.318 ± 0.007
0.353 ± 0.007
PubMed
89.72 ± 0.31
86.21 ± 0.31
85.72 ± 0.28
0.156 ± 0.005
0.215 ± 0.005
0.232 ± 0.004
Texas
91.31 ± 3.44
64.43 ± 6.80
71.80 ± 7.06
0.240 ± 0.111
0.576 ± 0.031
0.555 ± 0.038
Cornell
85.11 ± 5.12
52.34 ± 9.99
49.57 ± 12.63
0.231 ± 0.079
0.684 ± 0.080
0.631 ± 0.073
Appendix
Table 23: Node classification: test accuracy (%, ↑ ) and Brier score ( ↓ ) for DSS-GNN vs. standard GNN baselines. Mean ± SD over 10 splits. Best per dataset in bold .
Standalone DSS-GNN
DSS-Hybrid
GCN base
Dataset
Accuracy
Brier
Accuracy
Brier
Accuracy
Brier
Cora
86.63 ± 1.26
0.207 ± 0.019
84.33 ± 0.97
0.241 ± 0.015
84.20 ± 1.02
0.243 ± 0.016
Citeseer
80.20 ± 1.28
0.308 ± 0.010
74.71 ± 1.05
0.375 ± 0.009
74.32 ± 1.13
0.378 ± 0.009
PubMed
89.72 ± 0.31
0.156 ± 0.005
88.27 ± 0.56
0.177 ± 0.006
87.93 ± 0.56
0.181 ± 0.006
Texas
91.31 ± 3.44
0.240 ± 0.111
45.57 ± 7.68
0.737 ± 0.034
46.39 ± 7.55
0.737 ± 0.034
Cornell
85.11 ± 5.12
0.231 ± 0.079
55.11 ± 14.03
0.656 ± 0.142
47.66 ± 12.01
0.727 ± 0.068
Appendix
Table 24: Test accuracy (%) and Brier score, mean ± SD over the same 10 splits. The standalone column is Table 1 ; the hybrid and its GCN base are a matched pair (identical training, budgets, and splits; the only difference is the DSS residual).
Setting
Standalone AUROC
Standalone ID acc.
Hybrid AUROC
Hybrid ID acc.
Objective (standalone)
Cora / structure
79.79 ± 1.46 ( + BN)
69.50
94.32
n/r ∗
standard (no OOD exposure)
Cora / feature
88.15 ± 1.02 ( + BN)
69.17
97.60
n/r ∗
standard (no OOD exposure)
Cora / label
94.03 ± 1.09 ( + BN)
89.24
94.11
n/r ∗
standard (no OOD exposure)
Amazon-Photo / structure
96.53 ± 1.41 ( + BN)
91.59
99.69
n/r ∗
standard (no OOD exposure)
Amazon-Photo / feature
98.10 ± 0.11 ( + BN)
92.00
99.66
n/r ∗
standard (no OOD exposure)
Amazon-Photo / label
96.53 ± 0.10 ( + BN)
94.81
97.52
n/r ∗
standard (no OOD exposure)
Appendix
Table 25: AUROC (%) and ID test accuracy (%) on the 11 OOD settings. Standalone: 3 runs, mean ± SD; hybrid: Tables 2 – 3 (AUROC only for the nine node-OOD settings ∗ ). The last column gives the standalone training objective.
Setting
Standalone
Hybrid
Best baseline
GOOD-CBAS / color
82.38 ± 13.17 ( + BN)
88.57
TAR 87.29
GOOD-WebKB / university
36.41 ± 1.60 ( + BN)
40.77
TAR 30.83
GOOD-Twitch / language
60.28 ± 0.51 ( + BN)
61.21
TAR 57.20
GOOD-Cora / word
64.89 ± 0.21 ( + BN)
64.81
TAR 64.73
GOOD-Cora / degree
62.60 ± 0.59 ( + BN)
62.73
TAR 61.73
GOOD-Arxiv / time
65.43 ± 1.07 ( + BN)
66.76
TAR 66.08
Appendix
Table 26: Shifted (OOD) test accuracy (%) on the GOOD concept-shift settings. Standalone: mean ± SD over 3 seeds, configuration selected on the GOOD OOD-validation split (Appendix G.1 ); hybrid and best baseline from Table 4 .
Dataset
∣Δ Brier ∣
Disagree (%)
Dataset
∣Δ Brier ∣
Disagree (%)
Cora
0.0002
0.03
Roman-Emp.
0.0000
0.00
Citeseer
0.0000
0.00
Amz-Rat.
0.0000
0.00
PubMed
0.0000
0.02
Minesweeper
0.0003
0.21
Texas
0.0021
0.00
Tolokers
0.0000
0.00
Cornell
0.0000
0.00
Questions
0.0000
0.00
Wisconsin
0.0000
0.00
CS
0.0001
0.00
Appendix
Table 27: Quadrature readout versus mean-logit readout at identical checkpoints under the full 10-split protocol of Table 1 : absolute Brier difference and the fraction of test nodes on which argmaxpˉi and argmaxZi,0 disagree (mean over 10 splits). Accuracy differs by at most 0.04 points on every dataset. In hybrid mode the disagreement is 0.0000 on all 14 datasets.
Dataset
BatchNorm off (control)
Table 1
BatchNorm on
Cora
86.44 ± 1.56 / 0.216 ± 0.022
86.63 ± 1.26 / 0.207 ± 0.019
85.78 ± 1.76 / 0.220 ± 0.017
Citeseer
80.27 ± 1.18 / 0.307 ± 0.010
80.20 ± 1.28 / 0.308 ± 0.010
67.03 ± 1.85 / 0.480 ± 0.024
Wisconsin
93.62 ± 1.72 / 0.109 ± 0.029
93.25 ± 3.39 / 0.104 ± 0.047
91.75 ± 4.75 / 0.134 ± 0.061
Roman-Empire
78.82 ± 1.08 / 0.300 ± 0.012
78.84 ± 0.61 / 0.300 ± 0.007
78.60 ± 0.64 / 0.300 ± 0.009
Appendix
Table 28: Test accuracy (%) / Brier, mean ± SD over 10 splits, under the protocol of Table 1 at one calibration configuration per dataset, without and with per-chaos-channel BatchNorm.
Setting (covariate split)
ERM GCN
DSS-Hybrid (ERM)
GOOD-CBAS / color
49.05
53.33
GOOD-WebKB / university
13.49
29.10
GOOD-Twitch / language
42.78
52.64
GOOD-Cora / word
61.84
64.18
GOOD-Cora / degree
56.66
54.94
GOOD-Arxiv / time
44.80 ∗
70.39
Appendix
Table 29: Shifted test accuracy (%) under GOOD covariate shift: ERM-trained DSS-Hybrid (GCN encoder, validation-loss-selected configuration) versus the same-protocol ERM GCN. ∗ ERM baseline not converged within the shared budget (see text).
Symbol
Meaning
ω
the single shared standard Gaussian latent; never sampled, integrated out by quadrature
Ψn
normalized Hermite polynomial of order n ; orthonormal basis of L2(R,γ)
P
chaos truncation order; n,m=0,…,P index input/output chaos orders
Hn(ℓ)∈RN×dℓ
order- n chaos coefficient of the layer- ℓ embeddings ( 2 ); the embedding field is H(ℓ)(ω)=∑nHn(ℓ)Ψn(ω)
LG , Tk
rescaled Laplacian 2LG/λmax−I and Chebyshev polynomial of degree k (Section 2 )
Glp,Ghp
learned low/high-pass filter polynomials of degrees Klp,Khp with coefficients ck(lp),ck(hp) ( 3 )
Appendix
Table 30: Notation for the DSS layer, Eqs. ( 2 )–( 6 ).
Uncertainty quantification has become an important factor in understanding the data representations produced by Graph Neural Networks (GNNs). Despite their predictive capabilities being ever useful across industrial workspaces, the inherent uncertainty induced by the nature of the data is a huge mitigating factor to GNN performance. While aleatoric uncertainty is the result of noisy and incomplete stochastic data such as missing edges or over-smoothing, epistemic uncertainty arises from lack of knowledge about a system or model (e.g., a graph's topology or node feature representation), which can be reduced by gathering more data and information. In this paper, we propose an original new framework in which node-level epistemic uncertainty is modelled in a belief function (finite random set) formalism. The resulting Random-Set Graph Neural Networks have a belief-function head predicting a random set over the list of classes, from which both a precise probability prediction and a measure of epistemic uncertainty can be obtained. Extensive experiments on 9 different graph learning datasets, including real-world autonomous driving benchmarks as such Nuscene and ROAD, demonstrate RS-GNN's superior uncertainty quantification capabilities
Tommy Woodley, Shireen Kudukkil Manchingal, Matteo Tolloso +2
School of Engineering, Computing and Mathematics Oxford Brookes University · Department of Computer Science University of Pisa · Oxford Brookes Institute for Artificial Intelligence, Data Analysis and Systems (AIDAS)
Hypergraph neural networks have shown powerful capability in modeling higher-order relations, yet their predictive uncertainty remains underexplored. Unlike pairwise graphs, uncertainty in hypergraphs arises not only from noisy attributes and ambiguous labels, but also from variations in node-hyperedge incidence structures and complex higher-order dependencies. Existing approaches mainly estimate uncertainty from final predictions or rely on computationally expensive ensembles and Bayesian inference, limiting their ability to capture uncertainty evolution during representation learning. In this paper, we propose Hypergraph Neural Stochastic Diffusion(HyperNSD), a stochastic differential equation framework for uncertainty estimation on hypergraphs. HyperNSD models hypergraph representations as stochastic processes evolving over node-hyperedge incidence structures. A learnable drift function captures deterministic higher-order diffusion dynamics, while a learnable stochastic forcing function characterizes structural ambiguity and representation noise. Predictive uncertainty is directly quantified through the variability of stochastic representation trajectories, providing an intrinsic uncertainty measure beyond post-hoc confidence scores. We formulate HyperNSD with neural drift and diffusion networks, enabling joint learning of prediction and uncertainty propagation. Theoretical analyses establish well posedness, perturbation stability,permutation equivariance, and numerical convergence of the proposed stochastic dynamics. Experiments on multiple hypergraph benchmarks demonstrate that HyperNSD achieves reliable uncertainty estimation for out-of-distribution and misclassification detection while preserving competitive prediction accuracy. These results provide a principled stochastic-dynamical framework for trustworthy higher-order representation learning.
Out-of-distribution (OOD) generalization remains challenging for graph neural networks (GNNs), as graph distributions can vary substantially across time and domains. Supervised and self-supervised graph representation learning are guided by distinct objectives and offer different perspectives on graph representations. In this work, we study whether self-supervised representations (SSL) can provide complementary signals to improve supervised OOD node classification. We develop two backbone-agnostic frameworks that exploit such information at different stages of learning and prediction. Co-Train jointly learns supervised and SSL representations and adaptively integrates them during training, while Dual-Space Retrieval performs non-parametric prediction in the two representation spaces and combines their predictions through confidence-aware fusion at inference time. The supervised and SSL encoders are separately parameterized and need not share the same GNN architecture. We evaluate multiple GNN backbones and two distinct SSL objectives, DGI and GRACE, on four graph benchmarks spanning temporal and cross-domain distribution shifts. Extensive experiments show that Co-Train consistently outperforms strong supervised OOD baselines, while Dual-Space Retrieval achieves competitive performance as a flexible non-parametric alternative. Results across different backbones and SSL objectives, together with representation analyses and ablations, demonstrate that SSL representations provide complementary information to supervised representations and can improve OOD node classification across diverse settings.
Qingying Hao, Zikang Chen, Chuxuan Hu +4
ShanghaiTech University Shanghai, China · University of Illinois Urbana-Champaign Urbana, Illinois, USA · The Pennsylvania State University University Park, Pennsylvania, USA