Complementary Supervised and Self-Supervised Representations for Out-of-Distribution Graph Learning
Authors: Qingying Hao, Zikang Chen, Chuxuan Hu, Jinyuan Jia, Bo Li, Gang Wang, Carl Gunter
Organizations: ShanghaiTech University Shanghai, China · University of Illinois Urbana-Champaign Urbana, Illinois, USA · The Pennsylvania State University University Park, Pennsylvania, USA
Out-of-distribution (OOD) generalization remains challenging for graph neural networks (GNNs), as graph distributions can vary substantially across time and domains. Supervised and self-supervised graph representation learning are guided by distinct objectives and offer different perspectives on graph representations. In this work, we study whether self-supervised representations (SSL) can provide complementary signals to improve supervised OOD node classification. We develop two backbone-agnostic frameworks that exploit such information at different stages of learning and prediction. Co-Train jointly learns supervised and SSL representations and adaptively integrates them during training, while Dual-Space Retrieval performs non-parametric prediction in the two representation spaces and combines their predictions through confidence-aware fusion at inference time. The supervised and SSL encoders are separately parameterized and need not share the same GNN architecture. We evaluate multiple GNN backbones and two distinct SSL objectives, DGI and GRACE, on four graph benchmarks spanning temporal and cross-domain distribution shifts. Extensive experiments show that Co-Train consistently outperforms strong supervised OOD baselines, while Dual-Space Retrieval achieves competitive performance as a flexible non-parametric alternative. Results across different backbones and SSL objectives, together with representation analyses and ablations, demonstrate that SSL representations provide complementary information to supervised representations and can improve OOD node classification across diverse settings.
Figures & tables
Figure 1 . Overview of Co-Train. The supervised and SSL branches use separate encoders without parameter sharing and are jointly optimized. A learned gate adaptively incorporates the complementary SSL representation into the supervised representation for prediction.
Figure 2 . Overview of the Dual-Space Retrieval framework. Supervised and self-supervised GNNs construct separate embedding memory banks. At test time, neighbor- or centroid-based retrieval produces similarity-weighted prediction scores in each space, which are combined through confidence-aware dual-space score fusion.
Method (Backbone)
T1
T2
T3
T4
T5
T6
T7
T8
T9
ERM (SAGE)
91.74 ± 1.85
85.27 ± 1.39
77.42 ± 1.41
73.64 ± 1.45
75.24 ± 2.88
76.92 ± 4.81
80.68 ± 1.74
65.96 ± 3.97
49.03 ± 0.85
EERM (SAGE)
86.59 ± 2.71
81.03 ± 2.55
76.68 ± 1.39
72.24 ± 1.99
75.20 ± 2.20
78.28 ± 3.33
72.77 ± 4.73
63.83 ± 3.98
49.14 ± 1.01
LiSA (SAGE)
92.58 ± 7.91
85.23 ± 8.41
79.00 ± 4.70
74.45 ± 5.10
74.95 ± 1.75
68.46 ± 4.90
70.42 ± 2.74
55.12 ± 2.09
47.19 ± 0.92
MARIO (SAGE)
90.51 ± 4.12
84.73 ± 2.02
77.62 ± 2.45
72.05 ± 3.42
73.18 ± 4.65
77.35 ± 6.43
63.71 ± 8.36
65.02 ± 7.16
48.79 ± 0.03
DGI (SAGE)
56.39 ± 6.63
53.46 ± 5.17
55.13 ± 5.19
52.81 ± 6.09
59.09 ± 5.01
56.12 ± 4.01
57.98 ± 4.77
50.36 ± 1.54
49.88 ± 1.25
GRACE (SAGE)
90.07 ± 3.14
83.72 ± 1.35
78.11 ± 1.72
72.22 ± 2.13
75.15 ± 3.25
80.58 ± 2.98
77.34 ± 5.95
71.03 ± 4.09
48.70 ± 0.08
Table 1 . F1-score (%) on Elliptic across nine temporal test splits and different GNN backbones.
Method
GCN
GCNII
ES
FR
PTBR
RU
TW
ES
FR
PTBR
RU
TW
ERM
54.64 ± 4.43
52.32 ± 1.03
49.87 ± 6.46
51.29 ± 1.58
51.59 ± 3.34
63.21 ± 0.37
59.42 ± 0.28
61.19 ± 0.39
54.99 ± 0.41
58.19 ± 0.26
EERM
55.63 ± 4.21
53.78 ± 2.48
52.53 ± 6.74
51.87 ± 1.40
52.43 ± 2.83
63.38 ± 0.88
60.32 ± 0.50
61.22 ± 0.94
54.84 ± 0.65
58.12 ± 0.34
LiSA
56.51 ± 4.26
52.43 ± 1.68
56.76 ± 2.86
51.77 ± 1.41
52.11 ± 1.39
55.72 ± 1.34
54.95 ± 2.32
56.12 ± 1.76
51.43 ± 0.79
52.47 ± 0.73
MARIO
62.05 ± 0.94
58.81 ± 0.86
62.69 ± 0.79
53.52 ± 0.82
53.20 ± 0.82
61.78 ± 1.08
57.88 ± 0.72
62.87 ± 0.75
53.54 ± 0.64
53.61 ± 0.34
DGI
54.60 ± 4.31
54.38 ± 2.98
46.73 ± 4.69
50.97 ± 1.04
49.76 ± 2.18
59.65 ± 3.05
59.29 ± 2.21
55.91 ± 6.59
52.37 ± 1.17
51.13 ± 1.69
Table 3. ROC-AUC scores (%) on Twitch across domain splits with different GNN backbones.
Figure 3 . Accuracy comparison of supervised (Sup.), self-supervised (SSL), and Co-Train models on OGB-Arxiv across three temporal test periods.
Figure 4 . ROC-AUC comparsion of supervised (Sup.), self-supervised (SSL), and Retrieval models on five Twitch domains.
Figure 5 . Comparison of representation-separation scores for supervised, self-supervised (SSL), and Co-Train models with on OGB-Arxiv .
Figure 6 . t-SNE visualization of supervised, self-supervised (SSL), and Co-Train test-node representations on the 2018–2020 temporal split of OGB-Arxiv with GRACE.
Domain
Gate
Add.
Concat.
ES
65.88 ± 0.66
65.88 ± 0.26
65.77 ± 0.64
FR
64.89 ± 0.36
64.62 ± 0.49
64.64 ± 0.46
PTBR
64.40 ± 0.63
64.26 ± 0.86
64.03 ± 0.70
RU
56.57 ± 0.29
56.43 ± 0.41
56.49 ± 0.49
TW
59.07 ± 0.52
59.03 ± 0.71
58.82 ± 0.57
Avg.
62.16
62.04
61.95
Table 5 . ROC-AUC (%) of different Co-Train fusion operators across Twitch domains with GCNII.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Distribution Shift
Nodes
Edges
Classes
Metric
Elliptic ( Pareja et al., 2020 )
Temporal Evolution
203,769
234,355
2
F1-Score
OGB-Arxiv ( hu2020ogb )
Temporal Evolution
169,343
1,166,243
40
Accuracy
Twitch-explicit ( Rozemberczki et al., 2021 )
Cross-Domain Transfers
1,912 - 9,498
31,299 - 153,138
2
ROC-AUC
Facebook-100 ( Traud et al., 2012 )
Cross-Domain Transfers
769 - 41,536
16,656 - 1,590,655
2
Accuracy
Appendix
Table 6 . Dataset statistics, distribution shifts, and evaluation metrics.
Method
ES
FR
PTBR
RU
TW
ERM
62.28 ± 0.88
60.06 ± 1.02
60.30 ± 1.20
55.00 ± 0.40
57.55 ± 0.51
EERM
63.64 ± 0.74
62.30 ± 0.63
61.33 ± 1.28
55.51 ± 0.54
55.77 ± 0.38
LiSA
62.96 ± 0.41
61.44 ± 0.68
61.58 ± 1.01
55.25 ± 0.60
56.71 ± 1.45
MARIO
61.11 ± 1.20
59.79 ± 1.90
61.12 ± 0.91
54.77 ± 0.23
54.58 ± 1.34
DGI
61.19 ± 1.57
57.77 ± 0.85
56.59 ± 2.79
52.96 ± 0.63
54.70 ± 1.13
GRACE
63.36 ± 1.01
59.40 ± 1.20
63.95 ± 0.84
53.46 ± 0.74
54.39 ± 0.18
Appendix
Table 7. ROC-AUC scores (%) on Twitch-Explicit across domain splits with the GAT backbone.
Space
K=30
K=100
U
P
Top-1
Top-3
U
P
Top-1
Top-3
Supervised
5.42
43.06%
47.16%
72.83%
9.34
40.58%
47.19%
73.19%
DGI-SSL
10.89
23.31%
35.14%
56.22%
19.22
21.24%
34.88%
57.04%
GRACE-SSL
10.91
21.52%
28.97%
49.59%
18.64
20.52%
29.43%
50.60%
Appendix
Table 8. Neighbor-based retrieval results across supervised and SSL representation spaces on the validation set of OGB-Arxiv .
Branch
Before
After
Δ
GRACE SSL
26.78%
33.48%
+6.70 pp
Supervised
55.46%
60.15%
+4.69 pp
Appendix
Table 9 . Training-node label purity before and after feature smoothing on OGB-Arxiv with a GAT backbone.
Figure 7 . t-SNE visualization of supervised, self-supervised (SSL), and Co-Train test-node representations on the 2018–2020 temporal split of OGB-Arxiv with DGI.
Method
Total Runtime (s)
Fits in GPU Memory
ERM
70.83
Yes
EERM
>2 hours
Yes
LiSA
50.29
Yes
MARIO
206.19
Yes
Co-Train
43.75
Yes
Dual-Space Retrieval
69.16
Yes
Appendix
Table 10. Runtime and memory usage across methods on the OGB-Arxiv dataset using a backbone.
αssl
2014–2016
2016–2018
2018–2020
0
47.75 ± 0.22
47.07 ± 0.22
44.92 ± 0.30
0.01
47.71 ± 0.21
47.02 ± 0.21
45.04 ± 0.18
0.05
47.96 ± 0.23
47.25 ± 0.31
45.18 ± 0.31
0.1
47.91 ± 0.25
47.33 ± 0.29
45.25 ± 0.23
Appendix
Table 11 . Test accuracy (%) under different SSL fusion coefficients αssl on OGB-Arxiv using Retrieval-GRACE with a GAT backbone.
K
2014–2016
2016–2018
2018–2020
10
48.04 ± 0.13
47.32 ± 0.11
45.17 ± 0.20
20
47.83 ± 0.11
47.21 ± 0.13
45.07 ± 0.03
30
47.87 ± 0.08
47.19 ± 0.23
45.13 ± 0.04
40
47.71 ± 0.21
47.16 ± 0.32
45.33 ± 0.46
Appendix
Table 12 . Test accuracy (%) with different numbers of retrieved centroids K on OGB-Arxiv using Retrieval-GRACE with a GAT backbone.
Split
K=30
K=100
K=200
K=300
K=500
Test1
92.64 ± 1.17
92.62 ± 0.87
92.10 ± 0.91
92.34 ± 1.09
92.02 ± 1.32
Test2
84.34 ± 0.52
84.48 ± 0.70
84.29 ± 0.67
84.65 ± 0.77
84.47 ± 0.64
Test3
78.01 ± 1.10
78.27 ± 0.76
77.92 ± 0.64
78.01 ± 0.75
78.09 ± 1.06
Test4
74.44 ± 0.60
74.69 ± 0.41
74.41 ± 0.40
74.44 ± 0.45
74.45 ± 0.35
Test5
72.55 ± 0.72
72.98 ± 0.72
72.74 ± 0.68
72.64 ± 0.70
72.57 ± 0.99
Test6
80.33 ± 0.97
80.44 ± 1.05
80.89 ± 0.73
80.48 ± 1.08
80.54 ± 1.26
Appendix
Table 13. Test-wise sensitivity of Elliptic-SAGE-GRACE to the number of retrieved neighbors K . Each entry reports the mean F1 score (%) and standard deviation over 10 runs.
Test split
SSL-GCN
SSL-SAGE
2014-2016
50.70 ± 0.28
50.29 ± 0.34
2016-2018
49.41 ± 0.32
49.04 ± 0.48
2018-2020
46.71 ± 0.42
46.13 ± 0.50
Appendix
Table 14 . Test accuracy (%) on OGB-Arxiv with different SSL encoder architectures for Co-Train-DGI.
Test split
SSL-GCN
SSL-SAGE
Test1
94.57 ± 0.81
94.81 ± 0.58
Test2
87.17 ± 0.92
87.41 ± 1.09
Test3
79.50 ± 1.29
79.59 ± 1.16
Test4
74.39 ± 0.83
74.12 ± 1.40
Test5
79.05 ± 2.10
78.86 ± 2.00
Test6
84.98 ± 2.18
85.38 ± 1.93
Appendix
Table 15 . F1-score (%) of Co-Train-DGI to the SSL encoder architecture on Elliptic .
Method
Texas
JHU+CIT+AMH
BIN+DUK+PRI
WUSTL+BRD+CMU
ERM
56.25 ± 0.01
52.12 ± 5.52
55.99 ± 0.53
EERM
52.68 ± 3.19
52.05 ± 5.94
53.83 ± 5.33
LiSA
53.01 ± 4.59
55.26 ± 1.48
54.52 ± 3.88
MARIO
50.18 ± 3.80
55.77 ± 0.74
55.72 ± 1.42
DGI
46.46 ± 3.67
48.02 ± 1.38
46.19 ± 2.76
Appendix
Table 16. Accuracy (%) on the Texas of FB-100 under Different Training Graph Combinations.
Graph neural networks are widely used for node classification, but they remain vulnerable to out-of-distribution (OOD) shifts in node features and graph structure. Prior work established that methods trained with standard supervised learning (SL) objectives tend to capture spurious signals from either features and/or structure, leaving the model fragile under distributional changes. To address this, we propose TIDE, a novel and effective Tri-Component Information Decomposition framework that explicitly decomposes information into feature-specific, structure-specific and joint components. TIDE aims to preserve only the label-relevant part of the joint information while filtering out spurious feature- and structure-specific information, thereby enhancing the separation between in-distribution (ID) and OOD nodes. Beyond the framework, we provide theoretical and empirical analyses showing that an information bottleneck objective is preferable to standard SL for graph OOD detection, with higher ID confidence and a greater entropy gap between ID and OOD data. Extensive experiments across seven datasets confirm the efficacy of TIDE, achieving up to a 34% improvement in FPR95 over strong baselines while maintaining competitive ID accuracy.
Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as random graph signals: graph Fourier filters capture structural variation, and a scalar orthogonal-polynomial chaos coordinate captures latent stochastic variation. The resulting doubly-spectral stochastic (DSS) expansion supplies task-matched readouts from one representation: the mean coefficient encodes class evidence for the energy-based OOD score, the higher-order coefficients encode structured logit variation, and quadrature averaging over the chaos coordinate defines the single predictive distribution used for prediction and calibration. A capacity theorem shows that, under a full-rank feature assumption, a restricted subfamily matches the chaos coefficients of any Gaussian-latent random graph signal, with exponentially decaying truncation error under a growth condition; the task-level claims are established empirically. DSS-GNN has two deployment modes: standalone, or as a residual branch beside a deterministic encoder (DSS-Hybrid). Standalone DSS-GNN achieves the lowest Brier score among the compared uncertainty-aware baselines on all 14 node classification benchmarks without post-hoc correction; DSS-Hybrid achieves the best AUROC on most node-OOD settings, competitive cross-graph OOD detection, and the strongest shifted accuracy on all 7 GOOD concept-shift benchmarks under standard empirical risk minimization (ERM). Cross-evaluating both modes on all three tasks shows that each remains effective on the other's tasks, with documented exceptions, and yields explicit deployment guidance.
Fred Xu, Thomas Markovich, Florence Regol +1
Department of Computer Science University of California, Los Angeles · Block, Inc. · Mila, Quebec AI Institute
Node representation learning has advanced rapidly, yet most existing methods rely on per-dataset training and hyperparameter tuning. This dataset-specific optimization comes from the difficulty of designing reusable graph models that generalize across diverse graph datasets. In this work, we introduce Node4All, a node representation learner applicable to arbitrary graph datasets without any dataset-specific optimization. Node4All is built on two complementary ideas. At the architectural level, we introduce the Channel Graph Transformer (CGT), which enables a single fixed parameterization to process arbitrary graph datasets. At the learning level, we propose a self-supervised learning based on a series of synthetic graphs. Together, these components enable generalization beyond individual datasets, which is infeasible with existing architectures and learning frameworks. We extensively evaluate Node4All on node classification across 25 benchmarks against 21 baselines, covering both supervised and self-supervised methods. Despite all baselines being trained and optimized for each dataset, a single Node4All, applied uniformly across the datasets, achieves a competitive ranking of 5th among 21 baselines. Moreover, Node4All supports one-shot and in-context learning with an appropriate predictor and outperforms recent graph foundation models (GFMs) in these settings. These results demonstrate that Node4All not only achieves reusability across arbitrary graph datasets, but also remains an effective solution in practice. Code and model checkpoints are available in https://github.com/dooho00/node4all.
Dooho Lee, Jaemin Yoo
KAIST · Daejeon, Republic of Korea · Seoul National University +1