Complementary Supervised and Self-Supervised Representations for Out-of-Distribution Graph Learning
Authors: Qingying Hao, Zikang Chen, Chuxuan Hu, Jinyuan Jia, Bo Li, Gang Wang, Carl Gunter
Organizations: ShanghaiTech University Shanghai, China · University of Illinois Urbana-Champaign Urbana, Illinois, USA · The Pennsylvania State University University Park, Pennsylvania, USA
Out-of-distribution (OOD) generalization remains challenging for graph neural networks (GNNs), as graph distributions can vary substantially across time and domains. Supervised and self-supervised graph representation learning are guided by distinct objectives and offer different perspectives on graph representations. In this work, we study whether self-supervised representations (SSL) can provide complementary signals to improve supervised OOD node classification. We develop two backbone-agnostic frameworks that exploit such information at different stages of learning and prediction. Co-Train jointly learns supervised and SSL representations and adaptively integrates them during training, while Dual-Space Retrieval performs non-parametric prediction in the two representation spaces and combines their predictions through confidence-aware fusion at inference time. The supervised and SSL encoders are separately parameterized and need not share the same GNN architecture. We evaluate multiple GNN backbones and two distinct SSL objectives, DGI and GRACE, on four graph benchmarks spanning temporal and cross-domain distribution shifts. Extensive experiments show that Co-Train consistently outperforms strong supervised OOD baselines, while Dual-Space Retrieval achieves competitive performance as a flexible non-parametric alternative. Results across different backbones and SSL objectives, together with representation analyses and ablations, demonstrate that SSL representations provide complementary information to supervised representations and can improve OOD node classification across diverse settings.
Figures & tables
Figure 1 . Overview of Co-Train. The supervised and SSL branches use separate encoders without parameter sharing and are jointly optimized. A learned gate adaptively incorporates the complementary SSL representation into the supervised representation for prediction.
Figure 2 . Overview of the Dual-Space Retrieval framework. Supervised and self-supervised GNNs construct separate embedding memory banks. At test time, neighbor- or centroid-based retrieval produces similarity-weighted prediction scores in each space, which are combined through confidence-aware dual-space score fusion.
Method (Backbone)
T1
T2
T3
T4
T5
T6
T7
T8
T9
ERM (SAGE)
91.74 ± 1.85
85.27 ± 1.39
77.42 ± 1.41
73.64 ± 1.45
75.24 ± 2.88
76.92 ± 4.81
80.68 ± 1.74
65.96 ± 3.97
49.03 ± 0.85
EERM (SAGE)
86.59 ± 2.71
81.03 ± 2.55
76.68 ± 1.39
72.24 ± 1.99
75.20 ± 2.20
78.28 ± 3.33
72.77 ± 4.73
63.83 ± 3.98
49.14 ± 1.01
LiSA (SAGE)
92.58 ± 7.91
85.23 ± 8.41
79.00 ± 4.70
74.45 ± 5.10
74.95 ± 1.75
68.46 ± 4.90
70.42 ± 2.74
55.12 ± 2.09
47.19 ± 0.92
MARIO (SAGE)
90.51 ± 4.12
84.73 ± 2.02
77.62 ± 2.45
72.05 ± 3.42
73.18 ± 4.65
77.35 ± 6.43
63.71 ± 8.36
65.02 ± 7.16
48.79 ± 0.03
DGI (SAGE)
56.39 ± 6.63
53.46 ± 5.17
55.13 ± 5.19
52.81 ± 6.09
59.09 ± 5.01
56.12 ± 4.01
57.98 ± 4.77
50.36 ± 1.54
49.88 ± 1.25
GRACE (SAGE)
90.07 ± 3.14
83.72 ± 1.35
78.11 ± 1.72
72.22 ± 2.13
75.15 ± 3.25
80.58 ± 2.98
77.34 ± 5.95
71.03 ± 4.09
48.70 ± 0.08
Table 1 . F1-score (%) on Elliptic across nine temporal test splits and different GNN backbones.
Method
GCN
GCNII
ES
FR
PTBR
RU
TW
ES
FR
PTBR
RU
TW
ERM
54.64 ± 4.43
52.32 ± 1.03
49.87 ± 6.46
51.29 ± 1.58
51.59 ± 3.34
63.21 ± 0.37
59.42 ± 0.28
61.19 ± 0.39
54.99 ± 0.41
58.19 ± 0.26
EERM
55.63 ± 4.21
53.78 ± 2.48
52.53 ± 6.74
51.87 ± 1.40
52.43 ± 2.83
63.38 ± 0.88
60.32 ± 0.50
61.22 ± 0.94
54.84 ± 0.65
58.12 ± 0.34
LiSA
56.51 ± 4.26
52.43 ± 1.68
56.76 ± 2.86
51.77 ± 1.41
52.11 ± 1.39
55.72 ± 1.34
54.95 ± 2.32
56.12 ± 1.76
51.43 ± 0.79
52.47 ± 0.73
MARIO
62.05 ± 0.94
58.81 ± 0.86
62.69 ± 0.79
53.52 ± 0.82
53.20 ± 0.82
61.78 ± 1.08
57.88 ± 0.72
62.87 ± 0.75
53.54 ± 0.64
53.61 ± 0.34
DGI
54.60 ± 4.31
54.38 ± 2.98
46.73 ± 4.69
50.97 ± 1.04
49.76 ± 2.18
59.65 ± 3.05
59.29 ± 2.21
55.91 ± 6.59
52.37 ± 1.17
51.13 ± 1.69
Table 3. ROC-AUC scores (%) on Twitch across domain splits with different GNN backbones.
Figure 3 . Accuracy comparison of supervised (Sup.), self-supervised (SSL), and Co-Train models on OGB-Arxiv across three temporal test periods.
Figure 4 . ROC-AUC comparsion of supervised (Sup.), self-supervised (SSL), and Retrieval models on five Twitch domains.
Figure 5 . Comparison of representation-separation scores for supervised, self-supervised (SSL), and Co-Train models with on OGB-Arxiv .
Figure 6 . t-SNE visualization of supervised, self-supervised (SSL), and Co-Train test-node representations on the 2018–2020 temporal split of OGB-Arxiv with GRACE.
Domain
Gate
Add.
Concat.
ES
65.88 ± 0.66
65.88 ± 0.26
65.77 ± 0.64
FR
64.89 ± 0.36
64.62 ± 0.49
64.64 ± 0.46
PTBR
64.40 ± 0.63
64.26 ± 0.86
64.03 ± 0.70
RU
56.57 ± 0.29
56.43 ± 0.41
56.49 ± 0.49
TW
59.07 ± 0.52
59.03 ± 0.71
58.82 ± 0.57
Avg.
62.16
62.04
61.95
Table 5 . ROC-AUC (%) of different Co-Train fusion operators across Twitch domains with GCNII.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Distribution Shift
Nodes
Edges
Classes
Metric
Elliptic ( Pareja et al., 2020 )
Temporal Evolution
203,769
234,355
2
F1-Score
OGB-Arxiv ( hu2020ogb )
Temporal Evolution
169,343
1,166,243
40
Accuracy
Twitch-explicit ( Rozemberczki et al., 2021 )
Cross-Domain Transfers
1,912 - 9,498
31,299 - 153,138
2
ROC-AUC
Facebook-100 ( Traud et al., 2012 )
Cross-Domain Transfers
769 - 41,536
16,656 - 1,590,655
2
Accuracy
Appendix
Table 6 . Dataset statistics, distribution shifts, and evaluation metrics.
Method
ES
FR
PTBR
RU
TW
ERM
62.28 ± 0.88
60.06 ± 1.02
60.30 ± 1.20
55.00 ± 0.40
57.55 ± 0.51
EERM
63.64 ± 0.74
62.30 ± 0.63
61.33 ± 1.28
55.51 ± 0.54
55.77 ± 0.38
LiSA
62.96 ± 0.41
61.44 ± 0.68
61.58 ± 1.01
55.25 ± 0.60
56.71 ± 1.45
MARIO
61.11 ± 1.20
59.79 ± 1.90
61.12 ± 0.91
54.77 ± 0.23
54.58 ± 1.34
DGI
61.19 ± 1.57
57.77 ± 0.85
56.59 ± 2.79
52.96 ± 0.63
54.70 ± 1.13
GRACE
63.36 ± 1.01
59.40 ± 1.20
63.95 ± 0.84
53.46 ± 0.74
54.39 ± 0.18
Appendix
Table 7. ROC-AUC scores (%) on Twitch-Explicit across domain splits with the GAT backbone.
Space
K=30
K=100
U
P
Top-1
Top-3
U
P
Top-1
Top-3
Supervised
5.42
43.06%
47.16%
72.83%
9.34
40.58%
47.19%
73.19%
DGI-SSL
10.89
23.31%
35.14%
56.22%
19.22
21.24%
34.88%
57.04%
GRACE-SSL
10.91
21.52%
28.97%
49.59%
18.64
20.52%
29.43%
50.60%
Appendix
Table 8. Neighbor-based retrieval results across supervised and SSL representation spaces on the validation set of OGB-Arxiv .
Branch
Before
After
Δ
GRACE SSL
26.78%
33.48%
+6.70 pp
Supervised
55.46%
60.15%
+4.69 pp
Appendix
Table 9 . Training-node label purity before and after feature smoothing on OGB-Arxiv with a GAT backbone.
Figure 7 . t-SNE visualization of supervised, self-supervised (SSL), and Co-Train test-node representations on the 2018–2020 temporal split of OGB-Arxiv with DGI.
Method
Total Runtime (s)
Fits in GPU Memory
ERM
70.83
Yes
EERM
>2 hours
Yes
LiSA
50.29
Yes
MARIO
206.19
Yes
Co-Train
43.75
Yes
Dual-Space Retrieval
69.16
Yes
Appendix
Table 10. Runtime and memory usage across methods on the OGB-Arxiv dataset using a backbone.
αssl
2014–2016
2016–2018
2018–2020
0
47.75 ± 0.22
47.07 ± 0.22
44.92 ± 0.30
0.01
47.71 ± 0.21
47.02 ± 0.21
45.04 ± 0.18
0.05
47.96 ± 0.23
47.25 ± 0.31
45.18 ± 0.31
0.1
47.91 ± 0.25
47.33 ± 0.29
45.25 ± 0.23
Appendix
Table 11 . Test accuracy (%) under different SSL fusion coefficients αssl on OGB-Arxiv using Retrieval-GRACE with a GAT backbone.
K
2014–2016
2016–2018
2018–2020
10
48.04 ± 0.13
47.32 ± 0.11
45.17 ± 0.20
20
47.83 ± 0.11
47.21 ± 0.13
45.07 ± 0.03
30
47.87 ± 0.08
47.19 ± 0.23
45.13 ± 0.04
40
47.71 ± 0.21
47.16 ± 0.32
45.33 ± 0.46
Appendix
Table 12 . Test accuracy (%) with different numbers of retrieved centroids K on OGB-Arxiv using Retrieval-GRACE with a GAT backbone.
Split
K=30
K=100
K=200
K=300
K=500
Test1
92.64 ± 1.17
92.62 ± 0.87
92.10 ± 0.91
92.34 ± 1.09
92.02 ± 1.32
Test2
84.34 ± 0.52
84.48 ± 0.70
84.29 ± 0.67
84.65 ± 0.77
84.47 ± 0.64
Test3
78.01 ± 1.10
78.27 ± 0.76
77.92 ± 0.64
78.01 ± 0.75
78.09 ± 1.06
Test4
74.44 ± 0.60
74.69 ± 0.41
74.41 ± 0.40
74.44 ± 0.45
74.45 ± 0.35
Test5
72.55 ± 0.72
72.98 ± 0.72
72.74 ± 0.68
72.64 ± 0.70
72.57 ± 0.99
Test6
80.33 ± 0.97
80.44 ± 1.05
80.89 ± 0.73
80.48 ± 1.08
80.54 ± 1.26
Appendix
Table 13. Test-wise sensitivity of Elliptic-SAGE-GRACE to the number of retrieved neighbors K . Each entry reports the mean F1 score (%) and standard deviation over 10 runs.
Test split
SSL-GCN
SSL-SAGE
2014-2016
50.70 ± 0.28
50.29 ± 0.34
2016-2018
49.41 ± 0.32
49.04 ± 0.48
2018-2020
46.71 ± 0.42
46.13 ± 0.50
Appendix
Table 14 . Test accuracy (%) on OGB-Arxiv with different SSL encoder architectures for Co-Train-DGI.
Test split
SSL-GCN
SSL-SAGE
Test1
94.57 ± 0.81
94.81 ± 0.58
Test2
87.17 ± 0.92
87.41 ± 1.09
Test3
79.50 ± 1.29
79.59 ± 1.16
Test4
74.39 ± 0.83
74.12 ± 1.40
Test5
79.05 ± 2.10
78.86 ± 2.00
Test6
84.98 ± 2.18
85.38 ± 1.93
Appendix
Table 15 . F1-score (%) of Co-Train-DGI to the SSL encoder architecture on Elliptic .
Method
Texas
JHU+CIT+AMH
BIN+DUK+PRI
WUSTL+BRD+CMU
ERM
56.25 ± 0.01
52.12 ± 5.52
55.99 ± 0.53
EERM
52.68 ± 3.19
52.05 ± 5.94
53.83 ± 5.33
LiSA
53.01 ± 4.59
55.26 ± 1.48
54.52 ± 3.88
MARIO
50.18 ± 3.80
55.77 ± 0.74
55.72 ± 1.42
DGI
46.46 ± 3.67
48.02 ± 1.38
46.19 ± 2.76
Appendix
Table 16. Accuracy (%) on the Texas of FB-100 under Different Training Graph Combinations.