GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher-student alignment objective, we propose Graph Geometry-aware MLP (G^2MLP), a training-time distillation framework guided by Ollivier-Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G^2MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.
Figures & tables
Figure 1: Two spectral failure modes of GNN-to-MLP distillation. (a) Singular-value spectra: GLNN exhibits spectral underfit on sparse graphs and spectral overfit on dense graphs. (b) Rank-direction sketch. (c) Per-dataset rank gap.
Figure 2: Overview of G 2 MLP. A frozen GNN teacher and an MLP student produce embeddings and logits; precomputed Ollivier–Ricci curvature yields per-node weights αv and βv that allocate prediction-level and representation-level alignment during training.
Teacher
Graph-free students
Dataset
SAGE
MLP
GLNN
KRD
FF-G2M
G 2 MLP
Citeseer
70.49 ± 1.53
58.50 ± 1.86
71.22 ± 1.50
72.84 ± 1.70
72.85 ± 1.59
74.51 ± 2.35
Pubmed
75.56 ± 2.06
68.39 ± 3.09
75.59 ± 2.46
77.01 ± 3.11
76.56 ± 3.41
78.17 ± 2.75
Cora
80.64 ± 1.57
59.18 ± 1.60
80.26 ± 1.66
82.27 ± 1.31
82.38 ± 1.41
82.54 ± 1.87
A-computer
82.82 ± 1.37
67.62 ± 2.21
82.71 ± 1.18
82.87 ± 0.87
83.67 ± 1.04
83.87 ± 1.09
A-photo
90.85 ± 0.87
77.29 ± 1.79
91.95 ± 1.04
91.95 ± 1.44
92.51 ± 0.56
92.85 ± 0.78
Table 1: Node classification accuracy (%) under the transductive setting. Best student is in bold .
Teacher
Graph-free students
Dataset
Eval
SAGE
MLP
GLNN
KRD
FF-G2M
G 2 MLP
Citeseer
prod
68.06
58.49
69.08
71.38
71.89
71.67
ind
69.14 ± 2.99
59.31 ± 4.56
68.48 ± 2.38
69.78 ± 3.04
69.75 ± 3.16
71.85 ± 2.27
tran
67.79 ± 2.80
58.29 ± 1.94
69.23 ± 2.39
71.77 ± 2.81
72.12 ± 2.69
71.62 ± 1.31
Pubmed
prod
74.77
68.39
74.67
76.00
73.98
77.45
ind
75.07 ± 2.89
68.28 ± 3.25
74.52 ± 2.95
75.17 ± 3.11
73.49 ± 7.91
77.30 ± 1.48
Table 2: Node classification accuracy (%) under the production setting.
GT
GraphGPS
NAGphormer
Dataset
T
S
S†
T
S
S†
T
S
S†
Citeseer
71.16 ± 2.31
71.85 ± 2.45
72.35 ± 1.85
68.07 ± 2.62
69.38 ± 2.30
69.68 ± 2.05
70.09 ± 1.81
71.24 ± 1.64
71.82 ± 1.29
Pubmed
74.09 ± 2.48
74.61 ± 2.48
74.76 ± 2.83
71.77 ± 2.81
71.68 ± 2.61
72.03 ± 2.58
76.19 ± 2.86
76.85 ± 3.34
77.06 ± 3.49
Cora
78.53 ± 2.08
78.98 ± 1.90
79.21 ± 2.09
76.21 ± 1.91
76.06 ± 1.99
76.77 ± 1.97
80.88 ± 1.77
81.00 ± 1.45
81.48 ± 1.53
A-computer
83.89 ± 1.59
84.53 ± 1.42
84.72 ± 1.32
78.30 ± 1.70
78.99 ± 1.73
80.12 ± 1.55
82.78 ± 1.92
83.02 ± 1.87
83.31 ± 1.77
A-photo
90.94 ± 0.87
91.90 ± 0.56
92.06 ± 0.82
90.41 ± 1.10
91.03 ± 0.95
91.32 ± 0.93
91.45 ± 0.82
92.07 ± 0.74
92.49 ± 0.65
Table 3: Transformer-teacher to MLP distillation. T : teacher; S : GLNN student; S† : G 2 MLP student. Bold : student value ≥ teacher. Underline : S†>S .
Cora
Citeseer
Pubmed
Method
AUC
AP
AUC
AP
AUC
AP
SAGE (teacher, ref)
81.29 ± 1.99
80.54 ± 2.69
81.21 ± 3.15
81.80 ± 3.87
88.36 ± 1.65
86.83 ± 1.54
MLP
77.92 ± 1.78
76.23 ± 2.01
76.47 ± 2.08
76.13 ± 2.74
90.29 ± 0.38
89.34 ± 0.44
GLNN
78.61 ± 1.68
77.05 ± 1.76
83.32 ± 2.34
83.71 ± 2.53
90.03 ± 0.38
89.12 ± 0.32
KRD
79.20 ± 1.65
77.89 ± 2.31
83.31 ± 2.23
83.81 ± 2.42
90.32 ± 0.42
89.38 ± 0.51
G 2 MLP (Ours)
83.90 ± 1.10
83.36 ± 0.99
83.62 ± 2.33
84.32 ± 2.43
90.36 ± 1.29
89.96 ± 1.07
Table 4: Link prediction (multi-task regime). SAGE row: teacher reference. Bold : best graph-free student per column. Underline : exceeds the teacher.
Figure 3: Inference time on Arxiv ( ∼ 170K nodes).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
#Nodes
#Edges
#Feat.
#Cls.
dˉ
h
Cora
2,708
5,429
1,433
7
3.9
0.81
Citeseer
3,327
4,732
3,703
6
2.7
0.74
Pubmed
19,717
44,338
500
3
4.5
0.80
A-computer
13,752
245,861
767
10
35.8
0.78
A-photo
7,650
119,081
745
8
31.1
0.83
ogbn-arxiv
169,343
1,166,243
128
40
13.7
0.65
Appendix
Table 5: Benchmark statistics of the full graphs as published. dˉ is the mean degree; h is edge homophily. Following CPF, node-classification experiments on the five CPF datasets use the largest connected component (Cora 2,485, Citeseer 2,110, Pubmed 19,717, A-computer 13,381, A-photo 7,487 nodes).
Dataset
η
weight-decay
dropout
Cora
10−2
5×10−3
0.6
Citeseer
10−2
10−3
0.1
Pubmed
5×10−3
0
0.4
A-computer
10−3
2×10−3
0.3
A-photo
5×10−3
2×10−3
0.3
ogbn-arxiv
10−2
0
0.2
Appendix
Table 6: Student MLP optimizer hyperparameters.
Variant
Cora
Citeseer
Pubmed
GLNN baseline
80.26 ± 1.66
71.22 ± 1.50
75.59 ± 2.46
+W2 uniform (no curvature)
80.86 ± 2.30
71.07 ± 2.19
75.22 ± 2.15
+ curvature, logit-side only
81.78 ± 1.55
73.87 ± 2.09
77.27 ± 2.92
+ curvature, feature-side only
81.34 ± 1.88
73.99 ± 2.02
77.14 ± 2.93
Full G 2 MLP (both sides)
82.54 ± 1.87
74.51 ± 2.35
78.17 ± 2.75
Appendix
Table 7: Component ablation on transductive node classification. Baseline reproduced from Table 1 .
GraphGPS (dense, canonical)
GraphGPS (sparse, edge-softmax)
Dataset
T
S
S†
T
S
S†
Cora
76.21 ± 1.91
76.06 ± 1.99
76.77 ± 1.97
78.12 ± 2.12
77.86 ± 2.09
78.74 ± 1.87
Citeseer
68.07 ± 2.62
69.38 ± 2.30
69.68 ± 2.05
68.88 ± 2.25
70.29 ± 2.13
70.44 ± 2.13
Pubmed
71.77 ± 2.81
71.68 ± 2.61
72.03 ± 2.58
75.61 ± 2.13
76.08 ± 2.34
75.97 ± 2.21
A-computer
78.30 ± 1.70
78.99 ± 1.73
80.12 ± 1.55
81.84 ± 0.88
82.05 ± 1.18
82.55 ± 0.86
A-photo
90.41 ± 1.10
91.03 ± 0.95
91.32 ± 0.93
90.34 ± 1.21
91.12 ± 1.04
91.26 ± 1.02
Appendix
Table 8: G 2 MLP vs. GLNN with the GraphGPS teacher under both canonical dense ( N×N ) and sparse (edge-softmax) attention. Conventions as in Table 3 : T teacher, S GLNN student, S† G 2 MLP student. Bold : student value ≥ teacher. Underline : S†>S .
Cora
Citeseer
Pubmed
Method
AUC
AP
AUC
AP
AUC
AP
SAGE (teacher, ref)
81.29 ± 1.99
80.54 ± 2.69
81.21 ± 3.15
81.80 ± 3.87
88.36 ± 1.65
86.83 ± 1.54
GLNN
56.18 ± 5.79
56.81 ± 4.55
83.99 ± 2.03
84.23 ± 2.16
88.62 ± 1.57
87.17 ± 1.58
KRD
57.23 ± 5.53
57.41 ± 5.02
84.26 ± 2.56
84.46 ± 2.59
88.63 ± 1.56
87.18 ± 1.57
G 2 MLP (Ours)
81.77 ± 2.25
80.28 ± 2.84
82.89 ± 3.61
83.15 ± 4.39
88.82 ± 1.46
87.43 ± 1.35
Appendix
Table 9: Link prediction (pure-distillation regime, no edge-BCE). Conventions as in Table 4 .