GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher-student alignment objective, we propose Graph Geometry-aware MLP (G^2MLP), a training-time distillation framework guided by Ollivier-Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G^2MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.
Figures & tables
Figure 1: Two spectral failure modes of GNN-to-MLP distillation. (a) Singular-value spectra: GLNN exhibits spectral underfit on sparse graphs and spectral overfit on dense graphs. (b) Rank-direction sketch. (c) Per-dataset rank gap.
Figure 2: Overview of G 2 MLP. A frozen GNN teacher and an MLP student produce embeddings and logits; precomputed Ollivier–Ricci curvature yields per-node weights αv and βv that allocate prediction-level and representation-level alignment during training.
Teacher
Graph-free students
Dataset
SAGE
MLP
GLNN
KRD
FF-G2M
G 2 MLP
Citeseer
70.49 ± 1.53
58.50 ± 1.86
71.22 ± 1.50
72.84 ± 1.70
72.85 ± 1.59
74.51 ± 2.35
Pubmed
75.56 ± 2.06
68.39 ± 3.09
75.59 ± 2.46
77.01 ± 3.11
76.56 ± 3.41
78.17 ± 2.75
Cora
80.64 ± 1.57
59.18 ± 1.60
80.26 ± 1.66
82.27 ± 1.31
82.38 ± 1.41
82.54 ± 1.87
A-computer
82.82 ± 1.37
67.62 ± 2.21
82.71 ± 1.18
82.87 ± 0.87
83.67 ± 1.04
83.87 ± 1.09
A-photo
90.85 ± 0.87
77.29 ± 1.79
91.95 ± 1.04
91.95 ± 1.44
92.51 ± 0.56
92.85 ± 0.78
Table 1: Node classification accuracy (%) under the transductive setting. Best student is in bold .
Teacher
Graph-free students
Dataset
Eval
SAGE
MLP
GLNN
KRD
FF-G2M
G 2 MLP
Citeseer
prod
68.06
58.49
69.08
71.38
71.89
71.67
ind
69.14 ± 2.99
59.31 ± 4.56
68.48 ± 2.38
69.78 ± 3.04
69.75 ± 3.16
71.85 ± 2.27
tran
67.79 ± 2.80
58.29 ± 1.94
69.23 ± 2.39
71.77 ± 2.81
72.12 ± 2.69
71.62 ± 1.31
Pubmed
prod
74.77
68.39
74.67
76.00
73.98
77.45
ind
75.07 ± 2.89
68.28 ± 3.25
74.52 ± 2.95
75.17 ± 3.11
73.49 ± 7.91
77.30 ± 1.48
Table 2: Node classification accuracy (%) under the production setting.
GT
GraphGPS
NAGphormer
Dataset
T
S
S†
T
S
S†
T
S
S†
Citeseer
71.16 ± 2.31
71.85 ± 2.45
72.35 ± 1.85
68.07 ± 2.62
69.38 ± 2.30
69.68 ± 2.05
70.09 ± 1.81
71.24 ± 1.64
71.82 ± 1.29
Pubmed
74.09 ± 2.48
74.61 ± 2.48
74.76 ± 2.83
71.77 ± 2.81
71.68 ± 2.61
72.03 ± 2.58
76.19 ± 2.86
76.85 ± 3.34
77.06 ± 3.49
Cora
78.53 ± 2.08
78.98 ± 1.90
79.21 ± 2.09
76.21 ± 1.91
76.06 ± 1.99
76.77 ± 1.97
80.88 ± 1.77
81.00 ± 1.45
81.48 ± 1.53
A-computer
83.89 ± 1.59
84.53 ± 1.42
84.72 ± 1.32
78.30 ± 1.70
78.99 ± 1.73
80.12 ± 1.55
82.78 ± 1.92
83.02 ± 1.87
83.31 ± 1.77
A-photo
90.94 ± 0.87
91.90 ± 0.56
92.06 ± 0.82
90.41 ± 1.10
91.03 ± 0.95
91.32 ± 0.93
91.45 ± 0.82
92.07 ± 0.74
92.49 ± 0.65
Table 3: Transformer-teacher to MLP distillation. T : teacher; S : GLNN student; S† : G 2 MLP student. Bold : student value ≥ teacher. Underline : S†>S .
Cora
Citeseer
Pubmed
Method
AUC
AP
AUC
AP
AUC
AP
SAGE (teacher, ref)
81.29 ± 1.99
80.54 ± 2.69
81.21 ± 3.15
81.80 ± 3.87
88.36 ± 1.65
86.83 ± 1.54
MLP
77.92 ± 1.78
76.23 ± 2.01
76.47 ± 2.08
76.13 ± 2.74
90.29 ± 0.38
89.34 ± 0.44
GLNN
78.61 ± 1.68
77.05 ± 1.76
83.32 ± 2.34
83.71 ± 2.53
90.03 ± 0.38
89.12 ± 0.32
KRD
79.20 ± 1.65
77.89 ± 2.31
83.31 ± 2.23
83.81 ± 2.42
90.32 ± 0.42
89.38 ± 0.51
G 2 MLP (Ours)
83.90 ± 1.10
83.36 ± 0.99
83.62 ± 2.33
84.32 ± 2.43
90.36 ± 1.29
89.96 ± 1.07
Table 4: Link prediction (multi-task regime). SAGE row: teacher reference. Bold : best graph-free student per column. Underline : exceeds the teacher.
Figure 3: Inference time on Arxiv ( ∼ 170K nodes).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
#Nodes
#Edges
#Feat.
#Cls.
dˉ
h
Cora
2,708
5,429
1,433
7
3.9
0.81
Citeseer
3,327
4,732
3,703
6
2.7
0.74
Pubmed
19,717
44,338
500
3
4.5
0.80
A-computer
13,752
245,861
767
10
35.8
0.78
A-photo
7,650
119,081
745
8
31.1
0.83
ogbn-arxiv
169,343
1,166,243
128
40
13.7
0.65
Appendix
Table 5: Benchmark statistics of the full graphs as published. dˉ is the mean degree; h is edge homophily. Following CPF, node-classification experiments on the five CPF datasets use the largest connected component (Cora 2,485, Citeseer 2,110, Pubmed 19,717, A-computer 13,381, A-photo 7,487 nodes).
Dataset
η
weight-decay
dropout
Cora
10−2
5×10−3
0.6
Citeseer
10−2
10−3
0.1
Pubmed
5×10−3
0
0.4
A-computer
10−3
2×10−3
0.3
A-photo
5×10−3
2×10−3
0.3
ogbn-arxiv
10−2
0
0.2
Appendix
Table 6: Student MLP optimizer hyperparameters.
Variant
Cora
Citeseer
Pubmed
GLNN baseline
80.26 ± 1.66
71.22 ± 1.50
75.59 ± 2.46
+W2 uniform (no curvature)
80.86 ± 2.30
71.07 ± 2.19
75.22 ± 2.15
+ curvature, logit-side only
81.78 ± 1.55
73.87 ± 2.09
77.27 ± 2.92
+ curvature, feature-side only
81.34 ± 1.88
73.99 ± 2.02
77.14 ± 2.93
Full G 2 MLP (both sides)
82.54 ± 1.87
74.51 ± 2.35
78.17 ± 2.75
Appendix
Table 7: Component ablation on transductive node classification. Baseline reproduced from Table 1 .
GraphGPS (dense, canonical)
GraphGPS (sparse, edge-softmax)
Dataset
T
S
S†
T
S
S†
Cora
76.21 ± 1.91
76.06 ± 1.99
76.77 ± 1.97
78.12 ± 2.12
77.86 ± 2.09
78.74 ± 1.87
Citeseer
68.07 ± 2.62
69.38 ± 2.30
69.68 ± 2.05
68.88 ± 2.25
70.29 ± 2.13
70.44 ± 2.13
Pubmed
71.77 ± 2.81
71.68 ± 2.61
72.03 ± 2.58
75.61 ± 2.13
76.08 ± 2.34
75.97 ± 2.21
A-computer
78.30 ± 1.70
78.99 ± 1.73
80.12 ± 1.55
81.84 ± 0.88
82.05 ± 1.18
82.55 ± 0.86
A-photo
90.41 ± 1.10
91.03 ± 0.95
91.32 ± 0.93
90.34 ± 1.21
91.12 ± 1.04
91.26 ± 1.02
Appendix
Table 8: G 2 MLP vs. GLNN with the GraphGPS teacher under both canonical dense ( N×N ) and sparse (edge-softmax) attention. Conventions as in Table 3 : T teacher, S GLNN student, S† G 2 MLP student. Bold : student value ≥ teacher. Underline : S†>S .
Cora
Citeseer
Pubmed
Method
AUC
AP
AUC
AP
AUC
AP
SAGE (teacher, ref)
81.29 ± 1.99
80.54 ± 2.69
81.21 ± 3.15
81.80 ± 3.87
88.36 ± 1.65
86.83 ± 1.54
GLNN
56.18 ± 5.79
56.81 ± 4.55
83.99 ± 2.03
84.23 ± 2.16
88.62 ± 1.57
87.17 ± 1.58
KRD
57.23 ± 5.53
57.41 ± 5.02
84.26 ± 2.56
84.46 ± 2.59
88.63 ± 1.56
87.18 ± 1.57
G 2 MLP (Ours)
81.77 ± 2.25
80.28 ± 2.84
82.89 ± 3.61
83.15 ± 4.39
88.82 ± 1.46
87.43 ± 1.35
Appendix
Table 9: Link prediction (pure-distillation regime, no edge-BCE). Conventions as in Table 4 .
Learning local geometry enables graph neural networks (GNNs) to adapt how they compare and integrate neighborhood information. However, estimating geometry from aggregated representations can overlook variation among individual messages and dependencies across feature dimensions. We propose GeoF, a recurrent framework that jointly evolves node features and propagation geometry through message-passing feedback. Each node maintains a local symmetric positive-definite geometry, initialized from a structure-aware prototype atlas and parameterized in block log-triangular coordinates. At each step, the geometry determines neighborhood weights, while triangular frame transport maps transformed source messages into the target node's local coordinates before aggregation. Weighted second-order statistics of residuals between aligned messages and the transformed target state capture directional variation and within-block dependencies, yielding a geometric update target. A shared controller learns complementary corrections through task supervision. A bounded log-triangular update combines these corrections, the target, and the previous geometric state while preserving positive definiteness. The geometry governs subsequent propagation, closing the feedback loop. With parameters shared across recurrent steps, task-specific readouts support node classification, link prediction, and graph classification. Experiments on benchmark datasets show that GeoF consistently outperforms state-of-the-art GNN baselines.
Yingxu Wang, Kunyu Zhang, Xinwang Liu +4
The Chinese University of Hong Kong · The Education University of Hong Kong · National University of Defense Technology +3
Spectral graph sparsification is a classical tool for reducing graph complexity while preserving Laplacian quadratic forms. In graph neural networks (GNNs), sparsification is often used to accelerate computation while maintaining predictive performance. In this work, we study a complementary representation-level question: does sparsification preserve the geometry of learned embeddings? For polynomial-filter GNNs, we prove that any ε-spectral sparsifier induces O(ε) perturbations in polynomial graph filters, multilayer hidden representations, and their Gram matrices. These guarantees imply stability of squared pairwise distances, class means, and covariance structure in embedding space. We further establish finite-time training stability: under smoothness and boundedness assumptions, gradient descent on dense and sparsified graphs produces weight trajectories whose separation grows at most proportionally to the sparsification distortion. Empirically, effective-resistance sparsification validates the predicted perturbation chain on synthetic graphs and preserves hidden representation geometry on real datasets. In our experiments, the gram matrix and training dynamics show low divergence even under substantial sparsification, consistent with the predicted stability under spectral sparsification. Hidden Gram preservation strongly predicts neighborhood preservation and class-centroid stability across FashionMNIST, Cora, and Paul15. Together, these results show that spectral sparsification preserves not only graph operators, but also the representation geometry that supports downstream use of GNN embeddings for interpretability.
Sanjukta Krishnagopal
Department of Computer Science, University of California Santa Barbara
Graph neural networks (GNNs) can operate on large graphs but become infrastructure-sensitive at the scale of millions of nodes and typically require scalable training techniques for even larger graphs. This raises a central question: when can a model trained on a smaller, scaled-down replica of a graph be deployed on the full-resolution graph without retraining? We introduce a zero-shot transfer protocol in which a GNN is trained on a graph coarse-grained by geometric renormalization (GR), and the resulting weights are transferred directly to the original network. Across synthetic and real-world networks, training on GR scaled-down replicas preserves much of the original-scale predictive performance while significantly reducing training cost. We further find that learned representations and predictive trajectories remain aligned across scales. These findings suggest that structural similarity may be more important than network size in determining GNN transferability, opening a path toward scale-equivariant graph architectures.
Robert Jankowski, Pedro Almagro-Blanco, Marián Boguñá +2
TU Delft · University of Seville · University of Barcelona +2