Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed to operate across these settings within a single pretrained checkpoint. Following the prior-predictive perspective, we formulate graph learning as completion of a partially observed graph. We realize this task-general view through a common interface based on random walks, allowing the same model to operate across homogeneous and multi-relational graphs with varying features, labels, and relational schemas. Wander can increase its structural context at inference time without changing its learned parameters and, under suitable assumptions, universally approximates the corresponding Bayes-optimal predictor on bounded connected graphs. Empirically, a single pretrained checkpoint achieves state-of-the-art or highly competitive results across node classification, homogeneous link prediction, and knowledge-graph link prediction. Moreover, joint pretraining across graph modalities and tasks preserves performance in specialized settings while enabling positive transfer and the composition of separately learned capabilities at inference time.
Figures & tables
Graph Foundation [-1pt]Model Families
Input Graphs
Prediction Task
node [-2pt]features
edge [-2pt]types
node [-2pt]property
link
Ultra (2024), Flock (2025)
✗
✓
✗
✓
OpenGraph , AnyGraph (2024)
✓
✗
✓
✓
TS-Net (2026), GraphPFN (2025)
✓
✗
✓
✗
UniLP (2024), TFMLinker (2026)
✗
✗
✗
✓
Wander (ours)
✓
✓
✓
✓
Table 1 : Existing GFMs generalize across different subsets of graph modalities and prediction tasks. Wander uses a single model across both node- and edge-level tasks on homogeneous and multi-relational graphs.
Figure 1 : Overview of Wander . Each layer combines (1) intra-node attention across feature, label, and structure tokens; (2) in-context attention from each node to context nodes within the same channel; and (3) a stochastic structural update that processes sampled walks and aggregates proposals through a consensus protocol. After L layers, task-specific readouts map the query node’s label tokens to class probabilities or each candidate node’s structure token to a link probability.
Inductive ( e , r )
Inductive ( e )
Transductive
Mean
Mean rank
MRR
H@10
MRR
H@10
MRR
H@10
MRR
H@10
MRR
H@10
# graphs
23
18
13
Ultra
0.345
0.513
0.431
0.566
0.312
0.458
0.366
0.518
3.250
3.417
Trix
0.368
0.540
0.455
0.592
0.339
0.500
0.390
0.548
2.657
2.750
Flock
0.369
0.554
0.456
0.604
0.340
0.509
0.391
0.560
2.343
2.222
Wander
0.378
0.558
0.460
0.609
0.353
0.516
0.399
0.565
1.750
1.611
Table 2: Knowledge graph link prediction (MRR and Hits@10).
CiteSeer
Cora
PubMed
CS
DDI
Products Home
P2P Gnutella
Email Enron
Proteins Spec1
SOC Epinions
Mean
Mean rank
CN
29.83
33.30
11.02
49.09
3.45
56.94
1.46
52.88
15.65
15.73
26.94
5.00
RA
30.26
34.99
11.12
55.63
4.85
63.21
1.29
63.07
20.67
16.37
30.15
4.10
Buddy
44.32
40.03
25.72
48.28
9.42
58.66
7.10
50.18
13.67
10.78
30.82
4.00
NBFNet
34.21
45.98
22.88
62.17
16.34
68.19
9.29
69.56
39.85
19.70
38.82
1.90
AnyGraph
32.35
44.98
15.59
47.87
7.67
55.04
5.80
40.75
15.36
13.06
27.85
4.70
Wander
65.88
56.78
31.73
64.62
12.98
71.12
9.06
69.78
31.40
20.59
43.39
1.30
Table 3 : Homogeneous link prediction (Recall@20 %).
Web- KB
Air- ports
Wiki filt.
Actor
Plane- toid
Ama- zon
Wiki CS
Cit. Full
Roman Emp.
Amz. Rat.
Coau- thor
HM Cat.
Pokec 100k
City Netw.
Mean
Mean rank
# graphs
3
3
2
1
3
2
1
2
1
1
2
1
1
3
mean # nodes
206
573
1.6k
7.6k
8.6k
10.7k
11.7k
18.8k
22.7k
24.5k
26.4k
46.6k
100k
180k
GCN
51.3
46.9
37.4
33.1
70.8
87.9
79.4
66.2
79.6
51.2
92.0
66.4
78.4
38.2
60.64
3.12
TabPFNv3
73.3
32.0
44.0
39.0
67.2
83.1
73.8
59.2
65.7
49.1
90.8
42.1
43.2
47.0
58.66
3.69
GraphAny
63.0
40.2
29.2
29.1
74.7
86.9
75.4
64.2
64.2
42.5
91.7
37.1
44.0
18.9
54.85
3.96
\textscNodePFN∗
74.5
59.1
44.3
32.7
76.2
85.7
76.0
—
48.8
44.2
91.7
—
78.1
44.6
—
—
Table 4 : Node classification (test accuracy %, averaged over all available splits).
KG LP (MRR)
Hom. LP (R@20)
NC (Acc.)
Wander
0.399
43.39
68.20
Wander one task
0.398
40.37
67.69
Table 5 : Joint vs. single-task training.
edge types
fea- tures
Concept- Net100k
WN-v1- WN-v4
WN- 18RR
✗
✗
0.135
0.162
0.116
✗
✓
0.202
0.273
0.163
✓
✗
0.258
0.626
0.550
✓
✓
0.334
0.637
0.564
Table 6 : KG entity prediction (MRR) with vs. without node features and edge types.
Figure 2 : Predicted true-class probability on single grid taken from Grids for increasing walk lengths and context-aware walk filtering.
Grids
Real-world
GraphAny
20.0
34.0
NodePFN
26.6
41.2
GraphPFN
23.3
60.1
Wander
96.6
76.5
Wander (filt.)
100.0
78.3
Table 7 : Accuracy on Grids and four real-world graphs without node features.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : Structural coverage of the synthetic graph priors compared with the real-world evaluation graphs.
Figure 4 : Distributions of graph statistics for 1,000 graphs sampled from each synthetic prior, alongside the real-world evaluation graphs.
Phase 1
Phase 2
Phase 3
Phase 4
Epoch
1–100
101–200
201–300
301–350
Architecture
Hidden dimension d
128 (all phases)
Layers L
6 (all phases)
Heads (intra-node)
4 (all phases)
Heads (ICL)
4 (all phases)
Appendix
Table 8 : Pretraining hyperparameters across the four pretraining stages of Wander .
Dataset
Training Graph
Validation Graph
Test Graph
Entities
Rels
Triples
Entities
Rels
Triples
Valid
Entities
Rels
Triples
Test
FB-25
5190
163
91571
4097
216
17147
5716
4097
216
17147
5716
FB-50
5190
153
85375
4445
205
11636
3879
4445
205
11636
3879
FB-75
4659
134
62809
2792
186
9316
3106
2792
186
9316
3106
FB-100
4659
134
62809
2624
77
6987
2329
2624
77
6987
2329
WK-25
12659
47
41873
3228
74
3391
1130
3228
74
3391
1131
Appendix
Table 9 : Dataset statistics for inductive- e,r link prediction datasets. Triples are the number of edges given at training, validation, or test, respectively, whereas Valid and Test denote triples to be predicted in the validation and test graphs.
Dataset
Rels
Training Graph
Validation Graph
Test Graph
Entities
Triples
Entities
Triples
Valid
Entities
Triples
Test
FB-v1
180
1594
4245
1594
4245
489
1093
1993
411
FB-v2
200
2608
9739
2608
9739
1166
1660
4145
947
FB-v3
215
3668
17986
3668
17986
2194
2501
7406
1731
FB-v4
219
4707
27203
4707
27203
3352
3051
11714
2840
WN-v1
9
2746
5410
2746
5410
630
922
1618
373
Appendix
Table 10 : Dataset statistics for inductive- e link prediction datasets. Triples are the number of edges given at training, validation, or test, respectively, whereas Valid and Test denote triples to be predicted in the validation and test graphs.
Dataset
Entities
Rels
Train
Valid
Test
Entity Task
CoDEx Small
2034
42
32888
1827
1828
h/t
CoDEx Large
77951
69
551193
30622
30622
h/t
NELL995
74536
200
149678
543
2818
h/t
YAGO3-10
123182
37
1079040
5000
5000
h/t
WDsinger
10282
135
16142
2163
2203
h/t
NELL23k
22925
200
25445
4961
4952
h/t
Appendix
Table 11 : Dataset statistics for transductive link prediction datasets. Entity task denotes the entity-prediction task: h/t is predicting both heads and tails, and t is predicting only tails.
Dataset
Ultra
Trix
Flock
Wander
MRR
H@10
MRR
H@10
MRR
H@10
MRR
H@10
Inductive (entity, relation)
FB-25
0.388
0.640
0.393
0.650
0.404
0.664
0.404
0.663
FB-50
0.338
0.543
0.334
0.547
0.352
0.566
0.352
0.568
FB-75
0.403
0.604
0.401
0.611
0.418
0.622
0.408
0.617
FB-100
0.449
0.642
0.436
0.635
0.452
0.663
0.458
0.663
WK-25
0.316
0.532
0.305
0.496
0.280
0.491
0.288
0.477
Appendix
Table 12 : Zero-shot knowledge graph link prediction results. Best in each row is bold , second-best is underlined .
Dataset(s)
Base walks Kbase
Inference samples
Inductive (e,r)
FB-25, FB-50
16
16
FB-75, FB-100
8
16
WK-25, WK-75
4
16
WK-50, WK-100
16
16
NL-0, NL-25, NL-75
2
16
Appendix
Table 13 : Random-walk configuration used for knowledge-graph link prediction. We follow the entity-prediction evaluation configuration of Flock ( Kim et al., 2026 ) . Each Wander structural update samples 3Kbase walks. The walk length is S=128 for all datasets.
NBFNet
Buddy
Dataset
Depth
lr
#Neg.
Hidden
lr
#Neg.
CiteSeer
3
0.005
1
256
0.005
1
Cora
5
0.001
32
128
0.005
32
PubMed
5
0.005
1
256
0.001
256
CS
3
0.005
32
256
0.001
256
DDI
5
0.005
256
256
0.001
256
Appendix
Table 14 : Val-selected hyperparameters for the NBFNet and Buddy link-prediction baselines.
Dataset
#Nodes
#Edges
Avg. degree
Clustering Coeff.
Diameter
#Features
#Test sources (original)
#Test sources (filtered)
CiteSeer
3,327
4,552
2.7
0.141
28
3,703
423
175
Cora
2,708
5,278
3.9
0.241
18
1,433
399
192
PubMed
19,717
44,324
4.5
0.060
16
500
2,041
895
CS
18,333
81,894
8.9
0.343
24
6,805
1,973
1,377
DDI
4,267
1,201,400
563.1
0.576
5
–
1,587
1,587
P2P Gnutella06
8,717
31,525
7.2
0.007
9
–
2,107
2,107
Appendix
Table 15 : Statistics of the homogeneous link-prediction datasets. Test sources are unique heads of test edges. Original is the AnyGraph split. Remaining drops a test edge when either direction already appears in the training graph, which is stricter than AnyGraph’s directed protocol; a source is kept only if at least one test edge survives.
Celegans
USAir
NS
PB
Cora
CS
Facebook
R@20
min
R@20
min
R@20
min
R@20
min
R@20
min
R@20
min
R@20
min
UniLP
70.9
57
85.2
45
94.2
112
83.4
450
75.4
269
93.1
187
95.5
438
Wander (w/o feat.)
77.2
3
90.6
3
89.4
6
80.7
13
76.4
15
91.3
45
92.2
52
Wander (w/ feat.)
–
–
–
–
–
–
–
–
90.6
115
97.0
142
95.8
159
Appendix
Table 16 : Sampled Recall@20 and test-only wall-clock for UniLP vs. Wander on two random 70/10/20 splits. Per source: rank that source’s eval neighbors plus 100 sampled non-edges.
Group
Dataset
# nodes
# features
# classes
% train
avg. deg.
unbiased homophily
# official splits
WebKB
Cornell
183
1.7k
5
47.5%
3.0
-0.47
10
Texas
183
1.7k
5
47.5%
3.0
-0.81
10
Wisconsin
251
1.7k
5
47.8%
3.6
-0.29
10
Airports
Air Brazil
131
131
4
61.1%
15.3
0.00
–
Air Europe
399
399
4
20.1%
30.0
-0.12
–
Air USA
1.2k
1.2k
4
6.7%
22.9
0.43
–
Appendix
Table 17 : Statistics of considered node classification datasets. “# official splits” counts train/val/test masks shipped with the dataset. “–” denotes datasets without an official split; we use a custom protocol there (5 seeds, 20 labeled nodes per class).
Table 19 : Multiclass node classification results (test accuracy %, mean ± std over all available splits). NodePFN is shown but excluded from mean and mean rank (missing Full Cora and HM Categories). TabPFN v3 uses node features only (ignores the graph).
Grids
PubMed
Full DBLP
Co. CS
Co. Physics
GraphAny
20.0
18.0
44.7 ± 0.0
22.6 ± 0.0
50.5 ± 0.0
NodePFN
26.6 ± 0.0
39.5 ± 0.0
45.9 ± 1.0
28.3 ± 0.5
51.2 ± 0.3
GraphPFN
23.3 ± 0.7
49.9 ± 2.6
61.8 ± 1.6
59.2 ± 3.5
69.4 ± 1.2
Wander
96.6 ± 0.5
73.7 ± 0.7
65.9 ± 3.9
84.1 ± 0.6
82.2 ± 5.0
Wander (filtered)
100.0 ± 0.0
72.8 ± 0.5
70.0 ± 4.6
84.6 ± 0.6
85.8 ± 2.7
Appendix
Table 20 : Test accuracy on the synthetic Grids dataset and selected real-world graphs without node features.
Cora (LP)
Full DBLP (NC)
Wander
NBFNet
Wander
GCN
Grid search
–
13 min 36 s
–
1 min 28 s
Single batch (1 inference sample)
1.70 s
1.1 ms
1.13 s
2.0 ms
Single batch (16 inference sample)
27.2 s
–
18.1 s
–
Full eval (1 inference sample)
37 s
13 min 36 s
1 min 18 s
1 min 28 s
Full eval (16 inference samples)
10 min 58 s
–
21 min 16 s
–
Appendix
Table 21 : Wall-clock runtimes on a single NVIDIA RTX PRO 6000 GPU. Times are reported for one seed and, for Full DBLP, one data split.
Figure 5 : Compute–performance trade-offs obtained by reducing the number of walks, walk length, or inference samples relative to the default configuration. Results are averaged across the datasets of each benchmark.
Cycles
Majority
64.0
GCN
63.3
GraphPFN
67.4
Wander
97.5
Appendix
Table 22 : Accuracy on Cycles (finetuned).
Dataset
Ultra (ft.)
Trix (ft.)
Flock (ft.)
Wander
Wander (ft.)
MRR
H@10
MRR
H@10
MRR
H@10
MRR
H@10
MRR
H@10
Inductive (entity, relation)
FB-25
0.383
0.635
0.393
0.650
0.405
0.666
0.404
0.663
0.406
0.665
FB-50
0.334
0.538
0.334
0.547
0.357
0.570
0.352
0.568
0.345
0.562
FB-75
0.400
0.598
0.401
0.611
0.425
0.630
0.408
0.617
0.424
0.631
FB-100
0.444
0.643
0.436
0.633
0.460
0.668
0.458
0.663
0.458
0.666
WK-25
0.321
0.535
0.300
0.493
0.298
0.506
0.288
0.477
0.293
0.493
Appendix
Table 23 : Finetuned knowledge graph link prediction results. Best in each row is bold , second-best is underlined .
CiteSeer
Cora
PubMed
CS
DDI
Products Home
P2P Gnutella
Email Enron
Proteins Spec1
SOC Epinions
Mean
Mean rank
Wander
65.88
56.78
31.73
64.62
12.98
71.12
9.06
69.78
31.40
20.59
43.39
1.90
Wander (ft.)
69.20
60.90
39.70
70.00
11.90
72.60
9.30
71.90
36.60
21.30
46.34
1.10
Appendix
Table 24 : Finetuned homogeneous link prediction (Recall@20 %).
Group
Dataset
Zero-shot
Finetuned
WebKB
Cornell
74.77 ± 1.27
72.07 ± 1.27
Texas
75.68 ± 2.21
80.18 ± 1.27
Wisconsin
76.47 ± 0.00
71.24 ± 0.92
Airports
Air Brazil
78.21 ± 1.81
80.77 ± 0.00
Air Europe
57.29 ± 2.06
57.29 ± 0.29
Air USA
61.20 ± 0.08
61.62 ± 0.67
Appendix
Table 25 : Node classification finetuning versus zero-shot Wander on split 0 only (test accuracy %, mean ± std over 3 seeds).
Figure 6 : Effect of replacing Wander ’s synthetic attributed-graph prior with the NodePFN or GraphPFN prior. Points show results on individual real-world datasets; vertical differences are measured relative to the model trained with the Wander prior, in percentage points.
We propose Mochi, a Graph Foundation Model that addresses task unification and training efficiency by adopting a meta-learning based training framework. Prior models pre-train with reconstruction-based objectives such as link prediction, and assume that the resulting representations can be aligned with downstream tasks through a separate unification step such as class prototypes. We demonstrate through synthetic and real-world experiments that this procedure, while simple and intuitive, has limitations that directly affect downstream task performance. To address these limitations, Mochi pre-trains on few-shot episodes that mirror the downstream evaluation protocol, aligning the training objective with inference rather than relying on a post-hoc unification step. We show that Mochi, along with its more powerful variant Mochi++, achieves competitive or superior performance compared to existing Graph Foundation Models across 25 real-world graph datasets spanning node classification, link prediction, and graph classification, while requiring 8∼27 times less training time than the strongest baseline.
João Mattos, Arlei Silva
Computer Science Department · Ken Kennedy Institute
Graph foundation models face several fundamental challenges including transferability across diverse domains and data scarcity, which calls into question the very feasibility of creating such models. However, despite similar challenges, the tabular domain has recently witnessed the emergence of the first successful foundation models such as TabPFN. These models are based on the prior-data fitted networks (PFN) framework, in which models are pretrained on carefully designed synthetic datasets to make predictions in an in-context learning setting. Recently, G2T-FM, a framework that converts graph node-level tasks into tabular tasks, has made the first step towards adopting PFNs for graphs, yet it is limited to hand-crafted features and was never pretrained on graph data. In this work, we make the next step by proposing GraphPFN, a PFN-based model designed and pretrained specifically for graph node-level tasks. Following the PFN framework, we first design a prior distribution of synthetic attributed graphs by using a novel combination of multi-level stochastic block models and a preferential attachment process for structure generation and graph-aware structured causal models for attribute generation. Then, we augment the tabular foundation model LimiX with attention-based graph neighborhood aggregation layers and train it on millions of synthetic graphs sampled from our prior. On diverse real-world graph datasets with node-level tasks, GraphPFN achieves state-of-the-art results in both in-context learning and finetuning regimes, outperforming G2T-FM, prior GFMs, and task-specific GNNs trained from scratch. More broadly, GraphPFN shows the potential of PFN-based models for building graph foundation models.
Unlike vision and language domains, graph learning lacks a shared input space, as input features differ across graph datasets not only in semantics, but also in value ranges and dimensionality. This misalignment prevents graph models from generalizing across datasets, limiting their use as foundation models. In this work, we propose ALL-IN, a simple and theoretically grounded method that enables transferability across datasets with different input features. Our approach projects node features into a shared random space and constructs representations via covariance-based statistics, thus eliminating dependence on the original feature space. We show that the computed node-covariance operators and the resulting node representations are invariant in distribution to permutations of the input features. We further demonstrate that the expected operator exhibits invariance to general orthogonal transformations of the input features. Empirically, ALL-IN achieves strong performance across diverse node- and graph-level tasks on unseen datasets with new input features, without requiring architecture changes or retraining. These results point to a promising direction for input-agnostic, transferable graph models.
Moshe Eliasof, Krishna Sri Ipsit Mantri, Beatrice Bevilacqua +2
University of Cambridge · Ben-Gurion University of the Negev · Purdue University