NodeGround: A Node Classification Benchmark in the Graph Foundation Model Era
Organizations: Nums AI
Abstract
Can a pretrained graph model replace training and tuning a separate predictor for each dataset? Answering this requires evaluating prediction quality alongside computational cost. We present NodeGround, a node classification benchmark that puts graph foundation models (GFMs) and dataset-specific supervised learning under a common evaluation framework. The benchmark spans 51 datasets and evaluates six GFMs alongside 15 supervised methods under two label-availability regimes. Shared data partitions, validation-only model selection, controlled hyperparameter searches, and multiple predictive metrics make comparisons systematic, while workflow measurements account for adaptation, training, tuning, and inference. The results favor carefully tuned graph neural networks overall. GraphPFN reaches third place by Elo when more labels are available, yet its relative strengths vary substantially with dataset properties. Efficiency comparisons further qualify the benefits of pretrained reuse: GVT and GraphPFN appear on the Pareto frontiers when supervised methods are represented by their default and fully tuned configurations. Adding intermediate tuning budgets removes this advantage for GVT and leaves GraphPFN extending the estimated frontier in the label-rich setting alone. Thus, reusing pretrained parameters does not yet provide a broadly reliable route to either stronger predictions or cheaper workflows. We release the evaluation pipeline, run-level records, and an open leaderboard at https://github.com/nums-ai/nodeground.
Figures & tables
| Dataset coverage | Comparison control | Eval. robustness | Cost accounting | Transparency | ||||
| Benchmark | Number of datasets | HPO control | Val. selection | Data splits | Multi- metric | Compute cost | Raw results | Leader- board |
| Yang et al. (2016) † | 4 | ❍ | 1 | |||||
| Shchur et al. (2018) † | 8 | ❍ | 100 | |||||
| Pei et al. (2020) † | 9 | ❍ | ❍ | 10 | ❍ | |||
| Platonov et al. (2026) † | 4 | ❍ | ❍ | 1 | ❍ ❍ ❍ | |||
| You et al. (2020) | 18 | ❍ | 3 | ❍ | ||||
| Group | Methods |
|---|---|
| Supervised | MLP (feature-only) , GCN ( Kipf and Welling, 2016 ) , GraphSAGE ( Hamilton et al., 2017 ) , GAT ( Veličković et al., 2017 ) , APPNP ( Gasteiger et al., 2018 ) , SGC ( Wu et al., 2019 ) , GCNII ( Chen et al., 2020 ) , FAGCN ( Bo et al., 2021 ) , GPRGNN ( Chien et al., 2020 ) , LINKX ( Lim et al., 2021 ) , GATv2 ( Brody et al., 2021 ) , NodeFormer ( Wu et al., 2022 ) , SGFormer ( Wu et al., 2023 ) , NAGphormer ( Chen et al., 2022 ) , Polynormer ( Deng et al., 2024 ) |
| Foundation | G2T-FM ( Eremeev et al., 2025a ) , GraphPFN ( Eremeev et al., 2025b ) , NodePFN ( Choi et al., 2026 ) GraphAny ( Zhao et al., 2025 ) , GVT ( Lee et al., 2025 ) , Node4All ( Lee and Yoo, 2026 ) |
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Cohort | Definition | Split | Datasets | Methods | Idx. | Cells | Imputed |
|---|---|---|---|---|---|---|---|
| V1 | Full benchmark | 10/10/80 | 51 | 21 | 5 | 5355 | 144 |
| V1 | Full benchmark | 50/25/25 | 51 | 21 | 5 | 5355 | 244 |
| V2 | Complete case | 10/10/80 | 28 | 21 | 5 | 2940 | 0 |
| V2 | Complete case | 50/25/25 | 28 | 21 | 5 | 2940 | 0 |
| V3 | Restricted method | 10/10/80 | 51 | 15 | 5 | 3825 | 10 |
| V3 | Restricted method | 50/25/25 | 51 | 15 | 5 | 3825 | 10 |
| Comparison | Coefficient | 10/10/80 | 50/25/25 | Interpretation |
|---|---|---|---|---|
| Multiclass accuracy vs. Macro-F1 | Spearman | 0.978 | 0.989 | Strong agreement |
| Binary AUROC vs. positive precision | Kendall | 0.519 | 0.664 | Metric-sensitive ordering |
| Split | Rank–Elo | Rank–MI | Elo–MI | Leader | Max. shift | Same top four |
|---|---|---|---|---|---|---|
| 10/10/80 | 0.990 | 0.962 | 0.969 | GCNII | 4 | No |
| 50/25/25 | 0.993 | 0.972 | 0.955 | GraphPFN | 6 | No |
| 10/10/80 | 50/25/25 | |||
| Method | Significant wins | Significant losses | Significant wins | Significant losses |
| GCNII | 11 | 0 | 8 | 0 |
| GraphPFN | 1 | 0 | 7 | 0 |
| GPRGNN | 6 | 0 | 6 | 0 |
| FAGCN | 6 | 0 | 8 | 0 |
| GCN | 5 | 0 | 7 | 0 |
| 10/10/80 | 50/25/25 | |||||
|---|---|---|---|---|---|---|
| Method | (pp) | Cost at ( ) | (pp) | Cost at ( ) | ||
| APPNP | 4.5 | 25 | 14.0 | 6.0 | 25 | 14.8 |
| FAGCN | 3.0 | 10 | 5.6 | 2.9 | 25 | 14.8 |
| GAT | 7.3 | 25 | 13.9 | 10.0 | 50 | 29.4 |
| GATv2 | 6.9 | 25 | 10.2 | 10.0 | 50 | 20.1 |
| GCN | 5.5 | 50 | 27.7 | 8.2 | 25 | 14.7 |
| Source channel | Candidate records | Eligible records | Retained tasks |
|---|---|---|---|
| PyG | 94 | 48 | 32 |
| HF TAG | 11 | 11 | 11 |
| Heterophilous | 11 | 7 | 7 |
| OGB | 1 | 1 | 1 |
| Total | 117 | 67 | 51 |
| Check | Criterion | Excluded | Retained | Typical exclusions or merges |
|---|---|---|---|---|
| Scope | Homogeneous, single-domain, single-graph node classification. | 16 | 101 | PPI, WikipediaNetwork:Crocodile, Twitch:DE |
| Target | Verified categorical single label per evaluated node. | 7 | 94 | Yelp, AttributedGraphDataset:PPI, AttributedGraphDataset:Facebook |
| Feature | Eligible observable node attributes; no missing, identity-only, topology-only, or disallowed legacy representations. | 15 | 79 | KarateClub, Roman-empire, Minesweeper |
| Scale | ; ; . | 12 | 67 | NELL, Reddit, Reddit2 |
| Canonicalization | Merge eligible loader records representing the same benchmark task. | 16 | 51 | CitationFull:Cora, CitationFull:PubMed, AttributedGraphDataset:Cora |
| Canonical task | Records | Merged loader records |
|---|---|---|
| actor | 2 | Actor; heterophilous_graph:actor |
| amazon_ratings | 2 | HeterophilousGraphDataset:Amazon-ratings; heterophilous_graph:amazon_ratings |
| chameleon | 2 | WikipediaNetwork:Chameleon; heterophilous_graph:chameleon |
| citeseer | 2 | Planetoid:CiteSeer; AttributedGraphDataset:CiteSeer |
| cora | 2 | Planetoid:Cora; AttributedGraphDataset:Cora |
| cornell | 2 | WebKB:Cornell; heterophilous_graph:cornell |
| Dataset | Channel | Domain | Task | Citation | License |
|---|---|---|---|---|---|
| actor | Heteroph. | Web | Actor category classification | ( Pei et al., 2020 ) | MIT |
| amazon_computer | PyG | Commerce | Product category classification | ( McAuley et al., 2015 ; Shchur et al., 2018 ) | MIT |
| amazon_photo | PyG | Commerce | Product category classification | ( McAuley et al., 2015 ; Shchur et al., 2018 ) | MIT |
| amazon_ratings | Heteroph. | Commerce | Binned mean product-rating classification | ( Platonov et al., 2023b ) | MIT |
| amherst41 | PyG | Social | Student gender classification | ( Traud et al., 2012 ; Lim et al., 2021 ) | not stated |
| artnet-exp | PyG | Social | Explicit-content creator classification | ( Bazhenov et al., 2026 ) | Apache-2.0 |
| Property and definition | Subgroup Boundary #data |
|---|---|
| Number of nodes Number of graph nodes. | Tiny 9 Small 14 Medium 19 Large 9 |
| Average degree Mean number of neighbors per node. | Low 20 Medium 18 High 13 |
| Label homophily Label agreement along edges, adjusted for chance. | 8 – 15 28 |
| Number of features Number of columns in the node-feature matrix. | Low 15 Medium 31 High 5 |
| Feature nonzero fraction Fraction of sampled finite feature entries that are nonzero. | Sparse 16 Medium 17 Dense 18 |
| Number of classes Number of nonempty classes among labeled nodes. | Binary 11 3–10 32 8 |
| Quantity | Definition |
|---|---|
| Edge count | : number of unique undirected non-self-loop edges. Reciprocal orientations are counted once. |
| Density | for ; zero for . |
| Degree quantiles | and : median and 99th percentile of nonzero node degrees. |
| Hub concentration | : ratio of the 99th-percentile to the median nonzero degree. |
| Feature-similarity AUROC | AUROC of cosine feature similarity for distinguishing same-label from different-label node pairs. Sampling is specified below. |
| Labeled-node count | : number of nodes with valid target labels. |
| Number of nodes | Average degree | Label homophily | ||||
|---|---|---|---|---|---|---|
| Dataset | Group | Group | Group | |||
| actor | 7,600 | Small | 7.02 | Medium | 0.003 | – |
| amazon_computer | 13,752 | Medium | 35.76 | High | 0.682 | |
| amazon_photo | 7,650 | Small | 31.13 | High | 0.785 | |
| amazon_ratings | 24,492 | Medium | 7.60 | Medium | 0.140 | – |
| amherst41 | 2,235 | Small | 81.39 | High | 0.060 | – |
| Number of features | Nonzero fraction | Classes | Class imbalance | |||||
|---|---|---|---|---|---|---|---|---|
| Dataset | Group | Group | Group | Group | ||||
| actor | 932 | Medium | 0.006 | Sparse | 5 | 3–10 | 2.30 | Medium |
| amazon_computer | 767 | Medium | 0.343 | Medium | 10 | 3–10 | 17.73 | High |
| amazon_photo | 745 | Medium | 0.344 | Medium | 8 | 3–10 | 5.86 | Medium |
| amazon_ratings | 300 | Low | 1.000 | Dense | 5 | 3–10 | 8.49 | Medium |
| amherst41 | 1,193 | Medium | 0.005 | Sparse | 2 | Binary | 1.00 | Low |
| Dataset | Hub | Density | Feature AUROC | Label entropy | ||||
|---|---|---|---|---|---|---|---|---|
| actor | 26,659 | 4.00 | 51.00 | 12.75 | 9.23e-04 | 0.503 | 7,600 | 0.977 |
| amazon_computer | 245,861 | 23.00 | 245.00 | 10.65 | 2.60e-03 | 0.542 | 13,752 | 0.813 |
| amazon_photo | 119,081 | 22.00 | 201.32 | 9.15 | 4.07e-03 | 0.566 | 7,650 | 0.926 |
| amazon_ratings | 93,050 | 5.00 | 31.00 | 6.20 | 3.10e-04 | 0.502 | 24,492 | 0.877 |
| amherst41 | 90,954 | 70.00 | 297.98 | 4.26 | 3.64e-02 | 0.500 | 2,032 | 1.000 |
| artnet-exp | 280,348 | 2.00 | 135.00 | 67.50 | 2.21e-04 | 0.521 | 50,405 | 0.469 |
| Method | Family | Implementation used | Configuration set |
|---|---|---|---|
| MLP | Feature-only | Benchmark implementation (PyTorch) | default + 200 sampled |
| APPNP | GNN | Benchmark wrapper using PyG APPNP | default + 200 sampled |
| FAGCN | GNN | Benchmark wrapper using PyG FAConv | default + 200 sampled |
| GAT | GNN | Benchmark wrapper using PyG GATConv | default + 200 sampled |
| GATv2 | GNN | Benchmark wrapper using PyG GATv2Conv | default + 200 sampled |
| GCN | GNN | Benchmark wrapper using PyG GCNConv | default + 200 sampled |
| Method | Artifact and source | Frozen / trainable | Target operation |
|---|---|---|---|
| G2T-FM | Official G2T-FM implementation with the TabPFNv2 classifier checkpoint | TabPFNv2 frozen; no trainable head | Constructs graph-derived tabular features and supplies training labels as context. Uses n_epochs=0 and no fine-tuning. |
| GraphAny | Official graph_any_arxiv.pt checkpoint | Fusion MLP frozen; no trainable head | Forms feature and propagated channels, analytically solves label-conditioned linear channels from the training labels, and fuses their logits. |
| GraphPFN | Official GraphPFN 1.3 release | Complete pretrained model frozen; no trainable head | Runs graph-aware label-conditioned inference and selects between PCA disabled and 64 dimensions. |
| GVT | Official RGVT seed_42.pth checkpoint | RGVT encoder frozen; target MLP head trained | Computes representations at depths 1–8 and trains a separate validation-selected head at each depth. |
| Node4All | Official cgt_enc.pth checkpoint | CGT encoder frozen; target MLP head trained | Computes one canonical representation and trains a validation-selected target-specific MLP head. |
| NodePFN | Official v1.0.0 checkpoint_epoch_30.ckpt | Complete pretrained model frozen; no trainable head | Uses training labels as context and selects dimensionality reduction, component count, smoothing, and ensemble settings. |
| Method(s) | Default | Varied values |
| Shared supervised training | Adam; learning rate , weight decay , 2,500 epochs, patience 100 | learning rate ; weight decay ; patience |
| APPNP | hidden 64; feature dropout .5; ; ; propagation dropout 0 | hidden ; feature dropout ; ; ; propagation dropout ; shared training |
| FAGCN | hidden 32; feature dropout .5; 2 layers; ; edge dropout .5 | hidden ; feature and edge dropout ; layers ; ; shared training |
| GAT, GATv2 | hidden 64; one head; feature dropout .5; attention dropout .2; 2 layers; no normalization, residual, or pre-transform | hidden ; heads ; feature dropout ; attention dropout ; layers ; normalization ; residual and pre-transform booleans; shared training |
| GCN, GraphSAGE | hidden 64; feature dropout .5; 2 layers; no normalization, residual, or pre-transform | hidden ; feature dropout ; layers ; normalization ; residual and pre-transform booleans; shared training |
| GCNII | hidden 64; feature dropout .6; 16 layers; ; ; shared weights; no normalization, residual, or pre-transform | hidden ; feature dropout ; layers ; ; ; normalization, residual, and pre-transform choices; shared training |
| Method | Trial records | Cells | Default selected | Median ID | Modal selected settings |
| APPNP | 95,403/102,510 | 510/510 | 2.5% | 101.5 | hidden 512; feature drop .7; ; ; propagation drop 0 |
| FAGCN | 102,510/102,510 | 510/510 | 0.2% | 104.5 | hidden 128; feature drop .7; 4 layers; ; edge drop .3 |
| GAT | 102,510/102,510 | 510/510 | 0.2% | 97 | hidden 64; 4 heads; feature/attention drop .3/.2; 1 layer; layer norm; residual on; pre-transform off |
| GATv2 | 92,862/92,862 | 462/510 | 0.0% | 99 | hidden 64; 4 heads; feature/attention drop .3/.2; 1 layer; layer norm; residual on; pre-transform off |
| GCN | 102,510/102,510 | 510/510 | 1.0% | 102 | hidden 512; feature drop .3; 2 layers; layer norm; residual on; pre-transform off |
| GCNII | 100,500/100,500 | 500/510 | 3.4% | 99.5 | hidden 512; feature drop .7; 32 layers; ; ; layer norm; residual and pre-transform off |
| task | primary | secondary | selection loss |
|---|---|---|---|
| Multiclass | Accuracy | Macro-F1; cross-entropy | Validation cross-entropy |
| Binary | AUROC | Accuracy; class-1 precision; binary cross-entropy | Validation binary cross-entropy |
| Dataset | Task | Best, 10/10/80 | Best, 50/25/25 |
|---|---|---|---|
| actor | M (Acc.) | GraphPFN 38.15 0.56 | GraphPFN 39.64 0.34 |
| amazon_computer | M (Acc.) | GraphPFN 91.09 0.27 | GraphPFN 93.04 0.36 |
| amazon_photo | M (Acc.) | GraphPFN 94.95 0.33 | GraphPFN 96.25 0.47 |
| amazon_ratings | M (Acc.) | GraphPFN 43.78 0.29 | LINKX 50.63 1.07 |
| amherst41 | B (AUROC) | LINKX 76.50 1.36 | LINKX 89.42 1.73 |
| artnet-exp | B (AUROC) | GraphPFN 84.60 0.33 | GraphPFN 86.75 0.27 |
| Configuration | Mean rank 10/10/80 | Mean rank 50/25/25 | Rank 10/10/80 | Rank 50/25/25 |
|---|---|---|---|---|
| GCNII [hp_tuned] | 6.4216 | 7.1667 | 1 | 1 |
| GPRGNN [hp_tuned] | 6.8039 | 7.3333 | 2 | 2 |
| FAGCN [hp_tuned] | 8.2353 | 8.4608 | 3 | 3 |
| GCN [hp_tuned] | 9.9216 | 9.1373 | 4 | 4 |
| GAT [hp_tuned] | 10.8431 | 10.7157 | 5 | 9 |
| GraphSAGE [hp_tuned] | 11.2843 | 9.8627 | 6 | 7 |
| Method | Observed 10/10/80 | Imputed 10/10/80 | Observed 50/25/25 | Imputed 50/25/25 |
|---|---|---|---|---|
| GATv2 | 46 | 5 | 46 | 5 |
| GCNII | 50 | 1 | 50 | 1 |
| NAGphormer | 50 | 1 | 50 | 1 |
| NodeFormer | 41 | 10 | 41 | 10 |
| G2T-FM | 44 | 7 | 30 | 21 |
| GraphAny | 50 | 1 | 50 | 1 |