Relation-Aware Graph Foundation Model
Organizations: School of Data Science and Engineering, East China Normal University · Ant Group
Abstract
In recent years, large language models (LLMs) have demonstrated remarkable capability to generalize across diverse natural language processing tasks, inspiring the development of graph foundation models (GFMs) for large-scale pre-training. However, unlike language models with explicit token units, graphs lack a well-defined unit for generalization, making it challenging to design effective pre-training strategies. In this work, we propose REEF, a novel GFM framework that leverages relation tokens as the fundamental units. We construct a vocabulary of relation tokens to encode relational information within graphs. To accommodate diverse relations, we introduce two hypernetworks that adaptively generate the parameters of aggregators and classifiers in graph neural networks based on relation tokens. In addition, we design another hypernetwork to construct dataset-specific projectors and incorporate a dataset-level feature bias into the initial node representations, enhancing flexibility across different datasets with the same relation. Extensive experiments demonstrate that REEF consistently outperforms existing methods in both pre-training and transfer learning, highlighting its potential as a general-purpose graph foundation model.
Figures & tables
| Domain | Citation | WebKB | Amazon | Average | |||
|---|---|---|---|---|---|---|---|
| Dataset | Pubmed | Citeseer | Wisconsin | Texas | Photo | Acc | |
| GCN (ind) | 88.42 ( ) | 76.50 ( ) | 51.76 ( ) | 55.14 ( ) | 93.02 ( ) | 72.97 | 10.11 |
| GAT (ind) | 86.33 ( ) | 76.55 ( ) | 49.41 ( ) | 52.16 ( ) | 92.73 ( ) | 71.44 | 11.65 |
| GCN (joint) | 61.21 ( ) | 64.86 ( ) | 61.21 ( ) | 57.14 ( ) | 30.45 ( ) | 54.97 | 28.11 |
| GAT (joint) | 52.93 ( ) | 57.81 ( ) | 55.81 ( ) | 52.38 ( ) | 28.52 ( ) | 49.49 | 33.59 |
| RGCN (joint) | 80.30 ( ) | 71.77 ( ) | 72.09 ( ) | 80.95 ( ) | 79.03 ( ) | 76.83 | 6.25 |
| Method | Cora (Citation) | Cornell (WebKB) | Computers (Amazon) | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Acc | AUC | F1 | Acc | AUC | F1 | Acc | AUC | F1 | |
| GCN | |||||||||
| GAT | |||||||||
| GraphCL | |||||||||
| SimGRACE | |||||||||
| GCOPE+CL | |||||||||
| Domain | #Datasets | GAT | GraphAny | TS-Mean | REEF |
|---|---|---|---|---|---|
| Citation Networks | 4 | 64.23 | 70.14 | 63.39 | 71.08 |
| Co-authorship Networks | 4 | 77.63 | 81.82 | 81.03 | 83.00 |
| Social Networks | 6 | 69.00 | 69.03 | 68.31 | 70.25 |
| Co-purchase Graphs | 3 | 64.12 | 71.58 | 71.27 | 71.00 |
| Heterophilic Benchmarks | 6 | 72.85 | 69.34 | 68.53 | 76.77 |
| Air-traffic Graphs | 3 | 39.14 | 40.38 | 39.15 | 42.64 |
| Method | FT scale | BBBP | HIV |
|---|---|---|---|
| OFA Liu et al. (2023a) | N/A | – | 35.67 |
| MoMu Su et al. (2022) | N/A | 49.81 | 50.26 |
| Galactica Taylor et al. (2022) | N/A | 53.94 | 33.85 |
| GIMLET Zhao et al. (2023a) | 400M QA | 59.39 | 66.24 |
| GOFA Kong et al. (2024) | 100k QA | 54.91 | 53.02 |
| REEF | None | 54.80 | 63.17 |
| Method | Cora | History | Ratings |
|---|---|---|---|
| GraphMAE | 72.49 | 39.15 | 31.68 |
| LLaGA | 60.70 | 36.45 | 23.45 |
| OFA | 52.49 | 39.36 | 29.08 |
| OFA-FS | 42.10 | 17.50 | 20.50 |
| Prodigy (M) | 23.40 | 12.71 | 20.16 |
| Prodigy | 40.59 | 19.47 | 20.84 |
| Method | Pre-training datasets | Transfer datasets | Avg. | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| FB15K237 | WN18RR | Pubmed | Citeseer | Wisconsin | Texas | Photo | Cora | Cornell | Computers | Pre. | Trans. | All | |
| REEF-LM | 91.33 | 94.06 | 86.05 | 70.42 | 64.71 | 59.46 | 93.53 | 46.58 | 44.72 | 43.46 | 79.94 | 44.92 | 69.43 |
| REEF-FB | 89.90 | 92.18 | 86.23 | 70.27 | 70.59 | 54.05 | 92.16 | 40.07 | 44.51 | 42.09 | 79.34 | 42.22 | 68.21 |
| REEF-FP | 90.72 | 94.29 | 85.04 | 73.12 | 56.86 | 59.46 | 93.33 | 40.22 | 43.69 | 46.50 | 78.97 | 45.10 | 68.32 |
| REEF-AGU | 91.14 | 93.78 | 86.41 | 71.62 | 68.63 | 64.86 | 92.22 | 40.81 | 43.69 | 46.68 | 81.24 | 43.73 | 69.98 |
| REEF | 91.04 | 94.70 | 86.38 | 72.37 | 76.47 | 70.27 | 93.01 | 48.66 | 48.65 | 47.72 | 83.46 | 48.34 | 72.93 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Domain | #Nodes | #Edges | #Classes | Task | Type |
|---|---|---|---|---|---|---|
| FB15K237 | Knowledge | 14,541 | 310,116 | 237 | Link | |
| WN18RR | Knowledge | 40,943 | 93,003 | 11 | Link | |
| PubMed | Citation | 19,717 | 44,338 | 3 | Node | |
| Citeseer | Citation | 3,237 | 9,104 | 6 | Node | |
| Cora | Citation | 2,708 | 10,556 | 7 | Node | |
| Wisconsin | WebKB | 251 | 515 | 5 | Node |
| Dataset | Linear | GCN | GAT | GIN | DGI | BGRL | GraphMAE | GIANT | GFT | REEF |
|---|---|---|---|---|---|---|---|---|---|---|
| FB15K237 | 87.39 | 82.22 | 88.93 | 83.21 | 81.34 | 80.66 | 85.30 | 87.45 | 89.72 | 91.04 |
| WN18RR | 78.50 | 73.79 | 80.16 | 74.02 | 75.75 | 75.44 | 78.99 | 84.36 | 91.91 | 94.70 |
| Pre-training datasets | Cornell (#Labels = 5) | Computers (#Labels = 10) | ||||
|---|---|---|---|---|---|---|
| Acc | AUC | F1 | Acc | AUC | F1 | |
| Pubmed+Citeseer+Wisconsin+Texas+ | 48.65 7.76 | 70.16 2.82 | 36.49 6.77 | 47.72 5.61 | 84.80 3.22 | 44.03 2.49 |
| Photo+FB15K237+WN18RR (REEF) | ||||||
| Pubmed+Citeseer+FB15K237+WN18RR | 34.58 12.87 | 57.45 4.57 | 20.29 6.34 | 26.40 11.19 | 47.65 6.20 | 6.10 1.35 |
| Aggregator | Pre. | Trans. | All |
|---|---|---|---|
| Shared | 67.46 | 29.80 | 56.16 |
| Concat | 80.60 | 39.39 | 68.23 |
| Relation-ID | 83.24 | 33.51 | 68.32 |
| Cross-Attention | 65.10 | 34.29 | 55.86 |
| REEF (Hypernetwork) | 83.46 | 48.34 | 72.93 |
| Activation | FB15K237 | WN18RR | Wisconsin | Texas | Citeseer | Pubmed | Photo | Avg. |
|---|---|---|---|---|---|---|---|---|
| Linear | 91.04 | 94.70 | 76.47 | 70.27 | 72.37 | 86.38 | 93.01 | 83.46 |
| ReLU | 90.47 | 92.38 | 72.55 | 67.57 | 78.10 | 89.00 | 94.34 | 83.49 |
| GELU | 90.80 | 93.43 | 70.59 | 75.68 | 74.20 | 88.50 | 94.66 | 83.98 |
| Method | Paradigm | Trainable params | Training resources |
|---|---|---|---|
| GCOPE ( Zhao et al., 2024a ) | GNN-based | 93.3K | 34 min, 1 A800 80GB; 23.40 MiB peak |
| GOFA ( Kong et al., 2024 ) | LLM-based | 1.01B pretrain / 852.7M instruction tuning | Pre-training: 4 days, 4 A100 80GB |
| REEF | Hypernetwork + GNN-based | 177.4M | 8 h, 1 A800 80GB; 48.21 GiB peak |
| Dataset | H2GCN | FAGCN | GPR-GNN | LINKX | ACM-GCN | REEF |
|---|---|---|---|---|---|---|
| Cornell | 72.97 | 45.95 | 51.35 | 78.38 | 67.57 | 78.38 |
| Texas | 78.38 | 64.86 | 59.46 | 75.68 | 78.38 | 81.08 |
| Wisconsin | 82.35 | 66.67 | 70.59 | 62.75 | 70.59 | 84.31 |
| Chameleon | 57.46 | 62.06 | 61.18 | 72.59 | 60.75 | 56.80 |
| Actor | 37.96 | 36.32 | 34.74 | 29.67 | 37.24 | 33.82 |
| BlogCatalog | 80.65 | 86.52 | 89.32 | 60.91 | 74.98 | 76.93 |
| Dataset | Micro-F1 | Macro-F1 | AUROC |
|---|---|---|---|
| ACM | 84.94 | 85.03 | 94.46 |
| DBLP | 81.09 | 80.46 | 93.17 |
| Dataset | GAT | GraphAny | TS-Mean | REEF |
|---|---|---|---|---|
| Actor | 32.59 | 29.51 | 28.09 | 33.82 |
| AirBrazil | 35.38 | 36.15 | 39.23 | 44.44 |
| AirEU | 39.00 | 41.13 | 35.88 | 41.87 |
| AirUS | 43.03 | 43.86 | 42.34 | 41.62 |
| AmzComp | 70.94 | 82.00 | 81.37 | 77.71 |
| AmzPhoto | 80.78 | 90.18 | 90.18 | 93.83 |
| Domain | Relation Type | Description |
|---|---|---|
| Citation | Citation | A citation in a network represents the act of one scientific paper referencing another, embodying the flow and accumulation of knowledge. It reflects the topical relevance between papers and researchers’ acknowledgment of prior work, serving as a means to validate and support their own research while highlighting the collaborative nature of scientific inquiry. |
| Paper classification | Paper classification in graph learning focuses on predicting the categories or topics of individual papers within a graph. | |
| WebKB | Hyperlinks | A hyperlink represents a relationship or connection between nodes (e.g., entities such as documents, web pages, or words). It encapsulates semantic associations, enables information flow, and reflects the nature of interactions, which may vary in type, intensity, or direction. |
| Webpage classification | Webpage classification involves predicting the categories or types of individual webpages in a graph, often determined by the content, purpose, or role of the webpage within a network of hyperlinks. | |
| Amazon | Co-purchase | The “bought together” relationship between goods represents the co-purchasing behavior of customers, indicating a semantic association between products based on shared purchasing patterns. This relationship reflects the likelihood of two products being bought in conjunction, capturing their complementary nature or relevance in a shopping context. |
| Product classification | Product classification is a task in graph learning that predicts the categories or labels of product nodes, focusing on their classification based on attributes like type, function, or consumer category in a co-purchasing or recommendation graph. |