GraphMAS: A Systematic Benchmark of Multi-Agent Coordination for Graph Learning
Organizations: New York University Shanghai · New York University · Northwestern University
Abstract
LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph learning remains unexplored. Graph learning is a natural setting for such systems because useful evidence may arise from heterogeneous local, long-range, global structural, and semantic perspectives whose relevance varies across instances. Existing LLM-based graph learning approaches primarily rely on single-agent reasoning, while multi-agent coordination has been studied mainly in general reasoning settings. Consequently, it remains unclear whether multiple specialized agents can improve graph learning and how coordination strategies should be designed and evaluated. To address this gap, we introduce GraphMAS, a systematic benchmark of multi-agent coordination for graph learning. GraphMAS builds a shared pool of graph reasoning specialists and organizes coordination along two dimensions, inter-agent interaction and runtime adaptivity, yielding four paradigms and seven representative coordination methods. Under a unified protocol, we evaluate these methods across seven text-attributed graphs, three domains, and two graph learning tasks. We find that heterogeneous graph perspectives are complementary, and that coordinating specialists improves over individual specialists and single-agent graph reasoning, with gains from decomposing reasoning across specialists rather than from broader evidence access alone. However, richer inter-agent interaction does not reliably help, whereas instance-adaptive specialist selection yields the strongest accuracy-efficiency trade-off. We further show that coordination can be learned over a fixed specialist pool and transfers to held-out graphs. GraphMAS therefore provides a controlled evaluation framework and empirical principles for understanding when and how multi-agent coordination benefits graph learning.
Figures & tables
| Method | Link Prediction | Node Classification | Overall | |||||||||||||||
| arXiv | products | PubMed | sports | arXiv-23 | computers | Avg. | arXiv | products | PubMed | sports | arXiv-23 | computers | Avg. | Avg. Rank | Top 2 | |||
| LLM Reasoning | ||||||||||||||||||
| Chain-of-Thought | 76.1 | 63.8 | 80.4 | 55.3 | 64.8 | 78.0 | 63.7 | 68.9 | 56.5 | 66.1 | 91.6 | 60.1 | 53.8 | 60.1 | 59.4 | 63.9 | 10.8 | 0 |
| Agentic Methods | ||||||||||||||||||
| Search-o1 | 63.8 | 79.5 | 69.8 | 64.7 | 79.1 | 65.3 | 66.6 | 69.8 | 57.3 | 68.1 | 89.1 | 62.0 | 59.7 | 60.4 | 60.9 | 65.4 | 11.0 | 0 |
| Graph-CoT | 60.2 | 78.6 | 60.9 | 60.4 | 70.8 | 62.0 | 62.2 | 65.0 | 54.5 | 64.6 | 70.5 | 56.1 | 57.2 | 59.7 | 62.6 | 60.7 | 13.7 | 0 |
| Evidence | Role | LP | NC |
|---|---|---|---|
| (a) Specialist Design | |||
| ✗ | ✗ | 73.0 | 66.0 |
| ✗ | ✓ | 69.9 | 65.4 |
| ✓ | ✗ | 73.7 | 66.1 |
| ✓ | ✓ | 80.0 | 67.7 |
| (b) Reasoning decomposition | |||
| Node Classification | Link Prediction | ||||||||||||||||
| In-domain | Transfer | In-domain | Transfer | ||||||||||||||
| Category | Method | arXiv | products | PubMed | computers | sports | arXiv-23 | Avg. | arXiv | products | PubMed | computers | sports | arXiv-23 | Avg. | ||
| GNN | GCN | 53.2 | 62.5 | 73.9 | 69.4 | 4.1 | 4.4 | 1.8 | 38.5 | 74.6 | 59.9 | 49.8 | 72.4 | 50.1 | 49.9 | 54.8 | 58.8 |
| RevGAT | 47.1 | 48.4 | 81.5 | 68.8 | 3.9 | 17.5 | 0.7 | 38.3 | 73.6 | 57.1 | 58.6 | 69.9 | 49.9 | 49.3 | 49.9 | 58.3 | |
| GraphSAGE | 62.0 | 60.4 | 51.6 | 75.2 | 5.9 | 6.6 | 1.6 | 37.6 | 71.1 | 77.2 | 78.1 | 70.3 | 55.0 | 52.1 | 47.5 | 64.5 | |
| LLM-based GL | GraphPrompter | 55.2 | 68.8 | 90.3 | 60.6 | 54.8 | 30.6 | 23.6 | 54.8 | 91.7 | 91.4 | 87.5 | 79.9 | 61.5 | 85.2 | 89.2 | 83.8 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Dataset | #Nodes | #Edges | #Classes |
| Citation Network | ogbn-Arxiv | 169,343 | 1,166,245 | 40 |
| PubMed | 19,717 | 44,338 | 3 | |
| Arxiv-2023 | 46,198 | 78,548 | 40 | |
| E-commerce | ogbn-Products (subset) | 54,025 | 74,420 | 47 |
| Amazon-Sports | 173,055 | 1,946,555 | 13 | |
| Amazon-Computers | 87,229 | 808,310 | 10 |
| Hyperparameter | Value |
|---|---|
| Router initialization / reference | Qwen2.5-7B-Instruct |
| Frozen specialists | Qwen2.5-32B-Instruct |
| Training iterations | 150 |
| Trajectories per iteration | 128 |
| Maximum specialist calls | 4 |
| Actor / critic learning rate | / |
| Link Prediction | Node Classification | |||||
|---|---|---|---|---|---|---|
| Dataset | Best Specialist | Oracle | Best Specialist | Oracle | ||
| arXiv | 75.4 | 91.6 | 59.4 | 74.4 | ||
| products | 83.2 | 96.0 | 72.8 | 80.6 | ||
| PubMed | 81.2 | 98.2 | 91.1 | 94.8 | ||
| 67.1 | 83.5 | 68.3 | 74.8 | |||
| sports | 80.7 | 97.0 | 62.8 | 75.3 | ||
| Method | arXiv | products | PubMed | sports | arXiv-23 | computers | Avg. | |
|---|---|---|---|---|---|---|---|---|
| Link Prediction | ||||||||
| GraphSearch-R | 63.8 | 82.2 | 70.6 | 67.8 | 76.7 | 63.5 | 67.2 | 70.3 |
| GoA-max | 83.7 | 90.3 | 74.2 | 67.9 | 85.6 | 82.3 | 73.7 | 79.7 |
| Node Classification | ||||||||
| GraphSearch-R | 57.6 | 70.8 | 90.4 | 63.5 | 61.9 | 52.8 | 68.1 | 66.4 |
| GoA-max | 59.3 | 71.4 | 92.2 | 66.2 | 59.2 | 60.9 | 68.3 | 68.2 |
| Method | Evidence | Role | Link Prediction | Node Classification | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| arXiv | products | PubMed | sports | arXiv-23 | computers | Avg. | arXiv | products | PubMed | sports | arXiv-23 | computers | Avg. | |||||
| Majority Vote | Shared | Generic | 82 | 75 | 74 | 58 | 78 | 74 | 71 | 73.1 | 63 | 68 | 92 | 70 | 55 | 51 | 63 | 66.0 |
| Shared | Specialized | 83 | 77 | 70 | 62 | 76 | 80 | 66 | 73.4 | 64 | 66 | 91 | 72 | 55 | 51 | 62 | 65.9 | |
| Agent-specific | Generic | 82 | 80 | 78 | 56 | 83 | 82 | 76 | 76.7 | 68 | 69 | 89 | 70 | 59 | 54 | 64 | 67.6 | |
| Agent-specific | Specialized | 80 | 85 | 80 | 62 | 85 | 82 | 80 | 79.1 | 69 | 70 | 87 | 70 | 57 | 57 | 65 | 67.9 | |
| LLM Aggregation | Shared | Generic | 77 | 76 | 72 | 67 | 69 | 82 | 68 | 73.0 | 61 | 68 | 88 | 70 | 57 | 56 | 62 | 66.0 |
| Method | arXiv | products | PubMed | sports | arXiv-23 | computers | Avg. | |
| Link Prediction | ||||||||
| Single-agent reasoning | 78 | 83 | 76 | 57 | 75 | 83 | 70 | 74.6 |
| Multi-agent reasoning | 87 | 85 | 77 | 68 | 89 | 77 | 77 | 80.0 |
| Node Classification | ||||||||
| Single-agent reasoning | 65 | 68 | 88 | 64 | 57 | 54 | 63 | 65.6 |
| Multi-agent reasoning | 66 | 69 | 91 | 71 | 58 | 54 | 65 | 67.7 |
| Link Prediction | Node Classification | |||||
|---|---|---|---|---|---|---|
| Method | (54.5%) | (31.4%) | (14.1%) | (67.5%) | (19.0%) | (13.5%) |
| Majority | 88.7 | 68.9 | 53.9 | 79.7 | 49.7 | 39.5 |
| LLM Agg. | 88.7 | 71.4 | 66.5 | 79.7 | 49.4 | 39.1 |
| SOP | 84.3 | 64.6 | 56.6 | 78.7 | 49.5 | 36.1 |
| Debate | 88.7 | 64.3 | 50.1 | 79.7 | 47.7 | 38.5 |
| Routing | 88.7 | 71.1 | 69.4 | 79.7 | 50.0 | 39.7 |
| Method | Qwen2.5-7B | GLM-4-32B | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Link Prediction | Node Classification | Link Prediction | Node Classification | |||||||||||||
| sports | computers | Avg. | sports | computers | Avg. | sports | computers | Avg. | sports | computers | Avg. | |||||
| LLM Reasoning | ||||||||||||||||
| Chain-of-Thought | 52 | 60 | 60 | 57.3 | 44 | 32 | 52 | 42.7 | 49 | 65 | 54 | 56.0 | 63 | 68 | 54 | 61.7 |
| Agentic Graph Learning | ||||||||||||||||
| Graph-CoT | 48 | 48 | 45 | 47.0 | 40 | 26 | 46 | 37.3 | 48 | 51 | 57 | 52.0 | 50 | 59 | 54 | 54.3 |
| Method | Specialist ops | LLM calls | Input tokens | Output tokens | GPU-s |
|---|---|---|---|---|---|
| Node Classification | |||||
| Adaptive Routing | 1.80 | 7.12 | 6,836 | 787 | 21.28 |
| Majority Vote | 4.00 | 8.93 | 9,395 | 1,364 | 33.53 |
| LLM Aggregation | 4.00 | 9.94 | 10,536 | 1,433 | 36.13 |
| SOP Workflow | 4.00 | 9.18 | 11,662 | 1,324 | 33.53 |
| Multi-Agent Debate | 5.32 | 11.93 | 13,823 | 1,842 | 42.57 |
| Component | Stage | Core Prompt Instruction |
|---|---|---|
| Graph Specialist | Reasoning | You are the specialist role , a reasoning assistant for graph prediction. You have access to ONE graph search scope: assigned scope . Reason inside <think>...</think> ; use <search>...</search> when additional graph evidence is needed; after receiving <information>...</information> , reason over the new evidence before making the final prediction. |
| Proximal Neighborhood | Retrieval | Search the target node’s 1-hop and 2-hop neighbors using mode=local and hop=1|2 . |
| Distal Neighborhood | Retrieval | Search the target node’s 3-hop and 4-hop neighbors using mode=local and hop=3|4 . |
| Global Relevance | Retrieval | Search globally structure-relevant nodes selected by Personalized PageRank using mode=global . |
| Semantic Affinity | Retrieval | Search nodes whose attributes are semantically similar to the target node using mode=attribute . |
| Method | Stage | Core Prompt Instruction |
|---|---|---|
| LLM Aggregation | Aggregation | You are the coordinator of a multi-agent graph prediction system and have no search tool. Aggregate the specialist judgments and target information. Prefer agreement supported by relevant evidence; weight specialists by confidence and evidence quality. |
| SOP Workflow | Handoff | Earlier specialists reported: preceding specialist reports . Use their findings as context, then add evidence from YOUR own scope and correct earlier conclusions when your evidence disagrees. The final specialist additionally weighs all previous findings and produces the pipeline’s final prediction. |
| Multi-Agent Debate | Revision | You previously predicted previous prediction . Other specialists reported: peer responses . Re-examine the case from YOUR scope. You may search again within your scope; revise your answer if persuaded by another specialist, otherwise defend it with evidence from your scope. |
| Method | Stage | Core Prompt Instruction |
|---|---|---|
| Adaptive Routing | Routing | You are the adaptive router of a multi-agent graph prediction system. Available specialists are: specialist roles and retrieval scopes . After reading the specialist reports collected so far, either CALL another specialist to gather or corroborate evidence, or ANSWER if the evidence is already sufficient. At most remaining calls additional specialist calls are allowed. |
| Adaptive Routing | Final synthesis | The routing budget is exhausted. Commit to a final prediction using the collected specialist reports: consulted specialist reports . Do not invoke graph retrieval. |
| Method | Stage | Core Prompt Instruction |
|---|---|---|
| DyLAN | Ranking | Compare the candidate specialist responses according to evidence relevance, specificity, logical consistency, output validity, robustness to noisy graph context, and absence of unsupported assumptions. Do not solve the task or produce a new prediction. Rank the provided responses and retain the top- specialists. |
| DyLAN | Revision | You are the retained specialist role . Reconsider your previous response using your own specialist evidence and other retained responses . Preserve your assigned role and do not claim access to graph evidence outside your scope. Incorporate useful peer evidence and explicitly resolve disagreements. |
| DyLAN | Final selection | Compare the retained revised responses and select the single response best supported by its stated graph evidence. Do not create a new answer or combine responses. |
| Method | Stage | Core Prompt Instruction |
|---|---|---|
| Graph of Agents | Peer scoring | As the judge specialist , evaluate the other selected specialists according to evidence relevance, reasoning consistency, specificity, robustness, and output validity. Assign non-negative relevance scores that sum to . Do not score yourself or produce a new prediction. |
| Graph of Agents | Source-to-target | As the target specialist , refine your response using messages from higher-relevance specialists. Treat peer messages as communicated conclusions rather than graph evidence you directly retrieved. Incorporate useful information and reject unsupported claims. |
| Graph of Agents | Target-to-source | As the source specialist , use the revised responses of lower-relevance specialists to finalize your own response. Preserve your original evidence scope and critically incorporate useful corrections or consensus. |
| Graph of Agents | Pooling | Produce the final prediction from refined specialist responses and relevance scores . Use relevance scores as guidance rather than proof of correctness, and compare the stated evidence and reasoning. Do not invoke graph tools or introduce new graph evidence. |