LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph learning remains unexplored. Graph learning is a natural setting for such systems because useful evidence may arise from heterogeneous local, long-range, global structural, and semantic perspectives whose relevance varies across instances. Existing LLM-based graph learning approaches primarily rely on single-agent reasoning, while multi-agent coordination has been studied mainly in general reasoning settings. Consequently, it remains unclear whether multiple specialized agents can improve graph learning and how coordination strategies should be designed and evaluated. To address this gap, we introduce GraphMAS, a systematic benchmark of multi-agent coordination for graph learning. GraphMAS builds a shared pool of graph reasoning specialists and organizes coordination along two dimensions, inter-agent interaction and runtime adaptivity, yielding four paradigms and seven representative coordination methods. Under a unified protocol, we evaluate these methods across seven text-attributed graphs, three domains, and two graph learning tasks. We find that heterogeneous graph perspectives are complementary, and that coordinating specialists improves over individual specialists and single-agent graph reasoning, with gains from decomposing reasoning across specialists rather than from broader evidence access alone. However, richer inter-agent interaction does not reliably help, whereas instance-adaptive specialist selection yields the strongest accuracy-efficiency trade-off. We further show that coordination can be learned over a fixed specialist pool and transfers to held-out graphs. GraphMAS therefore provides a controlled evaluation framework and empirical principles for understanding when and how multi-agent coordination benefits graph learning.
Figures & tables
Figure 1: Performance gap between the best specialist and oracle on five LP tasks. The oracle is correct if any specialist is correct; the best specialist is selected per dataset. Exact results are reported in Appendix C.1 .
Figure 2: Overview of GraphMAS. A shared pool of graph reasoning specialists is coordinated across representative multi-agent paradigms organized by inter-agent interaction and runtime adaptivity . The timeline shows prior LLM-based graph learning and general multi-agent systems.
Table 1: Coordination methods instantiated in GraphMAS.
Method
Link Prediction
Node Classification
Overall
arXiv
products
PubMed
Reddit
sports
arXiv-23
computers
Avg.
arXiv
products
PubMed
Reddit
sports
arXiv-23
computers
Avg.
Avg. Rank
Top 2
LLM Reasoning
Chain-of-Thought
76.1
63.8
80.4
55.3
64.8
78.0
63.7
68.9
56.5
66.1
91.6
60.1
53.8
60.1
59.4
63.9
10.8
0
Agentic Methods
Search-o1
63.8
79.5
69.8
64.7
79.1
65.3
66.6
69.8
57.3
68.1
89.1
62.0
59.7
60.4
60.9
65.4
11.0
0
Graph-CoT
60.2
78.6
60.9
60.4
70.8
62.0
62.2
65.0
54.5
64.6
70.5
56.1
57.2
59.7
62.6
60.7
13.7
0
Table 2: Accuracy (%) of LLM reasoning, agentic baselines, graph specialists, and multi-agent methods, with all LLM-based methods using Qwen2.5-32B-Instruct. Avg. averages seven datasets; Avg. Rank averages ranks across 15 methods and 14 dataset–task settings. Top 2 counts settings reaching either of the two highest distinct accuracies. Best and second-best values are bold and underlined , respectively. See App. B for implementation details.
Evidence
Role
LP
NC
(a) Specialist Design
✗
✗
73.0
66.0
✗
✓
69.9
65.4
✓
✗
73.7
66.1
✓
✓
80.0
67.7
(b) Reasoning decomposition
Table 3: Multi-agent gains under LLM Aggregation . ✓/✗ denote specialized/shared evidence or roles. Results are average accuracy (%).
Figure 3: Accuracy–cost trade-offs averaged across NC and LP.
Figure 4: LP accuracy (%) by initial specialist agreement. Parentheses show each group’s share.
Node Classification
Link Prediction
In-domain
Transfer
In-domain
Transfer
Category
Method
arXiv
products
PubMed
computers
Reddit
sports
arXiv-23
Avg.
arXiv
products
PubMed
computers
Reddit
sports
arXiv-23
Avg.
GNN
GCN
53.2
62.5
73.9
69.4
4.1
4.4
1.8
38.5
74.6
59.9
49.8
72.4
50.1
49.9
54.8
58.8
RevGAT
47.1
48.4
81.5
68.8
3.9
17.5
0.7
38.3
73.6
57.1
58.6
69.9
49.9
49.3
49.9
58.3
GraphSAGE
62.0
60.4
51.6
75.2
5.9
6.6
1.6
37.6
71.1
77.2
78.1
70.3
55.0
52.1
47.5
64.5
LLM-based GL
GraphPrompter
55.2
68.8
90.3
60.6
54.8
30.6
23.6
54.8
91.7
91.4
87.5
79.9
61.5
85.2
89.2
83.8
Table 4: Accuracy (%) of learned graph-learning methods under a shared in-domain/transfer split. All LLM-based baselines use Qwen2.5-7B-Instruct; GraphMAS uses a 7B router with four frozen Qwen2.5-32B-Instruct specialists. In-domain/Transfer denote seen/unseen datasets during training. Our method is highlighted . Best and second-best results are bold and underlined .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Dataset
#Nodes
#Edges
#Classes
Citation Network
ogbn-Arxiv
169,343
1,166,245
40
PubMed
19,717
44,338
3
Arxiv-2023
46,198
78,548
40
E-commerce
ogbn-Products (subset)
54,025
74,420
47
Amazon-Sports
173,055
1,946,555
13
Amazon-Computers
87,229
808,310
10
Appendix
Table 5: Statistics of the seven text-attributed graphs used in GraphMAS. #Classes denotes the node-classification label space; link prediction is formulated as binary prediction.
Hyperparameter
Value
Router initialization / reference
Qwen2.5-7B-Instruct
Frozen specialists
Qwen2.5-32B-Instruct
Training iterations
150
Trajectories per iteration
128
Maximum specialist calls
4
Actor / critic learning rate
5×10−7 / 1×10−5
Appendix
Table 6: Main hyperparameters for learned-router training.
Link Prediction
Node Classification
Dataset
Best Specialist
Oracle
Δ
Best Specialist
Oracle
Δ
arXiv
75.4
91.6
+16.2
59.4
74.4
+15.0
products
83.2
96.0
+12.8
72.8
80.6
+7.8
PubMed
81.2
98.2
+17.0
91.1
94.8
+3.7
Reddit
67.1
83.5
+16.4
68.3
74.8
+6.5
sports
80.7
97.0
+16.3
62.8
75.3
+12.5
Appendix
Table 7: Accuracy (%) of the best individual specialist and the specialist oracle. Best Specialist selects the highest-performing specialist separately for each dataset, and Oracle counts an instance as correct if any specialist is correct. Δ denotes the Oracle–Best Specialist gap in percentage points. Avg. denotes the mean across the seven datasets.
Method
arXiv
products
PubMed
Reddit
sports
arXiv-23
computers
Avg.
Link Prediction
GraphSearch-R
63.8
82.2
70.6
67.8
76.7
63.5
67.2
70.3
GoA-max
83.7
90.3
74.2
67.9
85.6
82.3
73.7
79.7
Node Classification
GraphSearch-R
57.6
70.8
90.4
63.5
61.9
52.8
68.1
66.4
GoA-max
59.3
71.4
92.2
66.2
59.2
60.9
68.3
68.2
Appendix
Table 8: Accuracy (%) of baseline variants. Avg. denotes the mean across the seven datasets.
Method
Evidence
Role
Link Prediction
Node Classification
arXiv
products
PubMed
Reddit
sports
arXiv-23
computers
Avg.
arXiv
products
PubMed
Reddit
sports
arXiv-23
computers
Avg.
Majority Vote
Shared
Generic
82
75
74
58
78
74
71
73.1
63
68
92
70
55
51
63
66.0
Shared
Specialized
83
77
70
62
76
80
66
73.4
64
66
91
72
55
51
62
65.9
Agent-specific
Generic
82
80
78
56
83
82
76
76.7
68
69
89
70
59
54
64
67.6
Agent-specific
Specialized
80
85
80
62
85
82
80
79.1
69
70
87
70
57
57
65
67.9
LLM Aggregation
Shared
Generic
77
76
72
67
69
82
68
73.0
61
68
88
70
57
56
62
66.0
Appendix
Table 9: Per-dataset accuracy (%) under different specialist designs. Shared/Agent-specific denote common/specialist-specific evidence; Generic/Specialized denote general/specialist-specific role instructions. Bold and underlined values indicate the highest and second-highest accuracies within each method and column.
Method
arXiv
products
PubMed
Reddit
sports
arXiv-23
computers
Avg.
Link Prediction
Single-agent reasoning
78
83
76
57
75
83
70
74.6
Multi-agent reasoning
87
85
77
68
89
77
77
80.0
Node Classification
Single-agent reasoning
65
68
88
64
57
54
63
65.6
Multi-agent reasoning
66
69
91
71
58
54
65
67.7
Appendix
Table 10: Per-dataset accuracy (%) of single-agent and multi-agent reasoning over the same four evidence sources. Single-agent reasoning receives all evidence sources jointly in one reasoning context; multi-agent reasoning uses LLM Aggregation. Avg. denotes the mean across the seven datasets. The best results within each column are shown in bold .
Link Prediction
Node Classification
Method
4/4 (54.5%)
3/4 (31.4%)
≤2/4 (14.1%)
4/4 (67.5%)
3/4 (19.0%)
≤2/4 (13.5%)
Majority
88.7
68.9
53.9
79.7
49.7
39.5
LLM Agg.
88.7
71.4
66.5
79.7
49.4
39.1
SOP
84.3
64.6
56.6
78.7
49.5
36.1
Debate
88.7
64.3
50.1
79.7
47.7
38.5
Routing
88.7
71.1
69.4
79.7
50.0
39.7
Appendix
Table 11: Accuracy (%) conditioned on agreement among the four graph specialists. Percentages in parentheses denote the share of instances in each agreement group. The best and second-best results within each group are shown in bold and underlined respectively.
Method
Qwen2.5-7B
GLM-4-32B
Link Prediction
Node Classification
Link Prediction
Node Classification
Reddit
sports
computers
Avg.
Reddit
sports
computers
Avg.
Reddit
sports
computers
Avg.
Reddit
sports
computers
Avg.
LLM Reasoning
Chain-of-Thought
52
60
60
57.3
44
32
52
42.7
49
65
54
56.0
63
68
54
61.7
Agentic Graph Learning
Graph-CoT
48
48
45
47.0
40
26
46
37.3
48
51
57
52.0
50
59
54
54.3
Appendix
Table 12: Backbone ablation on three datasets using 100 randomly sampled instances per dataset and task. Results report accuracy (%). Avg. is the mean across the three datasets. Best and second-best distinct values within each backbone and column are bold and underlined respectively.
Method
Specialist ops
LLM calls
Input tokens
Output tokens
GPU-s
Node Classification
Adaptive Routing
1.80
7.12
6,836
787
21.28
Majority Vote
4.00
8.93
9,395
1,364
33.53
LLM Aggregation
4.00
9.94
10,536
1,433
36.13
SOP Workflow
4.00
9.18
11,662
1,324
33.53
Multi-Agent Debate
5.32
11.93
13,823
1,842
42.57
Appendix
Table 13: Per-question computational costs on node classification (NC) and link prediction (LP). GPU-s denotes allocated A100 GPU-seconds. The minimum within each task is shown in bold .
Component
Stage
Core Prompt Instruction
Graph Specialist
Reasoning
You are the specialist role , a reasoning assistant for graph prediction. You have access to ONE graph search scope: assigned scope . Reason inside <think>...</think> ; use <search>...</search> when additional graph evidence is needed; after receiving <information>...</information> , reason over the new evidence before making the final prediction.
Proximal Neighborhood
Retrieval
Search the target node’s 1-hop and 2-hop neighbors using mode=local and hop=1|2 .
Distal Neighborhood
Retrieval
Search the target node’s 3-hop and 4-hop neighbors using mode=local and hop=3|4 .
Global Relevance
Retrieval
Search globally structure-relevant nodes selected by Personalized PageRank using mode=global .
Semantic Affinity
Retrieval
Search nodes whose attributes are semantically similar to the target node using mode=attribute .
Appendix
Table 14: Core prompt instructions for GraphMAS graph reasoning specialists. All specialists share the same reasoning protocol and differ in their admissible retrieval scope.
Method
Stage
Core Prompt Instruction
LLM Aggregation
Aggregation
You are the coordinator of a multi-agent graph prediction system and have no search tool. Aggregate the specialist judgments and target information. Prefer agreement supported by relevant evidence; weight specialists by confidence and evidence quality.
SOP Workflow
Handoff
Earlier specialists reported: preceding specialist reports . Use their findings as context, then add evidence from YOUR own scope and correct earlier conclusions when your evidence disagrees. The final specialist additionally weighs all previous findings and produces the pipeline’s final prediction.
Multi-Agent Debate
Revision
You previously predicted previous prediction . Other specialists reported: peer responses . Re-examine the case from YOUR scope. You may search again within your scope; revise your answer if persuaded by another specialist, otherwise defend it with evidence from your scope.
Appendix
Table 15: Core prompt instructions for static coordination methods. LLM Aggregation combines independent responses, while SOP and Debate expose specialists to intermediate peer judgments.
Method
Stage
Core Prompt Instruction
Adaptive Routing
Routing
You are the adaptive router of a multi-agent graph prediction system. Available specialists are: specialist roles and retrieval scopes . After reading the specialist reports collected so far, either CALL another specialist to gather or corroborate evidence, or ANSWER if the evidence is already sufficient. At most remaining calls additional specialist calls are allowed.
Adaptive Routing
Final synthesis
The routing budget is exhausted. Commit to a final prediction using the collected specialist reports: consulted specialist reports . Do not invoke graph retrieval.
Appendix
Table 16: Core prompts for Adaptive Routing. Specialists are consulted sequentially, while the router adapts participation and stopping to the current instance.
Method
Stage
Core Prompt Instruction
DyLAN
Ranking
Compare the candidate specialist responses according to evidence relevance, specificity, logical consistency, output validity, robustness to noisy graph context, and absence of unsupported assumptions. Do not solve the task or produce a new prediction. Rank the provided responses and retain the top- k specialists.
DyLAN
Revision
You are the retained specialist role . Reconsider your previous response using your own specialist evidence and other retained responses . Preserve your assigned role and do not claim access to graph evidence outside your scope. Incorporate useful peer evidence and explicitly resolve disagreements.
DyLAN
Final selection
Compare the retained revised responses and select the single response best supported by its stated graph evidence. Do not create a new answer or combine responses.
Appendix
Table 17: Core prompts for the DyLAN instantiation. Ranking and revision are triggered only when the initial specialist responses do not satisfy the early-agreement criterion.
Method
Stage
Core Prompt Instruction
Graph of Agents
Peer scoring
As the judge specialist , evaluate the other selected specialists according to evidence relevance, reasoning consistency, specificity, robustness, and output validity. Assign non-negative relevance scores that sum to 1 . Do not score yourself or produce a new prediction.
Graph of Agents
Source-to-target
As the target specialist , refine your response using messages from higher-relevance specialists. Treat peer messages as communicated conclusions rather than graph evidence you directly retrieved. Incorporate useful information and reject unsupported claims.
Graph of Agents
Target-to-source
As the source specialist , use the revised responses of lower-relevance specialists to finalize your own response. Preserve your original evidence scope and critically incorporate useful corrections or consensus.
Graph of Agents
Pooling
Produce the final prediction from refined specialist responses and relevance scores . Use relevance scores as guidance rather than proof of correctness, and compare the stated evidence and reasoning. Do not invoke graph tools or introduce new graph evidence.
Appendix
Table 18: Core prompts for Graph of Agents, covering peer scoring, bidirectional message passing, and final pooling.
Agentic graph learning (AGL) has recently achieved promising results on graph reasoning tasks, where an agent powered by a large language model (LLM) sequentially samples the graph as evidence to support its final prediction. Existing methods either employ a single agent or orchestrate multiple role-based agents to reason and learn over the entire graph, but both essentially rely on a shared reasoning policy across different graph regions, which can be suboptimal for graphs with heterogeneous structural and semantic patterns. Inspired by the progress of multi-agent collaboration on complex reasoning tasks, a natural remedy is to let multiple agents own different memory and collaborate; however, applying this paradigm to graphs directly faces two challenges. First, existing AGL methods typically verbalize graph structures into natural-language descriptions for LLM agents, making the reasoning process sensitive to the ordering of structural information and thereby breaking the permutation-invariant nature of graphs. Second, incorporating increasingly large sampled neighborhoods leads to rapidly growing contexts. To address these challenges, this paper introduces a multi-agent agentic graph learning (i.e., MAAGL) framework. MAAGL partitions the graph into communities and assigns an independent agent to each community for region-specific specialization. MAAGL represents structural and semantic evidence separately. Structural evidence is summarized by a dynamically updated structural signature that is permutation-invariant and fixed in size, while semantic evidence is filtered to the top-k nodes ranked by relevance. Based on historical trajectories with similar signatures, agents estimate their confidence and trigger debate-style collaboration when needed. Extensive experiments on four benchmark datasets show that MAAGL outperforms SOTA AGL methods.
Liang Qu, Jianxin Li, Hua Wang
Edith Cowan University, Perth, Australia. · Victoria University, Melbourne, Australia.
Multi-agent reinforcement learning (MARL) is crucial for AI systems that operate collaboratively in distributed and adversarial settings, particularly in multi-domain operations (MDO). A central challenge in cooperative MARL is determining how agents should coordinate: existing approaches must either hand-specify graph topology, rely on proximity-based heuristics, or learn structure entirely from environment interaction; all of which are brittle, semantically uninformed, or data-intensive. We investigate whether large language models (LLMs) can generate useful coordination graph priors for MARL by using minimal natural language descriptions of agent observations to infer latent coordination patterns. These priors are integrated into MARL algorithms via graph convolutional layers within a graph neural network (GNN)-based pipeline, and evaluated on four cooperative scenarios from the Multi-Agent Particle Environment (MPE) benchmark against baselines spanning the full spectrum of coordination modeling, from independent learners to state-of-the-art graph-based methods. We further ablate across five compact open-source LLMs to assess the sensitivity of prior quality to model choice. Our results provide the first quantitative evidence that LLM-derived graph priors can enhance coordination and adaptability in dynamic multi-agent environments, and demonstrate that models as small as 1.5B parameters are sufficient for effective prior generation.
Nikunj Gupta, Rajgopal Kannan, Viktor Prasanna
University of Southern California, Los Angeles, California, United States · DEVCOM ARL Army Research Office, Los Angeles, California, United States
LLM-based multi-agent systems (MAS) typically optimize a single topology, restricting reasoning to a narrow trajectory and limiting comprehensive analytical capacity. Naively merging multiple topologies into a composite graph introduces redundant noise propagation across irrelevant connections, degrading solution quality. To address this dilemma, we propose \textbf{Hierarchical Sparse Coordination over a Union of Complementary Topologies for MAS (HELENA)}, a multi-agent framework that balances diverse reasoning paths with sparse task-dependent execution. \helena{} constructs a union MAS graph from complementary candidate topologies selected via Monte Carlo Tree Search and Determinantal Point Process, broadening the reasoning trajectory for comprehensive analysis of complex problems. A Hierarchical Sparse Coordination module then activates only a sparse subgraph at each step while agents exchange compressed latent briefs to suppress redundant noise propagation. Finally, a Local Self-Refinement stage identifies decision units with discrepancy evidence and rewrites them only when contrastive evidence simultaneously confirms a reliable solution-side failure and a challenger-side improvement. Experiments across eight benchmarks show that \helena{} achieves state-of-the-art results on all benchmarks, with an average gain of \pctup{3.47} over the strongest baseline and up to \pctup{10.34} on MMLU-Pro, achieving larger improvements on harder benchmarks at a reasonable additional cost.
Zhifang Mao, Linyao Zheng, Xuhang Shi +1
1XiaoLab · 2Xi’an Jiaotong University · 3Beijing University of Posts and Telecommunications