LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.
Figures & tables
Figure 1: Comparison of evaluation scope and metric dimensions between traditional benchmarks and MASTraceBench.
Figure 2: Benchmark design in MASTraceBench. Red dashed boxes indicate the proposal-level outputs evaluated by S(⋅) .
Benchmark
Setting
Task Score (TS)
Graded Proposal Scoring Function S(⋅)
MAPF
157 samples
Mean arrival rate
Fraction of robots reaching their goals
OptionGen
235 samples
Mean option-set score
LLM-judge score of the correct option and distractor set
EvidenceChain
100 samples
Mean chain score
Set F1 and ordering agreement with ground truth
Gomoku
Dynamic
Win rate
Self-formation score plus defense-interception score
Negotiation
Dynamic
Mean episode payoff ratio
Payoff ratio of proposed allocation
TacticDuel
Dynamic
Win rate
Rule-based tactical value of the proposed action
Table 1: Tasks and evaluation specifications in MASTraceBench. The first three tasks are static tasks that require one-shot solutions for independent samples, while tasks marked as Dynamic require MASs to repeatedly observe the current state and output actions until an episode ends. Additional implementation details are provided in the supplementary material.
Figure 3: Overview of CLEARS.
Method
MAPF
OptionGen
EvidenceChain
Gomoku
Negotiation
TacticDuel
TS
CG
TS
CG
TS
CG
TS
CG
TS
CG
TS
CG
Single-Agent
69.66
–
80.56
–
74.02
–
50.00
–
41.00
–
50.00
–
Consistency ( Wang et al. 2023 )
72.91
3.25
81.83
1.27
75.19
1.17
50.00
0.00
42.33
1.33
55.00
5.00
MAV ( Lifshitz et al. 2025 )
77.13
7.47
82.19
1.63
74.40
0.38
46.00
-4.00
43.21
2.21
54.00
4.00
Int. Debate ( Du et al. 2024 )
77.31
7.65
83.70
3.14
77.94
3.92
80.00
30.00
44.08
3.08
64.00
14.00
Adv. Debate ( Liang et al. 2024 )
67.50
-2.16
81.32
0.76
65.83
-8.19
48.00
-2.00
43.28
2.28
21.00
-29.00
Table 2: Performance of MAS methods across benchmarks. TS and CG denote Task Score (0–100 scale) and Collaboration Gain over the Single-Agent baseline; Int. Debate and Adv. Debate refer to interactive and adversarial debate.
Figure 4: GAB analysis. Each bar reports the percentage distribution across GAB labels.
Figure 5: AR analysis on EvidenceChain and Gomoku. For each sample, Proposer Agents are grouped as AR-Strong if their initial proposal achieves the best initial score.
Figure 6: AL analysis. Lower AL indicates stronger aggregation capability of the MAS.
Method
MAPF
Gomoku
Single-Agent
1.00×
1.00×
Reflection
2.40×
2.38×
Consistency
3.04×
3.04×
Int. Debate
5.14×
4.64×
AgentVerse
5.75×
6.35×
Adv. Debate
5.79×
5.92×
Table 3: Token cost multipliers over the Single-Agent baseline on MAPF and Gomoku.
Setting
Int. Debate
MAPR
GI ↑
GR ↓
WI ↑
SR ↓
GI ↑
GR ↓
WI ↑
SR ↓
Main (3P+1R)
4.4
39.8
54.9
20.3
11.3
44.7
71.1
39.4
5P+1R
3.0
49.0
56.2
20.1
9.0
50.0
63.1
39.9
7P+1R
4.0
53.0
57.4
30.5
3.0
50.0
61.1
37.5
3P+2R
11.0
48.0
63.4
39.7
16.0
38.0
72.8
33.8
3P+3R
7.0
41.0
70.6
32.9
6.0
42.0
67.9
34.8
Table 4: Configuration ablation on EvidenceChain. P and R denote the numbers of Proposers and interaction rounds, respectively; GI/GR denote GAB Improve/Regress; WI/SR denote Surpass+Reach+Partial and Regress for initially weak and strong agents, respectively. Under 3P+1R, Scale-Het. and Family-Het. denote scale- and family-heterogeneous teams composed of Qwen3-{32B,14B,8B} and {Qwen3-8B, Llama-3.1-8B-Instruct, Mistral-8B-Instruct}, respectively.
Figure 7: A representative MAPF sample with two robots in a 5×5 bottleneck map. The visualization shows the obstacles, initial robot positions, and corresponding goals.
Figure 8: A representative OptionGen sample. In addition to the source article and reference question and options, each sample provides instance-specific criteria describing the desired correct option and the plausibility and incorrectness of the distractors.
Comparison
Pearson r
Spearman ρ
LLM Judge vs. Human Mean
0.919
0.818
Human–Human Average
0.784
0.661
Table 5: Agreement between the LLM judge and human evaluation on 550 OptionGen outputs from 50 sampled benchmark instances. Human–Human Average denotes the mean pairwise correlation among the three annotators.
Figure 9: A representative EvidenceChain sample. The task requires identifying and ordering the relevant facts from the indexed candidate set, while the gold reasoning chain provides the reference ordering used for evaluation.
Pattern
Open Endpoints
Self
Defense
Five
2 / 1
5 / 5
2 / 5
Four
2 / 1
4 / 2
4 / 3
Three
2 / 1
3 / 2
3 / 2
Other
–
0
0
Table 6: Rule-based scoring of proposed Gomoku moves. Paired values correspond to patterns with two and one open endpoints, respectively. Patterns with no open endpoint or at most two consecutive stones receive zero.
Benchmark
ρ
pρ
r
pr
Gomoku
0.884
0.0003
0.945
<0.0001
TacticDuel
0.882
0.0003
0.951
<0.0001
Table 7: Correlation between the mean proposal score and TS across the 11 evaluated MAS methods. ρ and r denote Spearman’s rank correlation and Pearson’s correlation, respectively.
Action
Energy cost
Base effect
Conditional interaction
Quick
0
Deals 2 damage.
None.
Heavy
0
Deals 4 damage.
Deals 0 damage if the opponent selects Counter .
Guard
0
Deals no damage.
Reduces incoming non- Feint damage by 4, lower-bounded by 0.
Counter
0
Deals 1 damage.
Deals 4 damage if the opponent selects Heavy .
Feint
0
Deals 1 damage.
Deals 4 damage against Guard and bypasses its damage reduction.
Charge
0
Gains 2 energy.
Deals no damage in the current round.
Table 8: Action effects in TacticDuel. The two players’ selected actions are resolved simultaneously in each round.
Method
MAPF
OptionGen
EvidenceChain
Gomoku
Negotiation
TacticDuel
OV
FPA
OV
FPA
OV
FPA
OV
FPA
OV
FPA
OV
FPA
Single-Agent
98.1
–
100.0
–
100.0
–
97.0
–
99.5
–
100.0
–
Consistency
97.7
3.5
100.0
23.0
100.0
9.0
97.0
35.5
99.9
6.0
100.0
67.2
MAV
99.6
6.0
100.0
19.1
99.7
7.7
99.5
32.3
100.0
11.5
100.0
66.2
Int. Debate
99.6
58.2
99.4
77.2
100.0
50.7
98.3
88.8
99.7
75.0
100.0
89.7
Agg. Debate
100.0
71.3
98.2
79.2
100.0
48.0
98.6
86.4
99.9
81.5
100.0
90.2
Table 9: Output Validity (OV, %) and Final-Proposal Agreement (FPA, %) across methods and benchmarks. A dash indicates that FPA is not applicable or is trivially determined by the method configuration.
Figure 10: Agent Refinement distributions of Int. Debate and MAPR on MAPF, OptionGen, Negotiation, and TacticDuel.
Method
OptionGen
EvidenceChain
Negotiation
TacticDuel
Single-Agent
1.00×
1.00×
1.00×
1.00×
Consistency
3.00×
2.99×
2.94×
2.96×
MAV
20.46×
20.64×
19.69×
21.60×
Int. Debate
8.85×
8.09×
7.98×
7.37×
Agg. Debate
10.92×
9.82×
9.62×
8.98×
Adv. Debate
15.84×
13.60×
16.12×
15.34×
Table 10: Additional token cost multipliers over the Single-Agent baseline on the four benchmarks not included in the main-text token cost results.
Figure 11: Task-prompt template for MAPF, including the optimization priorities, collision rules, and coordination guidance.
Figure 12: Task-prompt template for OptionGen, specifying the generation of one correct option and three plausible but incorrect distractors.
Figure 13: Task-prompt template for EvidenceChain, requiring a complete, minimal, and logically ordered supporting evidence chain.
Figure 14: Task-prompt template for Gomoku, including the current board, legal-placement constraints, and tactical priorities.
Figure 15: Task-prompt template for Negotiation, including private valuations, payoff calculation, interaction rules, and strategic guidance.
Figure 16: Task-prompt template for TacticDuel, including the current combat state, action effects, legal actions, and round history.
Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across methods. Outcome-only benchmarks discard collaborations, whereas LLM-as-Judge evaluation requires additional, model-dependent inference and can vary with the LLM and rubric. We introduce a generalizable evaluation framework that maps native MAS traces into a shared space of unified collaboration graphs, enabling different methods to be evaluated under the same representation, reference set, and metric panel. Candidate graphs are compared with a query-specific reference forest. Each forest is a benchmark-provided collection of verified-success graphs: it records diverse ways in which representative MAS methods can complete the task, rather than prescribing a unique optimal process. Instantiating the framework as ForestBench, we filter 844 collaboration-necessary queries from seven public datasets, precompute ten successful target-conditioned reference graphs per query, and evaluate six representative MAS frameworks. Controlled backbone, reference-construction, and perturbation studies test the stability and scope of evaluation. Once the benchmark forests are built, ForestBench scores a trace in milliseconds without further LLM inference, providing a reusable structural basis for comparing diverse MAS collaboration traces.
Guo Chen, Ziwen Li, Reed Li +4
Southwest University Chongqing, China · Tencent Beijing, China · Institute of Computing Technology, Chinese Academy of Sciences Beijing, China
Large language model (LLM)-based Multi-agent systems (MAS) have shown promise in tackling complex collaborative tasks, where agents are typically orchestrated via role-specific prompts. While the quality of these prompts is pivotal, jointly optimizing them across interacting agents remains a non-trivial challenge, primarily due to the misalignment between local agent objectives and holistic system goals. To address this, we introduce MASPO, a novel framework designed to automatically and iteratively refine prompts across the entire system. A core innovation of MASPO is its joint evaluation mechanism, which assesses prompts not merely by their local validity, but by their capacity to facilitate downstream success for successor agents. This effectively bridges the gap between local interactions and global outcomes without relying on ground-truth labels. Furthermore, MASPO employs a data-driven evolutionary beam search to efficiently navigate the high-dimensional prompt space. Extensive empirical evaluations across 6 diverse tasks demonstrate that MASPO consistently outperforms state-of-the-art prompt optimization methods, achieving an average accuracy improvement of 2.9. We release our code at https://github.com/wangzx1219/MASPO.
Zhexuan Wang, Xuebo Liu, Li Wang +4
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China.
Multi-agent systems (MAS) built on large language models have shown growing promise, with their effectiveness resting on agents' ability to coordinate through text-based channels much as human teams do. Yet recent study suggests that MAS often falter not because agents lack individual task-solving ability, but because they lack collaborative competence: the capacity to establish common ground, maintain shared task understanding, balance individual and collective incentives, and repair misalignment as interaction unfolds. Decades of research in Computer-Supported Cooperative Work have characterized these requirements for human teams coordinating under constrained communication, yet existing MAS evaluations focus mainly on task outcomes or single-agent proficiency in reasoning, planning, and tool use. To enable a systematic analysis of agents' collaborative competence in MAS, we introduce CollabSim, a configurable simulation framework that combines a theory-grounded definition of collaborative capabilities, controlled manipulation of interaction conditions, and action-level probing of agents' internal states. Experiments across four LLMs show that CollabSim can capture condition effects, separate model performance patterns, and reveal task-dependent effects of agent design.