LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.
Figures & tables
Figure 1: Comparison of evaluation scope and metric dimensions between traditional benchmarks and MASTraceBench.
Figure 2: Benchmark design in MASTraceBench. Red dashed boxes indicate the proposal-level outputs evaluated by S(⋅) .
Benchmark
Setting
Task Score (TS)
Graded Proposal Scoring Function S(⋅)
MAPF
157 samples
Mean arrival rate
Fraction of robots reaching their goals
OptionGen
235 samples
Mean option-set score
LLM-judge score of the correct option and distractor set
EvidenceChain
100 samples
Mean chain score
Set F1 and ordering agreement with ground truth
Gomoku
Dynamic
Win rate
Self-formation score plus defense-interception score
Negotiation
Dynamic
Mean episode payoff ratio
Payoff ratio of proposed allocation
TacticDuel
Dynamic
Win rate
Rule-based tactical value of the proposed action
Table 1: Tasks and evaluation specifications in MASTraceBench. The first three tasks are static tasks that require one-shot solutions for independent samples, while tasks marked as Dynamic require MASs to repeatedly observe the current state and output actions until an episode ends. Additional implementation details are provided in the supplementary material.
Figure 3: Overview of CLEARS.
Method
MAPF
OptionGen
EvidenceChain
Gomoku
Negotiation
TacticDuel
TS
CG
TS
CG
TS
CG
TS
CG
TS
CG
TS
CG
Single-Agent
69.66
–
80.56
–
74.02
–
50.00
–
41.00
–
50.00
–
Consistency ( Wang et al. 2023 )
72.91
3.25
81.83
1.27
75.19
1.17
50.00
0.00
42.33
1.33
55.00
5.00
MAV ( Lifshitz et al. 2025 )
77.13
7.47
82.19
1.63
74.40
0.38
46.00
-4.00
43.21
2.21
54.00
4.00
Int. Debate ( Du et al. 2024 )
77.31
7.65
83.70
3.14
77.94
3.92
80.00
30.00
44.08
3.08
64.00
14.00
Adv. Debate ( Liang et al. 2024 )
67.50
-2.16
81.32
0.76
65.83
-8.19
48.00
-2.00
43.28
2.28
21.00
-29.00
Table 2: Performance of MAS methods across benchmarks. TS and CG denote Task Score (0–100 scale) and Collaboration Gain over the Single-Agent baseline; Int. Debate and Adv. Debate refer to interactive and adversarial debate.
Figure 4: GAB analysis. Each bar reports the percentage distribution across GAB labels.
Figure 5: AR analysis on EvidenceChain and Gomoku. For each sample, Proposer Agents are grouped as AR-Strong if their initial proposal achieves the best initial score.
Figure 6: AL analysis. Lower AL indicates stronger aggregation capability of the MAS.
Method
MAPF
Gomoku
Single-Agent
1.00×
1.00×
Reflection
2.40×
2.38×
Consistency
3.04×
3.04×
Int. Debate
5.14×
4.64×
AgentVerse
5.75×
6.35×
Adv. Debate
5.79×
5.92×
Table 3: Token cost multipliers over the Single-Agent baseline on MAPF and Gomoku.
Setting
Int. Debate
MAPR
GI ↑
GR ↓
WI ↑
SR ↓
GI ↑
GR ↓
WI ↑
SR ↓
Main (3P+1R)
4.4
39.8
54.9
20.3
11.3
44.7
71.1
39.4
5P+1R
3.0
49.0
56.2
20.1
9.0
50.0
63.1
39.9
7P+1R
4.0
53.0
57.4
30.5
3.0
50.0
61.1
37.5
3P+2R
11.0
48.0
63.4
39.7
16.0
38.0
72.8
33.8
3P+3R
7.0
41.0
70.6
32.9
6.0
42.0
67.9
34.8
Table 4: Configuration ablation on EvidenceChain. P and R denote the numbers of Proposers and interaction rounds, respectively; GI/GR denote GAB Improve/Regress; WI/SR denote Surpass+Reach+Partial and Regress for initially weak and strong agents, respectively. Under 3P+1R, Scale-Het. and Family-Het. denote scale- and family-heterogeneous teams composed of Qwen3-{32B,14B,8B} and {Qwen3-8B, Llama-3.1-8B-Instruct, Mistral-8B-Instruct}, respectively.
Figure 7: A representative MAPF sample with two robots in a 5×5 bottleneck map. The visualization shows the obstacles, initial robot positions, and corresponding goals.
Figure 8: A representative OptionGen sample. In addition to the source article and reference question and options, each sample provides instance-specific criteria describing the desired correct option and the plausibility and incorrectness of the distractors.
Comparison
Pearson r
Spearman ρ
LLM Judge vs. Human Mean
0.919
0.818
Human–Human Average
0.784
0.661
Table 5: Agreement between the LLM judge and human evaluation on 550 OptionGen outputs from 50 sampled benchmark instances. Human–Human Average denotes the mean pairwise correlation among the three annotators.
Figure 9: A representative EvidenceChain sample. The task requires identifying and ordering the relevant facts from the indexed candidate set, while the gold reasoning chain provides the reference ordering used for evaluation.
Pattern
Open Endpoints
Self
Defense
Five
2 / 1
5 / 5
2 / 5
Four
2 / 1
4 / 2
4 / 3
Three
2 / 1
3 / 2
3 / 2
Other
–
0
0
Table 6: Rule-based scoring of proposed Gomoku moves. Paired values correspond to patterns with two and one open endpoints, respectively. Patterns with no open endpoint or at most two consecutive stones receive zero.
Benchmark
ρ
pρ
r
pr
Gomoku
0.884
0.0003
0.945
<0.0001
TacticDuel
0.882
0.0003
0.951
<0.0001
Table 7: Correlation between the mean proposal score and TS across the 11 evaluated MAS methods. ρ and r denote Spearman’s rank correlation and Pearson’s correlation, respectively.
Action
Energy cost
Base effect
Conditional interaction
Quick
0
Deals 2 damage.
None.
Heavy
0
Deals 4 damage.
Deals 0 damage if the opponent selects Counter .
Guard
0
Deals no damage.
Reduces incoming non- Feint damage by 4, lower-bounded by 0.
Counter
0
Deals 1 damage.
Deals 4 damage if the opponent selects Heavy .
Feint
0
Deals 1 damage.
Deals 4 damage against Guard and bypasses its damage reduction.
Charge
0
Gains 2 energy.
Deals no damage in the current round.
Table 8: Action effects in TacticDuel. The two players’ selected actions are resolved simultaneously in each round.
Method
MAPF
OptionGen
EvidenceChain
Gomoku
Negotiation
TacticDuel
OV
FPA
OV
FPA
OV
FPA
OV
FPA
OV
FPA
OV
FPA
Single-Agent
98.1
–
100.0
–
100.0
–
97.0
–
99.5
–
100.0
–
Consistency
97.7
3.5
100.0
23.0
100.0
9.0
97.0
35.5
99.9
6.0
100.0
67.2
MAV
99.6
6.0
100.0
19.1
99.7
7.7
99.5
32.3
100.0
11.5
100.0
66.2
Int. Debate
99.6
58.2
99.4
77.2
100.0
50.7
98.3
88.8
99.7
75.0
100.0
89.7
Agg. Debate
100.0
71.3
98.2
79.2
100.0
48.0
98.6
86.4
99.9
81.5
100.0
90.2
Table 9: Output Validity (OV, %) and Final-Proposal Agreement (FPA, %) across methods and benchmarks. A dash indicates that FPA is not applicable or is trivially determined by the method configuration.
Figure 10: Agent Refinement distributions of Int. Debate and MAPR on MAPF, OptionGen, Negotiation, and TacticDuel.
Method
OptionGen
EvidenceChain
Negotiation
TacticDuel
Single-Agent
1.00×
1.00×
1.00×
1.00×
Consistency
3.00×
2.99×
2.94×
2.96×
MAV
20.46×
20.64×
19.69×
21.60×
Int. Debate
8.85×
8.09×
7.98×
7.37×
Agg. Debate
10.92×
9.82×
9.62×
8.98×
Adv. Debate
15.84×
13.60×
16.12×
15.34×
Table 10: Additional token cost multipliers over the Single-Agent baseline on the four benchmarks not included in the main-text token cost results.
Figure 11: Task-prompt template for MAPF, including the optimization priorities, collision rules, and coordination guidance.
Figure 12: Task-prompt template for OptionGen, specifying the generation of one correct option and three plausible but incorrect distractors.
Figure 13: Task-prompt template for EvidenceChain, requiring a complete, minimal, and logically ordered supporting evidence chain.
Figure 14: Task-prompt template for Gomoku, including the current board, legal-placement constraints, and tactical priorities.
Figure 15: Task-prompt template for Negotiation, including private valuations, payoff calculation, interaction rules, and strategic guidance.
Figure 16: Task-prompt template for TacticDuel, including the current combat state, action effects, legal actions, and round history.