While multi-agent systems (MAS) excel at complex reasoning, they are vulnerable to errors that intermediate agents introduce and downstream agents build upon. Auditing intermediate messages before they propagate requires an explicit standard, yet evaluation rubrics are typically authored by domain experts or written against a reference answer, neither of which is available for an unseen message at test time. We present MASRubric, a MAS information flow auditing framework with failure-distilled pitfall rubrics. Offline, trajectories on which the MAS has failed are automatically distilled into a reusable bank of pitfall criteria, each describing a recurrent error by its underlying misconception, the reasoning situations in which it arises, and the check that would expose it. Online, the criteria applicable to each intermediate message are retrieved from this off-the-shelf bank and checked one by one, and the resulting satisfaction rate decides whether the message is broadcast, returned to its author with diagnostic feedback for revision, or withheld. Empirical results demonstrate that MASRubric enhances MAS performance on both fixed and dynamic frameworks, achieving average accuracy gains of up to 2.83 points on math reasoning benchmarks and 1.74 points on code generation benchmarks. Further analysis shows that the retrieved criteria vary systematically with task types, and that the audit effort tracks task difficulty. Moreover, the bank transfers without re-mining to a system with a stronger backbone, which makes more adaptive and more efficient use of it. Our code and dataset are released at https://github.com/TonySY2/MASRubric.
Figures & tables
Figure 1: Existing rubrics (left) versus MASRubric (right) on each requirement.
Figure 2: Overview of MASRubric . The lower block shows the offline stage, where failure trajectories of a MAS are distilled by a rubric miner into criteria and compacted by dual-stage redundancy elimination into a reusable criterion bank. The upper block shows the online stage, where a context-specific rubric is retrieved for each intermediate message, the auditing agent judges the message criterion by criterion, and the resulting satisfaction rate decides whether the message passes and is broadcast, is revised by its author with diagnostic feedback, or is withheld.
Method
Easy Tasks
Hard Tasks
All Avg
GSM8K
MATH500
AQuA
AMC23
Avg
OlymB
AIME24
AIME25
OlymE
OlymH
Avg
Single Agent
87.64
74.80
83.86
62.50
77.20
47.56
13.33
20.00
20.00
16.00
23.38
47.30
+ CoT
93.71
79.40
84.65
67.50
81.32
43.85
20.00
23.33
24.00
15.00
25.24
50.16
Fixed-MAS
93.18
77.40
85.04
65.00
80.15
49.04
26.67
20.00
29.00
15.00
27.94
51.15
+ Self-Refine
93.03
77.00
83.46
65.00
79.62
50.67
23.33
20.00
27.00
22.00
28.60
51.28
+ Multi-TAG
93.78
76.20
83.07
70.00
80.76
47.85
26.67
23.33
19.00
17.00
26.77
50.77
Table 1: Performance of our method and the baselines across mathematical benchmarks, conducted within the fixed and dynamic MAS frameworks. The blue section (left) represents relatively easy tasks, while the orange section (right) indicates harder ones. Notably, MASRubric demonstrates the most significant performance improvements on these harder datasets. “OlymB”, “OlymE”, and “OlymH” represent OlympiadBench, OlymMATH Easy, and OlymMATH Hard, respectively.
Method
MBPP
HumE
CodeC
LiveC
Average
Single Agent
61.87
85.71
7.27
29.50
46.09
Dynamic-MAS
65.76
85.09
6.67
29.00
46.63
+ Self-Refine
64.20
83.23
6.06
32.50
46.50
+ Multi-TAG
66.15
80.75
6.06
29.50
45.61
+ RaR-Gen
64.98
82.61
6.06
29.25
45.73
+ MASRubric
68.48
85.71
7.27
32.00
48.37
Table 2: Performance comparison of our method against baselines across code domain benchmarks. “HumE”, “CodeC”, “LiveC” represent HumanEval, CodeContests, and LiveCodeBench.
Method
GSM8K
MATH500
AQuA
AMC23
OlymB
AIME24
AIME25
OlymE
OlymH
Average
MASRubric
91.66
79.60
83.86
70.00
52.44
30.00
26.67
32.00
17.00
53.69
(I) Max Revision Rounds ( Tmax , Default: 3)
1 Round
91.28
81.60
85.43
65.00
47.70
33.33
23.33
24.00
17.00
52.08
2 Rounds
91.74
80.80
86.22
62.50
49.63
33.33
23.33
25.00
18.00
52.28
4 Rounds
90.83
79.60
82.28
75.00
48.00
30.00
20.00
29.00
11.00
51.75
(II) Number of Retrieved Criteria ( Kctx , Default: 5)
Table 3: Results of the ablation study on the Dynamic-MAS framework.
Method
GSM8K
MATH500
AQuA
AMC23
OlymB
AIME24
AIME25
OlymE
OlymH
Average
Single Agent
87.64
79.40
84.19
62.50
52.15
23.33
20.00
20.00
16.00
49.47
Dynamic-MAS
93.48
80.80
86.22
72.50
51.41
26.67
20.00
31.00
11.00
52.56
+ MASRubric
93.93
82.20
87.01
77.50
53.93
43.33
26.67
27.00
15.00
56.29
Table 4: Performance with Qwen3-14B as the backbone, using the criterion bank mined from the Qwen3-8B system without re-mining.
Figure 3: Distribution of audit rounds needed to pass across benchmarks.
Figure 5: Run-time behavior of the auditing agent across the math suite, ordered from easy to hard. The Qwen3-14B system reuses the criterion bank mined from Qwen3-8B.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen3-8B
Qwen3-14B
Benchmark
Mean
0
1
2
3
4
5
Mean
0
1
2
3
4
5
GSM8K
2.840
0.3
2.5
24.0
62.6
7.4
3.3
2.535
5.7
5.9
35.9
40.3
5.7
6.4
MATH500
2.859
0.8
4.1
24.9
56.5
6.0
7.7
2.970
12.8
2.9
14.9
36.3
10.0
23.1
AQuA
2.552
1.7
6.7
37.7
45.6
4.9
3.3
2.702
14.3
4.3
20.4
35.8
8.1
17.1
AMC23
2.868
0.3
2.3
22.7
64.5
5.3
4.9
2.830
19.6
2.2
11.5
33.3
8.9
24.4
OlymB
3.013
0.8
2.5
21.5
56.7
6.9
11.7
3.371
10.1
1.9
10.4
31.1
11.3
35.2
Appendix
Table 5: Mean and distribution of the criterion number actually put in force per audit. Both systems use the same criterion bank, mined from the Qwen3-8B system.
Qwen3-8B
Qwen3-14B
Benchmark
Mean
0
1
2
3
Mean
0
1
2
3
GSM8K
0.077
94.2
4.4
0.8
0.5
0.082
92.4
7.2
0.3
0.1
MATH500
0.186
87.4
8.3
2.4
1.8
0.067
94.4
4.8
0.6
0.2
AQuA
0.271
82.2
11.5
3.2
3.1
0.149
87.4
10.8
1.4
0.5
AMC23
0.325
79.2
12.9
4.2
3.8
0.125
90.0
7.9
1.7
0.4
OlymB
0.481
70.1
17.7
6.3
6.0
0.238
81.1
15.5
2.0
1.4
Appendix
Table 6: Average revision rounds actually consumed per message. Column 0 denotes the acceptance at the first audit; the last column merges messages accepted after three repairs with those ultimately withheld.
Method
Math Avg
Total tok.
Tok./inst.
Cost
Δ Acc.
( ×106 )
( ×103 )
vs. vanilla
Dynamic-MAS
50.86
76.43
25.1
1.00 ×
—
+ Multi-TAG
51.31
69.71
22.9
0.91 ×
+0.45
+ PRM
51.94
159.07
52.2
2.08 ×
+1.08
+ RaR-Gen
51.94
144.85
47.5
1.90 ×
+1.08
+ Self-Refine
48.89
211.07
69.2
2.76 ×
− 1.97
Appendix
Table 7: End-to-end token accounting on the Dynamic-MAS math suite. “Generic Criterion” replaces the retrieved rubric with the single domain-wide criterion of Appendix A.2 .
Figure 6: Average accuracy versus token consumption on Math benchmarks by the dynamic MAS framework.
Domain
Dataset
Size
Test Set
Math
GSM8K ( Cobbe et al., 2021 )
1,319
MATH-500 Lightman et al. (2024)
500
AQuA ( Ling et al., 2017 )
254
AMC23
40
OlympiadBench ( He et al., 2024 )
675
Appendix
Table 8: Dataset statistics
Figure 7: An example criterion from the constructed bank for the math domain. The Diagnostic Definition and Applicability Condition drive retrieval and selection, while the Applicability Condition and Inspection Directive are the fields handed to the auditing agent at judgment time.
Figure 8: The generic criterion designed for the math domain.
Figure 9: The prompt template for the auditing agent in the math domain. The Area of Concern block is instantiated per criterion: the applicability condition and the inspection directive are concatenated and substituted into the slot.
Figure 10: The prompt template for the auditing agent in the code domain.
Figure 11: The prompt template for the rubric miner during offline bank construction.
Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually locate agent mistakes by inspecting agent trajectories (i.e., {\em failure attribution}) and provide feedback to refine the outputs (i.e., {\em repair}). Despite some recent work in MAS failure attribution, automated mechanisms to recover from such mistakes remain largely unexplored. To bridge this gap, we propose MARS, a search-based framework that formulates MAS repair as a Monte Carlo Tree Search (MCTS) process and navigates the vast space of potential repairs via diagnosis-guided expansion with taxonomy-augmented evaluation. Unlike standard MCTS, which evaluates a complete simulation via full rollout, MARS evaluates the agent trajectory using partial rollout to reduce token consumption. Furthermore, we introduce StateMAS, a large-scale MAS repair benchmark with 1,310 replayable multi-agent failure trajectories spanning four types of agent architectures and four LLM backbones. Experiments on StateMAS demonstrate that MARS consistently outperforms state-of-the-art methods, achieving an absolute improvement from 3.0% to 12.1% across all settings, while maintaining a comparable token consumption cost. The ablation study further confirms that taxonomy-augmented evaluation and diagnosis-guided expansion are critical to achieving these performance gains.
Multi-Agent Systems (MAS) built on Large Language Models (LLMs) often exhibit high variance in their reasoning trajectories. Process verification, which evaluates intermediate steps in trajectories, has shown promise in general reasoning settings, and has been suggested as a potential tool for guiding coordination of MAS; however, its actual effectiveness in MAS remains unclear. To fill this gap, we present MAS-ProVe, a systematic empirical study of process verification for multi-agent systems (MAS). Our study spans three verification paradigms (LLM-as-a-Judge, reward models, and process reward models), evaluated across two levels of verification granularity (agent-level and iteration-level). We further examine five representative verifiers and four context management strategies, and conduct experiments over six diverse MAS frameworks on multiple reasoning benchmarks. We find that process-level verification does not consistently improve performance and frequently exhibits high variance, highlighting the difficulty of reliably evaluating partial multi-agent trajectories. Among the methods studied, LLM-as-a-Judge generally outperforms reward-based approaches, with trained judges surpassing general-purpose LLMs. We further observe a small performance gap between LLMs acting as judges and as single agents, and identify a context-length-performance trade-off in verification. Overall, our results suggest that effective and robust process verification for MAS remains an open challenge, requiring further advances beyond current paradigms. Code is available at https://github.com/Wang-ML-Lab/MAS-ProVe.
Multi-agent systems (MAS) were expected to overcome the limitation of single-agent systems (SAS) through collaboration. However, under typicality conditions on the task's constraint graph and bounded inter-agent communication, we prove that the success probability of a MAS is closely tied to the connectivity of task constraints, where each agent has limited information-processing capacity. Specifically, the success probability decays exponentially with an information bottleneck that emerges from partitioning the task's constraint graph among agents. We define this quantity as the \emph{minimum cut cost} Cmin of the potential constraint graph of each task. This information-theoretic bound applies to both open systems with external feedback and closed systems without. We validate our theory on both synthetic experiments and real-world empirical data from SWE-bench submissions. From our framework, effective MAS design should incorporate task-inherent constraints alongside engineering optimization, and when \Cmin is high, practitioners should restructure tasks rather than simply scaling agents or communication.
Shi Pan, Ming Luo
University College London London, UK · University of Bristol Bristol, UK