While multi-agent systems (MAS) excel at complex reasoning, they are vulnerable to errors that intermediate agents introduce and downstream agents build upon. Auditing intermediate messages before they propagate requires an explicit standard, yet evaluation rubrics are typically authored by domain experts or written against a reference answer, neither of which is available for an unseen message at test time. We present MASRubric, a MAS information flow auditing framework with failure-distilled pitfall rubrics. Offline, trajectories on which the MAS has failed are automatically distilled into a reusable bank of pitfall criteria, each describing a recurrent error by its underlying misconception, the reasoning situations in which it arises, and the check that would expose it. Online, the criteria applicable to each intermediate message are retrieved from this off-the-shelf bank and checked one by one, and the resulting satisfaction rate decides whether the message is broadcast, returned to its author with diagnostic feedback for revision, or withheld. Empirical results demonstrate that MASRubric enhances MAS performance on both fixed and dynamic frameworks, achieving average accuracy gains of up to 2.83 points on math reasoning benchmarks and 1.74 points on code generation benchmarks. Further analysis shows that the retrieved criteria vary systematically with task types, and that the audit effort tracks task difficulty. Moreover, the bank transfers without re-mining to a system with a stronger backbone, which makes more adaptive and more efficient use of it. Our code and dataset are released at https://github.com/TonySY2/MASRubric.
Figures & tables
Figure 1: Existing rubrics (left) versus MASRubric (right) on each requirement.
Figure 2: Overview of MASRubric . The lower block shows the offline stage, where failure trajectories of a MAS are distilled by a rubric miner into criteria and compacted by dual-stage redundancy elimination into a reusable criterion bank. The upper block shows the online stage, where a context-specific rubric is retrieved for each intermediate message, the auditing agent judges the message criterion by criterion, and the resulting satisfaction rate decides whether the message passes and is broadcast, is revised by its author with diagnostic feedback, or is withheld.
Method
Easy Tasks
Hard Tasks
All Avg
GSM8K
MATH500
AQuA
AMC23
Avg
OlymB
AIME24
AIME25
OlymE
OlymH
Avg
Single Agent
87.64
74.80
83.86
62.50
77.20
47.56
13.33
20.00
20.00
16.00
23.38
47.30
+ CoT
93.71
79.40
84.65
67.50
81.32
43.85
20.00
23.33
24.00
15.00
25.24
50.16
Fixed-MAS
93.18
77.40
85.04
65.00
80.15
49.04
26.67
20.00
29.00
15.00
27.94
51.15
+ Self-Refine
93.03
77.00
83.46
65.00
79.62
50.67
23.33
20.00
27.00
22.00
28.60
51.28
+ Multi-TAG
93.78
76.20
83.07
70.00
80.76
47.85
26.67
23.33
19.00
17.00
26.77
50.77
Table 1: Performance of our method and the baselines across mathematical benchmarks, conducted within the fixed and dynamic MAS frameworks. The blue section (left) represents relatively easy tasks, while the orange section (right) indicates harder ones. Notably, MASRubric demonstrates the most significant performance improvements on these harder datasets. “OlymB”, “OlymE”, and “OlymH” represent OlympiadBench, OlymMATH Easy, and OlymMATH Hard, respectively.
Method
MBPP
HumE
CodeC
LiveC
Average
Single Agent
61.87
85.71
7.27
29.50
46.09
Dynamic-MAS
65.76
85.09
6.67
29.00
46.63
+ Self-Refine
64.20
83.23
6.06
32.50
46.50
+ Multi-TAG
66.15
80.75
6.06
29.50
45.61
+ RaR-Gen
64.98
82.61
6.06
29.25
45.73
+ MASRubric
68.48
85.71
7.27
32.00
48.37
Table 2: Performance comparison of our method against baselines across code domain benchmarks. “HumE”, “CodeC”, “LiveC” represent HumanEval, CodeContests, and LiveCodeBench.
Method
GSM8K
MATH500
AQuA
AMC23
OlymB
AIME24
AIME25
OlymE
OlymH
Average
MASRubric
91.66
79.60
83.86
70.00
52.44
30.00
26.67
32.00
17.00
53.69
(I) Max Revision Rounds ( Tmax , Default: 3)
1 Round
91.28
81.60
85.43
65.00
47.70
33.33
23.33
24.00
17.00
52.08
2 Rounds
91.74
80.80
86.22
62.50
49.63
33.33
23.33
25.00
18.00
52.28
4 Rounds
90.83
79.60
82.28
75.00
48.00
30.00
20.00
29.00
11.00
51.75
(II) Number of Retrieved Criteria ( Kctx , Default: 5)
Table 3: Results of the ablation study on the Dynamic-MAS framework.
Method
GSM8K
MATH500
AQuA
AMC23
OlymB
AIME24
AIME25
OlymE
OlymH
Average
Single Agent
87.64
79.40
84.19
62.50
52.15
23.33
20.00
20.00
16.00
49.47
Dynamic-MAS
93.48
80.80
86.22
72.50
51.41
26.67
20.00
31.00
11.00
52.56
+ MASRubric
93.93
82.20
87.01
77.50
53.93
43.33
26.67
27.00
15.00
56.29
Table 4: Performance with Qwen3-14B as the backbone, using the criterion bank mined from the Qwen3-8B system without re-mining.
Figure 3: Distribution of audit rounds needed to pass across benchmarks.
Figure 5: Run-time behavior of the auditing agent across the math suite, ordered from easy to hard. The Qwen3-14B system reuses the criterion bank mined from Qwen3-8B.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen3-8B
Qwen3-14B
Benchmark
Mean
0
1
2
3
4
5
Mean
0
1
2
3
4
5
GSM8K
2.840
0.3
2.5
24.0
62.6
7.4
3.3
2.535
5.7
5.9
35.9
40.3
5.7
6.4
MATH500
2.859
0.8
4.1
24.9
56.5
6.0
7.7
2.970
12.8
2.9
14.9
36.3
10.0
23.1
AQuA
2.552
1.7
6.7
37.7
45.6
4.9
3.3
2.702
14.3
4.3
20.4
35.8
8.1
17.1
AMC23
2.868
0.3
2.3
22.7
64.5
5.3
4.9
2.830
19.6
2.2
11.5
33.3
8.9
24.4
OlymB
3.013
0.8
2.5
21.5
56.7
6.9
11.7
3.371
10.1
1.9
10.4
31.1
11.3
35.2
Appendix
Table 5: Mean and distribution of the criterion number actually put in force per audit. Both systems use the same criterion bank, mined from the Qwen3-8B system.
Qwen3-8B
Qwen3-14B
Benchmark
Mean
0
1
2
3
Mean
0
1
2
3
GSM8K
0.077
94.2
4.4
0.8
0.5
0.082
92.4
7.2
0.3
0.1
MATH500
0.186
87.4
8.3
2.4
1.8
0.067
94.4
4.8
0.6
0.2
AQuA
0.271
82.2
11.5
3.2
3.1
0.149
87.4
10.8
1.4
0.5
AMC23
0.325
79.2
12.9
4.2
3.8
0.125
90.0
7.9
1.7
0.4
OlymB
0.481
70.1
17.7
6.3
6.0
0.238
81.1
15.5
2.0
1.4
Appendix
Table 6: Average revision rounds actually consumed per message. Column 0 denotes the acceptance at the first audit; the last column merges messages accepted after three repairs with those ultimately withheld.
Method
Math Avg
Total tok.
Tok./inst.
Cost
Δ Acc.
( ×106 )
( ×103 )
vs. vanilla
Dynamic-MAS
50.86
76.43
25.1
1.00 ×
—
+ Multi-TAG
51.31
69.71
22.9
0.91 ×
+0.45
+ PRM
51.94
159.07
52.2
2.08 ×
+1.08
+ RaR-Gen
51.94
144.85
47.5
1.90 ×
+1.08
+ Self-Refine
48.89
211.07
69.2
2.76 ×
− 1.97
Appendix
Table 7: End-to-end token accounting on the Dynamic-MAS math suite. “Generic Criterion” replaces the retrieved rubric with the single domain-wide criterion of Appendix A.2 .
Figure 6: Average accuracy versus token consumption on Math benchmarks by the dynamic MAS framework.
Domain
Dataset
Size
Test Set
Math
GSM8K ( Cobbe et al., 2021 )
1,319
MATH-500 Lightman et al. (2024)
500
AQuA ( Ling et al., 2017 )
254
AMC23
40
OlympiadBench ( He et al., 2024 )
675
Appendix
Table 8: Dataset statistics
Figure 7: An example criterion from the constructed bank for the math domain. The Diagnostic Definition and Applicability Condition drive retrieval and selection, while the Applicability Condition and Inspection Directive are the fields handed to the auditing agent at judgment time.
Figure 8: The generic criterion designed for the math domain.
Figure 9: The prompt template for the auditing agent in the math domain. The Area of Concern block is instantiated per criterion: the applicability condition and the inspection directive are concatenated and substituted into the slot.
Figure 10: The prompt template for the auditing agent in the code domain.
Figure 11: The prompt template for the rubric miner during offline bank construction.