Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM systems evolve: tasks may leak into training data and thus be memorized by LLMs, traces and anomaly patterns expire as backbones evolve, and labels must be provided reliably for each refresh. To address these challenges, we present MAADBench (MA: multi-agent; AD: anomaly detection), the first refreshable MAS AD benchmark designed for diverse, evolving LLM backbones underlying the agents. MAADBench combines (1) sampled-and-coupled generative tasks over an approximately 10^37-task space to mitigate task leakage, (2) refreshable trace generation under configurable LLM backbones, and (3) automated provision of cost-free, deterministic step-level labels for fine-grained AD evaluation. Beyond offering the paradigm itself, we run MAADBench with five state-of-the-art LLM backbones and release the MAADBench-Full dataset with 5,200 step-labeled traces. Benchmarking 25 AD methods on the MAADBench dataset reveals substantial limitations in current approaches: they rely heavily on supervision, struggle with subtle MAS-specific anomalies, and lack robustness across LLM backbones. These gaps point to a rich research agenda for MAS-specific anomaly detection, with MAADBench providing a systematic and refreshable testbed for method development and evaluation. We open-source MAADBench-Full at https://huggingface.co/datasets/hww123/MAADBench-full.
Figures & tables
Figure 1: Overview of the MAS AD benchmarking pipeline (left) and the key design challenges (right) that must be addressed to achieve a long-lasting, scalable, and credible MAS AD benchmark.
Benchmark
CH1: Task Leakage
CH2: Trace
CH3: Labeling Fidelity
AD
Generative
Space †
Expiration
Step-level
Relabeling Cost \lx@sectionsign
Quality Guarantee
Support
MAST ( Cemri et al., 2025 )
× (fixed)
1,642
× (non-refresh.)
× (trace)
× (human/LLM)
Limited
TRAIL ( Deshpande et al., 2025 )
× (fixed)
148
× (non-refresh.)
× (trace)
× (human/LLM)
Limited
Silent Failures* ( Pathak et al., 2025 )
× (fixed)
5,169
× (non-refresh.)
× (trace)
× (human)
Limited
Who&When ( Zhang et al., 2025a )
× (fixed)
184
× (non-refresh.)
✓ (step)
× (human/LLM)
Limited
AgenTracer ( Zhang et al., 2026 )
× (fixed)
2,500+
× (non-refresh.)
✓ (step)
× (LLM-assisted)
Limited
Table 1: Comparison of SOTA MAS trace-analysis benchmarks across three core challenges. MaadBench addresses all three to support systematic MAS AD development and evaluation.
Figure 2: MAADBench Instantiation and MAS Execution. 1) Workflow: The execution workflow follows the same hierarchy: at the room level, the MAS resolves cross-puzzle dependencies; at the puzzle level, it executes atomic tasks with input-output dependencies. 2) Top-level Composite Task (Room): The MAS first needs to reason about the correct puzzle execution order under cross-puzzle dependencies, as a planning task. 3) Intermediate Composite Task (Puzzle): Each selected puzzle needs to be solved through a sequence of interdependent agent actions examining different capabilities. The Solve_Clue tasks are sampled from GSM-Hard (Math) and LiveCodeBench-Execution (Code), while the remaining actions are generated from parameterized templates with algorithmic ground truth. 4) Overall Task Difficulty: We define four room difficulty levels: Easy, Medium, Hard, and Extreme, by increasing the number of puzzles and the complexity of their cross-puzzle dependencies.
LLM
# Room ( × Temp)
Room-Level SR
Puzzle-Level
Easy
Medium
Hard
Extreme
Overall
Attempted / Total
SR
Claude Opus 4.8
400 ( × 1)
70.0%
63.0%
29.0%
30.0%
48.0%
805 / 1000
61.0%
Claude Sonnet 4.6
400 ( × 3)
52.3%
35.3%
11.0%
19.3%
29.5%
2065 / 3000
41.4%
DeepSeek-R1
400 ( × 3)
66.7%
42.0%
12.3%
22.0%
35.8%
2054 / 3000
43.9%
GPT-4.1
400 ( × 3)
46.3%
19.0%
3.0%
0.0%
17.1%
1521 / 3000
21.0%
GPT-5.4
400 ( × 3)
60.0%
27.0%
10.3%
0.3%
24.4%
1689 / 3000
30.0%
Table 2: Room-Level and Puzzle-Level MAS Success Rate ( ↑ ) Across LLM Types on MaadBench-Full . Best and worst values are highlighted. All rooms are fixed across all LLM instantiations to ensure fair comparison. Success Rate (SR) is verified by the final Room/Puzzle output.
LLM
SEL._PUZZLE
OBS._CLUE
SOL._CLUE
OBS._INSTRU
TOOL_CALL
APL._DELTA
OBS._PUZZLE
Overall
(Planning)
(Distraction)
(Math)
(Distraction)
(Tool Use)
(Reasoning)
(Verification)
Claude Opus 4.8
94.8%
99.8%
80.4%
97.8%
73.1%
75.8%
98.9%
89.1%
Claude Sonnet 4.6
97.1%
97.5%
80.5%
95.4%
64.9%
60.1%
100.0%
85.0%
DeepSeek-R1
100.0%
46.2%
63.0%
73.3%
60.8%
64.1%
97.0%
68.4%
GPT-4.1
70.9%
97.2%
71.3%
94.3%
63.2%
41.5%
98.4%
78.3%
GPT-5.4
65.6%
98.3%
78.7%
93.5%
67.0%
53.3%
99.9%
82.2%
Table 3: Success Rate ( ↑ ) of Actions on MaadBench-Full . For every action, we annotate the primary LLM capability it evaluates.
Figure 3: Anomaly rate by capability × LLMs.
LLM
Overall
Fmt Err
SEL._PUZZLE
OBS._CLUE
SOL._CLUE
OBS._INSTRU
TOOL_CALL
APL._DELTA
OBS._PUZZLE
Claude Opus 4.8
9.0%
9.1%
100.0%
0.0%
10.2%
16.7%
14.5%
0.0%
0.0%
Claude Sonnet 4.6
7.7%
18.8%
100.0%
0.0%
7.9%
42.1%
10.5%
0.0%
–
DeepSeek-R1
30.4%
34.1%
–
–
1.5%
50.0%
7.5%
0.0%
100.0%
GPT-4.1
12.2%
10.9%
100.0%
13.9%
13.5%
10.8%
25.5%
0.0%
80.0%
GPT-5.4
11.4%
18.5%
100.0%
10.3%
4.2%
20.0%
2.5%
0.0%
0.0%
Table 4: Silent failure rate by model and failure type on MaadBench-Full . Formatting errors ( Fmt Err ), e.g. , invalid JSON outputs, separated from action failures. Each column reports fraction of silent failures among all failures of corresponding type. “–” indicates no such failure occurred.
Figure 4: Action-level anomaly propagation by LLMs. Row : failed source action, Column : subsequent failed actions. Colored cells : propagation rates, as fraction of source failures followed by a downstream failure of target action. Darker colors indicate stronger propagation. “–” cells : no source failures. Cross-hatched cells : invalid action pairs.
Figure 5: Overall AUCROC of AD methods categorized by supervision level and method family. LLM zero-shot methods are grouped as unsupervised, while LLM few-shot methods are grouped as semi-supervised. Dashed line : A random baseline with AUCROC=0.5 as reference. Significance vs. Random : ** p<0.01 , * p<0.05 , ns p≥0.05 . More metrics with statistical tests are reported in Table 9 and 10 in Appendix H .
Figure 6: AUCROC of AD methods by actions (full results with std in Table 11 Appendix H ). Rows: Six action types with success rate; Bottom row: avg AUCROC per method across actions. Columns: Method AUC. Rightmost column: avg AUCROC per action across methods.
Figure 7: AD performance for different anomaly types. Evaluated Left : on test sets split by LLM backbone. Right : on test sets split by anomaly type (silent: puzzle solved; loud: puzzle failed).
Figure 8: Recall on anomaly propagation chains. In-propagation (solid) : recall on an action inside propagation chains; In-isolation (dashed) : global recall on the same actions as reference. The delta of the two is annotated. − means the anomalies are harder to detect inside a propagation chain, + otherwise.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Annotation method
Est. cost/trace
TRAIL ( Deshpande et al., 2025 )
Human, 5-person pipeline
\approx\38$
MAST ( Cemri et al., 2025 )
LLM-as-judge ( o1 ) + human QA
{>}\1$
Who&When ( Zhang et al., 2025a )
3 experts, 3-round consensus
\approx\9.2$
Silent Failures ∗ ( Pathak et al., 2025 )
Script + human GT per prompt †
small (unquantified)
AgenTracer ( Zhang et al., 2026 )
Counterfactual replay ‡
unreported
MaadBench (ours)
Deterministic, generated ground-truth
\mathbf{\approx\0}$
Appendix
Table 5: Estimated per-trace labeling cost across representative trace analysis benchmarks. H=\20$ /hr is used for human annotation estimates.
Difficulty
Puzzles
Fake Items/Clues
DAG Variants
Instance Space
Easy
1
1
—
[106,109]
Medium
2
3
—
[1011,1017]
Hard
3
5
—
[1017,1026]
Extreme
4
3
2,376
[1026,1037]
Appendix
Table 6: Task instance space per difficulty level. Each range reflects the least flexible instrument (Scale, 100×4 configurations) as the lower bound and the most flexible (Clock, 86,400×3 ) as the upper bound. Extreme additionally multiplies by 2,376 valid DAG configurations.
Model
Traces
API Calls
Input Tokens
Output Tokens
Price (in/out per 1M)
Total Cost
DeepSeek-R1
1,200
10,267
19.96M
13.05M
0.55/2.19
$39.57
Claude Sonnet 4.6
1,200
10,329
20.26M
5.72M
3.00/15.00
$146.63
Claude Opus 4.8
400
4,025
10.27M
1.68M
5.00/25.00
$93.44
GPT-4.1
1,200
7,605
13.27M
2.81M
2.00/8.00
$48.98
GPT-5.4
1,200
8,445
14.01M
2.04M
2.50/15.00
$65.71
Total
5,200
40,671
77.77M
25.31M
–
$394.32
Appendix
Table 7: API cost for generating MaadBench-Full across five SOTA LLM backbones, aggregated over all sampled configurations (domains, temperatures, and tool/no-tool variants). Prices are based on official API rates at experiment time. Claude Opus 4.8 currently covers only the default temperature configuration (400 traces); all other backbones cover 1,200 traces each.
Figure 9: Distribution of failures across five LLM backbones, broken down by MAST failure modes (FM-1.1–FM-3.3, light bars) and MaadBench ’s extended failure modes (FM-4.1, FM-4.2 and FM-5, dark bars). FM-4.2 (Reasoning Error) and FM-5 (Error Propagation) dominate failures across all backbones.
Action
Error code
Failure mode
Select_Puzzle
puzzle_locked
FM-1.1 Disobey task spec.
already_solved
FM-1.3 Step repetition
leads_to_deadend
FM-2.3 Task derailment
puzzle_not_found
FM-2.6 Reason–action mismatch
Observe_Clue
clue_select_wrong (GT id ∈ output reasoning content)
FM-2.6 Reason–action mismatch
clue_select_wrong (GT id ∈/ output reasoning content)
FM-4.1 Deception suscept.
Appendix
Table 8: Deterministic mapping from MaadBench action-level error codes to MAST failure modes. Cascade rule: if any upstream action on the same puzzle already failed, the downstream failure is labeled as FM-5 Error Propagation .
Figure 10: AUCROC vs. Normalized runtime GPU cost of semi-supervised methods.
Supervision
Method
Method Family
F1
Acc
AUC
Bal. Acc
-
Random
Random
.176±.021
.705±.021
.494±.013
.498±.013
Unsupervised
Isolation Forest
Tabular
.281±.041
.743±.012
.543±.039
.562±.020
KNN
Tabular
.239±.032
.728±.024
.488±.038
.537±.021
LOF
Tabular
.056±.009
.662±.027
.378±.026
.425±.011
Gemma4-E2B (zero-shot)
LLM
.296±.013
.746±.014
.560±.014
.571±.005
Gemma4-E4B (zero-shot)
LLM
.332±.027
.758±.018
.596±.020
.593±.017
Appendix
Table 9: Overall anomaly detection performance across four supervision settings. Dark blue cells indicate p<0.01 and light blue cells indicate p<0.05 relative to Random performance.
Supervision
Method
Method Family
F1
Acc
AUC
Bal. Acc
-
Random
Random
-
-
-
-
Unsupervised
Isolation Forest
Tabular
.001
.006
.022
<.001
KNN
Tabular
.004
.081
.631
.006
LOF
Tabular
1.000
.988
1.000
1.000
Gemma4-E2B (zero-shot)
LLM
<.001
.005
<.001
<.001
Gemma4-E4B (zero-shot)
LLM
<.001
.002
<.001
<.001
Appendix
Table 10: One-tailed two-sample t-test p-values (Method > Random) for overall anomaly detection performance across four performance metrics.
Family
Method
SP
OC
SC
OI
TC
AD
OP
Avg
(87.0%)
(96.5%)
(76.0%)
(91.9%)
(62.5%)
(57.3%)
(99.5%)
Unsupervised
iForest
0.49 ± 0.13
0.54 ± 0.02
0.60 ± 0.06
0.56 ± 0.04
0.51 ± 0.04
0.53 ± 0.07
0.47 ± 0.24
0.53 ± 0.05
kNN
0.37 ± 0.17
0.52 ± 0.07
0.68 ± 0.09
0.58 ± 0.04
0.54 ± 0.09
0.61 ± 0.03
0.52 ± 0.18
0.55 ± 0.10
LOF
0.71 ± 0.16
0.49 ± 0.05
0.62 ± 0.07
0.52 ± 0.05
0.57 ± 0.08
0.59 ± 0.03
0.56 ± 0.23
0.58 ± 0.07
Gemma4-E2B
0.24 ± 0.10
0.29 ± 0.08
0.62 ± 0.02
0.71 ± 0.05
0.40 ± 0.07
0.61 ± 0.02
0.41 ± 0.27
0.47 ± 0.18
Gemma4-E4B
0.19 ± 0.03
0.46 ± 0.06
0.63 ± 0.02
0.65 ± 0.03
0.53 ± 0.05
0.54 ± 0.04
0.72 ± 0.11
0.53 ± 0.17
Appendix
Table 11: AUCROC per (method × action), mean ± std over available train-test splits. Action success rate shown in parentheses. Per-column maxima in bold
Figure 11: Per-method recall on multi-anomaly propagation chains, broken down by step role. Rows : step role, including Origin (the first anomalous step), Mid (intermediate anomalous steps), and Down (the last anomalous step); Bottom row : per-method average across roles, indicating overall propagation detection performance. Columns : AD methods grouped by supervision level. Rightmost column : per-role average across methods. Middle steps have the highest average recall across methods ( 0.47 ), while origin and downstream steps are substantially harder ( 0.29 and 0.29 , respectively). Overall, Random Forest achieves the best average propagation-role recall ( 0.59 ), followed by XGBoost ( 0.57 ) and SVM ( 0.54 ). Per role, DOMINANT performs best on origin steps (recall 0.53 ), OC-SVM performs best on middle steps (recall 0.70 ), and DevNet performs best on downstream steps (recall 0.77 ).
Department of Computer Science, Worcester Polytechnic Institute, MA, USA · School of Computer Science, Fudan University, Shanghai, China · Institute of Big Data, Fudan University, Shanghai, China +1