As AI agents surpass human performance, it becomes exceedingly hard for system designers to evaluate them directly and understand their failure modes. Consequently, agents themselves are being deployed extensively to evaluate, judge, and provide feedback on model traces. But this raises an important question: how can we trust the judge? In this work, we propose an agent-guided method to find weaknesses of agentic judges that expose interpretable failure mechanisms. Our method focuses on mathematical reasoning and proceeds in two stages: first, we deploy adversarial agents to mutate a set of sound proofs by introducing errors, attempting to misguide judges---in other words, injecting errors that judges are unable to catch. Then, we distill these attempts into a small set of mutation strategies which allow us to analyze the failure modes of the judges. To ensure that these strategies are not overfit to the initial set of proofs, we evaluate them by applying the mutation strategies to a held-out set of proofs and querying the same judge. We deploy our method on GPT-5.6-sol and Claude Opus 5, paired with their agent orchestrators, Codex and Claude Code, respectively. These are used both as mutators to introduce errors and as judges to evaluate correctness of mathematical reasoning. We find that across all the agentic judges, we are able to distill mutation strategies that consistently bypass their evaluations, thereby enabling us to ascertain actionable failure modes. Our analysis also reveals that judge reliability degrades at the frontier: errors in Olympiad-level proofs or graduate-level mathematical texts are detected more consistently, whereas flaws in research-level manuscripts are more likely to escape detection.
Figures & tables
Figure 2 : Overview of our mutation strategy learning and evaluation pipeline. (a) Long-running Mutator agents generate candidate errors; independent checking and judging identify successful bypasses, and both checks feed back to the mutator. (b) Distillation summarizes all attempt traces into a reusable strategy file. (c) On held-out proofs, strategy-guided and unguided mutators are evaluated through the same MutationChecker – Judge – JudgeChecker pipeline.
Figure 3 : Blind-judge miss rates during discovery and held-out evaluation. Error bars show descriptive 95% Wilson score intervals without adjustment for clustering. Both panels report the same pipeline instantiated with different models.
Figure 4 : (a) Cross-model judges: We take ten candidates with the most original judge misses from the strategy-guided evaluation run for each model on each dataset, and evaluate them with the other model used as a judge. Error bars show descriptive 95% Wilson score intervals without adjustment for clustering by source proof. (b) Higher-reasoning judges evaluated on six of the most successful mutations per dataset produced by their lower reasoning counterparts; each model judged its own mutations, so the two bars in a group cover different mutations. Two Claude Opus 5 (max) calls on Claude Code exited without producing a result after 128k output tokens.
Figure 4Figure 5
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Valid mutations
Completed reviews
Other-error flags
Mean/review
Mean if caught
Mean if missed
Olympiad
200
600
974
1.62
1.62
2.00
GraduateCourses
98
293
4,947
16.88
17.19
15.74
OpenAI-TCS
44
131
27
0.21
0.19
0.22
ArXivMath
50
150
907
6.05
5.50
7.14
Appendix
Table 1 : Other-error flags in the frozen GPT-5.6-sol evaluation. Ambiguous detection outcomes are included in the overall columns but excluded from the caught/missed breakdown.
GPT-5.6-sol judge / Claude mutations
Claude Opus 5 judge / GPT mutations
Dataset
Caught
Missed
Caught
Missed
Olympiad
7
3
8
2
GraduateCourses
5
5
9
1
OpenAI-TCS
3
7
5
5
ArXivMath
3
7
8
2
Total
18
22
30
10
Appendix
Table 2 : Cross-model evaluation of selected strategy-guided mutations. Each direction contains ten mutations per dataset.
Figure 5 : Mean blind-judge runtime by dataset and model, at medium reasoning effort. Error bars are percentile 95% bootstrap confidence intervals over calls.
Dataset
Judge
Calls
Mean (95% CI)
Median
Max
Excluded
Olympiad
GPT-5.6-sol
600
0.61 (0.59–0.64)
0.56
2.1
0
Claude Opus 5
597
0.37 (0.36–0.38)
0.36
0.8
0
GraduateCourses
GPT-5.6-sol
301
3.11 (3.03–3.19)
3.10
8.0
5
Claude Opus 5
293
4.73 (4.61–4.84)
4.67
7.7
8
OpenAI-TCS
GPT-5.6-sol
131
5.49 (5.20–5.81)
5.02
11.5
1
Claude Opus 5
116
13.42 (11.67–15.29)
10.71
43.4
22
Appendix
Table 3 : Blind-judge runtime in minutes per completed call. Excluded calls failed or did not finish.
Figure 6 : Blind-judge misses by mathematical area in the frozen evaluation. Bars show miss rates, and labels give misses over available review slots; the two mutation arms are pooled, invalid mutations are excluded, and ambiguous reviews do not count as misses. Error bars show descriptive 95% Wilson score intervals without adjustment for clustering of reviews within mutations or mutations within proofs.
Figure 7 : Mutation properties by outcome for valid GPT-5.6-sol candidates. Bars show means for successful and unsuccessful mutations; error bars are percentile 95% bootstrap confidence intervals. The intervals are descriptive and do not adjust for clustering by source proof or repeated mutation mechanism. Pooled mean edit sizes are 43.7 characters for successful mutations ( n=55 ) and 59.4 for unsuccessful mutations ( n=137 ).
Dataset
Generated
Valid
Distinct logical errors
Singleton mutations
Olympiad
200
200
132 (66.0%)
102 (51.0%)
GraduateCourses
100
98
70 (71.4%)
52 (53.1%)
OpenAI-TCS
50
44
36 (81.8%)
31 (70.5%)
ArXivMath
50
50
32 (64.0%)
21 (42.0%)
Total
400
392
270 (68.9%)
206 (52.6%)
Appendix
Table 4 : Semantic uniqueness of frozen GPT-5.6-sol evaluation mutations.
Dataset
Generated
Valid
Distinct logical errors
Singleton mutations
Olympiad
200
199
110 (55.3%)
79 (39.7%)
GraduateCourses
100
99
79 (79.8%)
63 (63.6%)
ArXivMath
50
50
36 (72.0%)
28 (56.0%)
Total
350
348
225 (64.7%)
170 (48.9%)
Appendix
Table 5 : Semantic uniqueness of frozen Claude Opus 5 evaluation mutations.