Cost-Effective Automated Judging of Natural-Language Mathematical Proofs
Organizations: Department of Computer Science, Dartmouth College, Hanover, NH, USA.
Abstract
Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sample of IMO-GradingBench, two of three cheap judges (GPT-OSS-120B, DeepSeek-V4-Flash) and their three-model consensus are statistically no worse than the frontier (Claude Opus 4.7, Gemini 3.1 Pro) on agreement with human pass/fail decisions, at 4-100 lower cost. On the full 1000-instance benchmark, the choice of consensus rule over the three judges is a precision/recall dial: unanimous (all-three-pass) rules reach the highest precision (0.855), majority vote the highest recall (0.912); across four replicate runs the unanimous rule is also the steadiest. No rule won outright; the dial replicated on a held-out 600-instance split and on the independent ProofBench. In this domain, cheap judges are competitive with the frontier at one to two orders of magnitude lower cost, and unanimity is the right setting when false positives are costly.
Figures & tables
| Judge | Reasoning | Pass-agree [95% CI] | 4-run | F1 | Spearman | Cost / 200 | |
|---|---|---|---|---|---|---|---|
| GPT-OSS-120B | xhigh | 0.875 [0.830, 0.920] | 0.851 | 0.716 | 0.806 | 0.623 | $0.32 |
| DeepSeek-V4-Flash | default | 0.860 [0.815, 0.905] | 0.881 | 0.664 | 0.763 | 0.660 | $0.70 |
| Cheap consensus (trio) | — | 0.855 [0.805, 0.905] | 0.874 | 0.673 | 0.779 | 0.654 | $1.73 |
| Claude Opus 4.7 | adaptive | 0.855 [0.805, 0.900] | — | 0.680 | 0.785 | 0.715 | $32.45 |
| Gemini 3.1 Pro | high | 0.840 [0.790, 0.890] | — | 0.645 | 0.761 | 0.704 | $28.61 |
| Gemini 3.1 Pro | default | 0.840 [0.790, 0.890] | — | 0.655 | 0.771 | 0.714 | $7.07 |
| System | Pass-agree [95% CI] | Prec. | Rec. | F1 | Spearman | |
|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 0.873 [0.852, 0.893] | 0.717 | 0.815 | 0.812 | 0.814 | 0.732 |
| GPT-OSS-120B ( xhigh ) | 0.842 [0.820, 0.864] | 0.668 | 0.716 | 0.889 | 0.793 | 0.704 |
| Gemma-4-31B ( high ) | 0.801 [0.776, 0.826] | 0.604 | 0.640 | 0.953 | 0.766 | 0.741 |
| Majority vote (trio) | 0.863 [0.842, 0.883] | 0.711 | 0.744 | 0.912 | 0.819 | 0.765 |
| All-three-pass | 0.879 [0.860, 0.898] | 0.725 | 0.855 | 0.777 | 0.814 | 0.733 |
| Fail-everything | 0.659 [0.629, 0.689] | 0.000 | — | 0.000 | 0.000 | — |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Reason. tok. | Agree | F1 | Cost | |
|---|---|---|---|---|---|
| Default | 0.840 | 0.771 | 0.714 | $7.07 | |
| High | 0.840 | 0.761 | 0.704 | $28.61 |
| System | orig | rep1 | rep2 | rep3 | mean | std | Self-maj. | Self-all-3 |
|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 0.860 | 0.861 | 0.898 | 0.905 | 0.881 | 0.024 | 0.901 | 0.895 |
| GPT-OSS-120B xhigh | 0.875 | 0.863 | 0.829 | 0.835 | 0.851 | 0.022 | 0.847 | 0.888 |
| Gemma-4-31B high | 0.795 | 0.820 | 0.810 | 0.814 | 0.810 | 0.011 | 0.824 | 0.849 |
| Majority vote (trio) | 0.855 | 0.875 | 0.890 | 0.875 | 0.874 | 0.014 | — | — |
| All-three-pass | 0.895 | 0.895 | 0.903 | 0.915 | 0.902 | 0.009 | — | — |
| System | Pass-agree | Prec. | Rec. | F1 | |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 0.878 | 0.737 | 0.843 | 0.824 | 0.833 |
| GPT-OSS-120B ( xhigh ) | 0.837 | 0.664 | 0.729 | 0.887 | 0.800 |
| Gemma-4-31B ( high ) | 0.807 | 0.622 | 0.662 | 0.973 | 0.788 |
| DeepSeek + GPT-OSS | 0.873 | 0.723 | 0.861 | 0.783 | 0.820 |
| DeepSeek + Gemma | 0.878 | 0.737 | 0.843 | 0.824 | 0.833 |
| GPT-OSS + Gemma | 0.868 | 0.724 | 0.786 | 0.882 | 0.832 |
| System | Pass-agree | Prec. | Rec. | F1 | |
|---|---|---|---|---|---|
| GPT-OSS-120B ( xhigh ) | 0.883 | 0.684 | 0.681 | 0.862 | 0.761 |
| DeepSeek-V4-Flash | 0.878 | 0.609 | 0.781 | 0.606 | 0.683 |
| Gemma-4-31B ( high ) | 0.903 | 0.730 | 0.741 | 0.851 | 0.792 |
| DeepSeek + GPT-OSS | 0.887 | 0.626 | 0.846 | 0.585 | 0.692 |
| DeepSeek + Gemma | 0.897 | 0.656 | 0.877 | 0.606 | 0.717 |
| GPT-OSS + Gemma | 0.917 | 0.754 | 0.815 | 0.798 | 0.806 |
| Disputed set | Sides with judges | Sides with human | |
|---|---|---|---|
| False positives (judges pass, human fails) | 25 | 10 (40%) | 15 (60%) |
| False negatives (judges fail, human passes) | 15 | 8 (53%) | 7 (47%) |
| Combined | 40 | 18 (45%) | 22 (55%) |
| Reference | System | Instance | Cluster | 4-run, cluster |
|---|---|---|---|---|
| Opus | GPT-OSS | ✓ | ✓ | ✓ |
| DeepSeek | ✓ | ✓ | ✓ | |
| Trio | ✓ | ✓ | ✓ | |
| Gemma | ||||
| Gem. high | GPT-OSS | ✓ | ✓ | ✓ |
| DeepSeek | ✓ | ✓ | ✓ |
| System | Pass-agree [95% CI] | Prec. | Rec. | F1 | Spearman | |
|---|---|---|---|---|---|---|
| DeepSeek + GPT-OSS | 0.878 [0.859, 0.897] | 0.723 | 0.852 | 0.777 | 0.813 | 0.733 |
| DeepSeek + Gemma | 0.877 [0.857, 0.897] | 0.725 | 0.824 | 0.812 | 0.818 | 0.738 |
| GPT-OSS + Gemma | 0.866 [0.845, 0.886] | 0.712 | 0.765 | 0.877 | 0.817 | 0.744 |