A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards
Organizations: ETH Zurich · Handshake AI
Abstract
When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifier agreement predicts post-training performance on such tasks. In this paper, we explore this question with over 11k H100 GPU-hours, across HealthBench and PRBench tasks in medical, legal, and finance domains. Across the tested domains, Qwen3 trainees (1.7B-8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes. We compare two low-cost choices retrospectively -- a cost-reducing choice and a balanced choice -- with estimated grading cost reductions of 98.8%-99.7% relative to the golden grading protocols and average post-training score gaps of 1-3 points from the best evaluated training verifier. These averages include larger losses in individual settings; they do not establish that verifier choices are interchangeable.
Figures & tables
| Domain | Train ( ) | Hard eval-only ( ) | Total eval ( ) |
| HealthBench | |||
| PRBench legal | |||
| PRBench finance |
| Verifier | Calls/task | Tokens/task | USD/1k tasks |
| Claude Opus 4.6 (golden) | 11.6 | 21,321 | 133.20 |
| GPT-5.5 (golden) | 11.6 | 17,093 | 99.70 |
| Gemma 4 31B | 4.3 | 5,498 | 0.77 |
| Gemma 4 26B | 4.3 | 5,514 | 0.45 |
| GPT oss 20B | 4.3 | 6,157 | 0.08 |
| Overall | HealthBench | PR-finance | PR-legal | ||||||
| Variant | Metric | all | hard | all | hard | all | hard | all | hard |
| Cheapest verifier (no quality floor) | Score gap | +0.030 | +0.026 | +0.062 | +0.054 | +0.011 | +0.011 | +0.017 | +0.013 |
| Uplift | 77% | 86% | 46% | 70% | 95% | 95% | 89% | 93% | |
| Cost reduction | 99.7% | 99.7% | 99.4% | 99.5% | 99.8% | 99.8% | 99.8% | 99.8% | |
| Cost-reducing Pareto verifier | Score gap | +0.030 | +0.023 | +0.062 | +0.046 | +0.011 | +0.011 | +0.017 | +0.013 |
| Uplift | 77% | 88% | 46% | 76% | 95% | 95% | 89% | 93% | |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
| Optimizer | Adam |
| Learning rate | (constant) |
| Weight decay | |
| Effective KL penalty | None |
| Algorithm | GRPO (group-std normalized advantages) |
| Rollouts per prompt |
| Trainee | GPUs/node | Tensor parallel | Approx. GPU-hours per cell |
| Qwen3 1.7B | 4 | 1 | 12 |
| Qwen3 4B | 4 | 2 | 18 |
| Qwen3 8B | 6 | 2 | 24 |
| Agreement | Cohen’s | |
| Pairwise | ||
| Claude Opus 4.6 vs. Gemini 3.1 Pro | ||
| Claude Opus 4.6 vs. GPT-5.5 | ||
| GPT-5.5 vs. Gemini 3.1 Pro | ||
| vs. -of- majority vote | ||
| Claude Opus 4.6 | ||
| HealthBench score | |||||||||
| Training verifier | Accuracy | Precision | Recall | F1 | Spearman | Intelligence | 8B | 4B | 1.7B |
| Gemma 4 31B | 0.85 | 0.80 | 0.87 | 0.83 | 0.78 | 1115 | 0.34 | 0.31 | 0.26 |
| GPT oss 120B | 0.82 | 0.80 | 0.78 | 0.79 | 0.72 | 832 | 0.31 | – | 0.20 |
| GPT oss 20B | 0.81 | 0.77 | 0.78 | 0.78 | 0.71 | 549 | 0.28 | 0.26 | 0.21 |
| Gemini 3 flash preview | 0.81 | 0.75 | 0.84 | 0.79 | 0.75 | 1205 | 0.32 | 0.32 | 0.24 |
| Gemma 4 26B | 0.79 | 0.74 | 0.78 | 0.76 | 0.68 | 1013 | 0.34 | 0.31 | 0.27 |
| HealthBench score | ||||||||
| Training verifier | Accuracy | Precision | Recall | F1 | Spearman | 8B | 4B | 1.7B |
| Gemma 4 31B | 0.86 | 0.77 | 0.90 | 0.83 | 0.83 | 0.42 | 0.41 | 0.33 |
| GPT oss 120B | 0.83 | 0.75 | 0.83 | 0.79 | 0.73 | 0.43 | – | 0.28 |
| GPT oss 20B | 0.84 | 0.74 | 0.90 | 0.81 | 0.87 | 0.37 | 0.36 | 0.26 |
| Grok 4.20 | 0.78 | 0.72 | 0.70 | 0.71 | 0.76 | 0.46 | – | – |
| Gemini 3 flash preview | 0.82 | 0.71 | 0.89 | 0.79 | 0.82 | 0.44 | 0.41 | 0.32 |
| Training verifier | Accuracy | Precision | Recall | F1 | Spearman | PRBench finance score |
| Gemma 4 31B | 0.91 | 0.82 | 0.78 | 0.80 | 0.79 | 0.46 |
| GPT oss 120B | 0.90 | 0.82 | 0.71 | 0.76 | 0.82 | 0.38 |
| Gemma 4 26B | 0.89 | 0.77 | 0.73 | 0.75 | 0.71 | 0.45 |
| Gemini 3 flash preview | 0.88 | 0.72 | 0.82 | 0.77 | 0.79 | 0.45 |
| GPT oss 20B | 0.86 | 0.71 | 0.65 | 0.68 | 0.70 | 0.45 |
| Qwen3 235B | 0.83 | 0.59 | 0.80 | 0.68 | 0.63 | 0.45 |
| Training verifier | Accuracy | Precision | Recall | F1 | Spearman | PRBench finance score |
| Gemma 4 31B | 0.93 | 0.83 | 0.85 | 0.84 | 0.89 | 0.42 |
| GPT oss 120B | 0.91 | 0.78 | 0.71 | 0.74 | 0.88 | 0.35 |
| Gemma 4 26B | 0.91 | 0.77 | 0.78 | 0.78 | 0.70 | 0.42 |
| GPT oss 20B | 0.89 | 0.74 | 0.66 | 0.70 | 0.70 | 0.42 |
| Gemini 3 flash preview | 0.89 | 0.68 | 0.80 | 0.74 | 0.85 | 0.42 |
| Qwen3 Next 80B | 0.85 | 0.59 | 0.85 | 0.70 | 0.78 | 0.42 |
| Training verifier | Accuracy | Precision | Recall | F1 | Spearman | PRBench legal score |
| Gemma 4 26B | 0.90 | 0.82 | 0.81 | 0.81 | 0.81 | 0.41 |
| Gemma 4 31B | 0.90 | 0.78 | 0.90 | 0.83 | 0.85 | 0.42 |
| GPT oss 120B | 0.89 | 0.77 | 0.76 | 0.76 | 0.79 | 0.37 |
| GPT oss 20B | 0.85 | 0.76 | 0.70 | 0.73 | 0.76 | 0.41 |
| Gemini 3 flash preview | 0.87 | 0.72 | 0.89 | 0.79 | 0.83 | 0.43 |
| Qwen3 235B | 0.84 | 0.66 | 0.88 | 0.75 | 0.80 | 0.40 |
| Training verifier | Accuracy | Precision | Recall | F1 | Spearman | PRBench legal score |
| GPT oss 120B | 0.93 | 0.83 | 0.69 | 0.75 | 0.60 | 0.35 |
| Gemma 4 26B | 0.91 | 0.83 | 0.65 | 0.73 | 0.40 | 0.39 |
| GPT oss 20B | 0.91 | 0.81 | 0.66 | 0.73 | 0.49 | 0.39 |
| Gemma 4 31B | 0.94 | 0.81 | 0.87 | 0.84 | 0.76 | 0.39 |
| Gemini 3 flash preview | 0.91 | 0.73 | 0.84 | 0.78 | 0.69 | 0.40 |
| Qwen3 235B | 0.88 | 0.64 | 0.82 | 0.72 | 0.64 | 0.37 |
| Split | Median | Max | |||
| non-hard | |||||
| non-hard-eval | |||||
| hard |
| Split | Median | Max | |||
| non-hard | |||||
| non-hard-eval | |||||
| hard |
| Claude | Gemini | GPT | |||||
| Variant | Metric | all | hard | all | hard | all | hard |
| Cheapest verifier (no quality floor) | Score gap | +0.050 | +0.039 | +0.021 | +0.027 | +0.019 | +0.012 |
| Uplift | 65% | 82% | 82% | 85% | 84% | 92% | |
| Cost reduction | 99.8% | 99.9% | 99.6% | 99.7% | 99.7% | 99.6% | |
| Cost-reducing Pareto verifier | Score gap | +0.050 | +0.039 | +0.021 | +0.027 | +0.019 | +0.004 |
| Uplift | 65% | 82% | 82% | 85% | 84% | 97% | |
| Verifier | Calls | Input tok. | Output tok. | USD |
| Claude Opus 4.6 (golden) | ||||
| GPT-5.5 (golden) | ||||
| Gemma 4 31B | ||||
| Gemma 4 26B (balanced) | ||||
| GPT oss 20B (cost-reducing) |
| Strategy | Family | Score gap | Uplift | Cost reduction | Pareto |
| Always Gemma 4 26B (balanced Pareto, paper pick) | quality-first | 95% | 98.8% | 13/18 | |
| Always Gemma 4 31B | quality-first | 91% | 98.0% | 3/18 | |
| Pareto frontier argmax | balanced | 88% | 98.1% | 4/18 | |
| Cheapest in top-3 by top-3 by | balanced | 86% | 98.5% | 8/18 | |
| Argmax | balanced | 86% | 98.5% | 7/18 | |
| Argmax Spearman | quality-first | 85% | 97.6% | 6/18 |