An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
Organizations: Handshake AI · ETH Zurich
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used for post-training Large Language Models, but practical verifiers can make errors. We study how the rate and structure of reward noise affect RLVR in code generation, with a preliminary scientific-reasoning check. In multi-seed Qwen3 8B experiments on MBPP, mean validation reward over two fixed late evaluations is within 1 percentage point of the clean baseline at the tested resampled group-rollout noise rates through 20%, and within about 2 points at 30%. Confidence intervals allow larger losses; these point estimates do not establish a general tolerance threshold. We also examine full-program pass@k, four controlled noise structures, two model-based verifiers, and policy models from three families spanning 4B-9B parameters. We derive conditional advantage distributions for symmetric and asymmetric group noise, including retained format penalties, and show why clipping limits simple gradient-scaling arguments. The analysis identifies information preserved by whole-group corruption and limits on interpreting our asymmetric sweep as a precision-recall comparison. Overall, the results indicate that imperfect verification can support effective RLVR in the tested settings, while aggregate error rates alone do not characterize the learning signal.
Figures & tables
| Model | Setup | Seeds | Best | Final | Steps-to-best |
| Base model (before training) | |||||
| GLM4 9B | Base model | ||||
| Qwen3 8B | Base model | ||||
| Baselines (ground-truth unit tests) | |||||
| GLM4 9B | Baseline | ||||
| Qwen3 8B | Baseline | ||||
| Setup | Best |
| Base model | |
| No noise | |
| Noise | |
| Noise |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | Value |
| Optimization | |
| Optimizer | Adam |
| Learning rate | |
| LR schedule | Constant |
| Weight decay | |
| , | , |
| Model | Hyperparameters | Seeds | Best | Final |
| GLM4 9B | ||||
| top- , , (main) | ||||
| top- , , | ||||
| top- , , | ||||
| top- , , | ||||
| Qwen3 8B | ||||
| Noise rate | Peak score (mean SD) | (pp) | 95% CI (pp) | -value | |
| 0 | 4 | — | — | — | |
| 0.05 | 3 | -1.24 | 0.3601 | ||
| 0.10 | 6 | -0.63 | 0.5246 | ||
| 0.15 | 4 | -0.81 | 0.3085 | ||
| 0.20 | 4 | -0.42 | 0.5269 | ||
| 0.25 | 4 | -1.66 | 0.0284 |
| Noise rate | pass@1 | pass@4 | pass@8 | pass@16 | |
| 0 | 2 | ||||
| 0.10 | 2 | ||||
| 0.15 | 3 | ||||
| 0.20 | 2 | ||||
| 0.25 | 2 | ||||
| 0.30 | 2 |
| FNR FPR | 0.05 | 0.15 | 0.30 | 0.40 | 0.50 |
| 0.05 | 0.892 | 0.906 | 0.905 | 0.887 | 0.878 |
| 0.15 | 0.900 | 0.899 | 0.893 | 0.896 | 0.886 |
| 0.30 | 0.900 | 0.893 | 0.869 | 0.875 | 0.864 |
| 0.40 | 0.891 | 0.901 | 0.895 | 0.851 | 0.823 † |
| 0.50 | 0.896 | 0.860 | 0.888 | 0.843 | 0.781 ‡ |
| Noise rate | Best through 259 | At step 259 | Mean of 239, 259 | late (pp) [95% CI] | |
| 0 | 4 | — | |||
| 0.05 | 3 | ||||
| 0.10 | 6 | ||||
| 0.15 | 4 | ||||
| 0.20 | 4 | ||||
| 0.25 | 4 |
| Qwen2.5-Math-1.5B | Qwen2.5-Math-7B | DeepSeek-R1-Distill-Qwen-1.5B | ||||
| Reference | Range | Reference | Range | Reference | Range | |
| Math | ||||||
| MinervaMath | ||||||
| Olympiad Bench | ||||||
| AIME 2024 | ||||||
| AMC 2023 | ||||||