cs.LGApr 9, 2026

An Imperfect Verifier is Good Enough: Learning with Noisy Rewards

Authors: Andreas Plesner, Francisco Guzmán, Anish Athalye

Organizations: Handshake AI · ETH Zurich

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) is widely used for post-training Large Language Models, but practical verifiers can make errors. We study how the rate and structure of reward noise affect RLVR in code generation, with a preliminary scientific-reasoning check. In multi-seed Qwen3 8B experiments on MBPP, mean validation reward over two fixed late evaluations is within 1 percentage point of the clean baseline at the tested resampled group-rollout noise rates through 20%, and within about 2 points at 30%. Confidence intervals allow larger losses; these point estimates do not establish a general tolerance threshold. We also examine full-program pass@k, four controlled noise structures, two model-based verifiers, and policy models from three families spanning 4B-9B parameters. We derive conditional advantage distributions for symmetric and asymmetric group noise, including retained format penalties, and show why clipping limits simple gradient-scaling arguments. The analysis identifies information preserved by whole-group corruption and limits on interpreting our asymmetric sweep as a precision-recall comparison. Overall, the results indicate that imperfect verification can support effective RLVR in the tested settings, while aggregate error rates alone do not characterize the learning signal.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

    Sep 28, 2026Christian Moya, Elliott Thornley, Guang LinReinforcement Learning With Verifiable RewardVerifiable Rewards

  2. Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

    May 27, 2026Saurabh Dash, Pierre Clavier, John Dang +4Reinforcement Learning With Verifiable RewardVerifiable Rewards