cs.LGOct 1, 2025

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Authors: Xin-Qiang Cai, Wei Wang, Feng Liu, Tongliang Liu, Gang Niu, Masashi Sugiyama

Organizations: RIKEN AIP, Tokyo, Japan · The University of Tokyo, Tokyo, Japan · The University of Melbourne, Melbourne, Australia

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) replaces costly human labeling with automated verifiers. To reduce verifier hacking, many RLVR systems binarize rewards to {0,1}\{0,1\}, but imperfect verifiers inevitably introduce \emph{false negatives} (rejecting correct answers) and \emph{false positives} (accepting incorrect ones). We formalize verifier unreliability as a stochastic reward channel with asymmetric noise rates ρ0ρ_0 and ρ1ρ_1 -- the FP rate and the FN rate, respectively. From this abstraction we derive two lightweight corrections: (i) a \emph{backward} correction that yields an unbiased surrogate reward and thus an unbiased policy-gradient estimator in expectation, and (ii) a \emph{forward} correction that reweights score-function terms so the expected update aligns with the clean gradient direction and requires only the FN rate. We implement both as lightweight hooks in a group relative policy optimization pipeline, both corrections improve RLVR for math reasoning under synthetic and real verifier noise, with the forward variant being more stable under heavier noise. Finally, an appeals mechanism with a lightweight LLM verifier estimates the FN rate online and further improves performance.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. An Imperfect Verifier is Good Enough: Learning with Noisy Rewards

    Apr 9, 2026Andreas Plesner, Francisco Guzmán, Anish AthalyeReinforcement Learning With Verifiable RewardCode Generation

  2. Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

    Sep 28, 2026Christian Moya, Elliott Thornley, Guang LinReinforcement Learning With Verifiable RewardVerifiable Rewards