Reinforcement Learning with Verifiable Rewards (RLVR) replaces costly human labeling with automated verifiers. To reduce verifier hacking, many RLVR systems binarize rewards to {0,1}, but imperfect verifiers inevitably introduce \emph{false negatives} (rejecting correct answers) and \emph{false positives} (accepting incorrect ones). We formalize verifier unreliability as a stochastic reward channel with asymmetric noise rates ρ0 and ρ1 -- the FP rate and the FN rate, respectively. From this abstraction we derive two lightweight corrections: (i) a \emph{backward} correction that yields an unbiased surrogate reward and thus an unbiased policy-gradient estimator in expectation, and (ii) a \emph{forward} correction that reweights score-function terms so the expected update aligns with the clean gradient direction and requires only the FN rate. We implement both as lightweight hooks in a group relative policy optimization pipeline, both corrections improve RLVR for math reasoning under synthetic and real verifier noise, with the forward variant being more stable under heavier noise. Finally, an appeals mechanism with a lightweight LLM verifier estimates the FN rate online and further improves performance.
Figures & tables
Figure 1: Verifier-noise flow in RLVR. An AI agent produces candidate solutions that are scored by automated verifiers. While verifiers would yield false negatives ( 3612 vs. 31 , reaching 38% rates ( Xu et al., 2025 ) ) and false positives (mislead by “Let’s solve it step by step…”, reaching 35%−66.8% rates ( Zhao et al., 2025 ) ), confusing the agent; applying our backward/forward corrections restores correct signals.
Ri←1−ρ^0−ρ^1R~i−ρ^0.
Algorithm 1 Noisy Policy Gradient with Backward Correction (PGBC)
wR~i←{ρ^1−1,ρ^1,if R~i=0,if R~i=1.
Algorithm 2 Noisy Policy Gradient with Forward Correction (PGFC)
Figure 2: Synthetic-Noise Results (pass@1) with 16 samples and 5 random seeds. Base : baseline without RL; Oracle : Training with clean rewards; Noise : Training with noisy verifier rewards; Noise_BC : Training with noise under backward correction; Noise_FC : Training with noise under forward correction.
Figure 3: Synthetic-Noise Results (pass@8) with 16 samples and 5 random seeds. Base : baseline without RL; Oracle : Training with clean rewards; Noise : Training with noisy verifier rewards; Noise_BC : Training with noise under backward correction; Noise_FC : Training with noise under forward correction.
Dataset
AIME2024
AIME2025
AMC2023
MATH500
Minerva MATH
Olympiad Bench
Average
Qwen2.5-Math-1.5B
Base
6.0 ± 1.9
4.0 ± 0.6
34.2 ± 0.2
47.5 ± 0.3
5.1 ± 0.4
25.1 ± 0.4
20.3 ± 0.6
Rule
15.0 ± 0.4
5.6 ± 0.6
50.3 ± 0.6
69.4 ± 0.4
17.8 ± 0.6
31.6 ± 0.0
31.6 ± 0.4
LLM-as-Judge
10.9 ± 1.3
4.7 ± 1.0
42.1 ± 1.8
63.0 ± 0.7
15.9 ± 0.7
25.3 ± 0.5
27.0 ± 1.0
Appeals
11.9 ± 0.6
5.8 ± 1.2
47.8 ± 1.2
68.3 ± 0.1
16.7 ± 0.6
29.8 ± 0.1
30.1 ± 0.6
Appeals+PGFC (Ours)
20.3 ± 0.0
10.7 ± 1.7
53.3 ± 1.4
68.6 ± 0.8
16.5 ± 0.4
32.9 ± 0.2
33.7 ± 0.8
Table 1: Real-world pass@1. Rule : uncorrected rule-based rewards; LLM-as-Judge : direct LLM-judge rewards; Appeals : rule-based reward plus LLM appeals on negative samples without gradient correction; Appeals+PGFC : appeals-based FN-rate estimation plus forward correction.
Figure 4: Robustness results. (a) Backward correction (BC) with ρ^0 fixed and sweeping ρ^1 ; (b) Backward correction (BC) with ρ^1 fixed and sweeping ρ^0 ; (c) Forward correction (FC) with ρ^0 fixed and sweeping ρ^1 .
Method
MATH500
AIME2024
AIME2025
AMC2023
Minerva Math
OlympiadBench
Average
Base
47.5 / 62.8
6.0 / 32.4
4.0 / 16.7
34.2 / 79.3
5.1 / 15.5
25.1 / 31.3
20.3 / 39.7
Noise
40.7 / 60.5
7.5 / 26.5
2.2 / 15.8
30.0 / 72.4
3.8 / 11.6
21.5 / 25.8
17.6 / 35.4
PGBC
67.7 / 67.9
11.5 / 33.0
7.4 / 22.4
49.1 / 81.8
20.7 / 21.1
30.5 / 31.7
31.2 / 43.0
PGFC
68.1 / 69.9
12.8 / 32.4
8.0 / 22.6
49.6 / 79.5
21.8 / 23.8
31.1 / 33.6
31.9 / 43.6
Table 2: Format-dependent non-iid verifier-noise setting on Qwen2.5-Math-1.5B. Each cell reports pass@1 / pass@8.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Training dynamics under synthetic iid noise. PGBC and PGFC change the group reward standard deviation, normalized advantage standard deviation, and gradient norm relative to uncorrected noisy training, illustrating that the correction is active inside advantage construction rather than only changing endpoint scores.
Dataset
AIME2024
AIME2025
AMC2023
MATH500
Minerva MATH
Olympiad Bench
Average
Qwen2.5-Math-1.5B
Base
32.4 ± 0.5
16.7 ± 0.7
79.3 ± 1.4
62.8 ± 1.3
15.5 ± 1.1
31.3 ± 0.8
39.7 ± 1.0
Rule
34.3 ± 1.2
19.9 ± 1.1
80.7 ± 1.4
69.6 ± 0.5
18.0 ± 0.6
32.8 ± 1.2
42.6 ± 1.0
LLM-as-Judge
29.6 ± 1.1
15.4 ± 0.4
80.0 ± 0.6
63.4 ± 1.2
16.2 ± 0.8
28.6 ± 0.3
38.9 ± 0.7
Appeals
30.5 ± 0.2
20.5 ± 0.7
80.7 ± 0.9
68.9 ± 1.3
17.6 ± 1.0
30.5 ± 0.5
41.5 ± 0.8
Appeals+PGFC (Ours)
31.0 ± 0.2
20.0 ± 0.6
82.2 ± 0.4
69.8 ± 1.1
18.2 ± 0.5
33.3 ± 0.3
42.4 ± 0.5
Appendix
Table 3: Pass@8 under real-world noise.
Method
Avg pass@1
Avg pass@8
Appeal rate
Flip rate
TinyV calls
Rule
31.6±0.4
42.6±1.0
0.000
0.0000
0
Appeals
30.1±0.6
41.5±0.8
0.081
0.2288
7,963
Appeals+PGFC
33.7±0.8
42.4±0.5
0.093
0.2745
8,255
Appendix
Table 4: Appeals statistics for Qwen2.5-Math-1.5B real-world verifier noise. Appeals+PGFC uses the same appeal stream both for recovery and for online FN-rate estimation.
Method
MATH500
AIME2024
AIME2025
AMC2023
Minerva Math
OlympiadBench
Average
Stronger checker
68.1 / 69.2
12.3 / 18.4
7.2 / 12.3
47.9 / 59.7
17.8 / 18.6
30.9 / 33.1
30.7 / 35.2
Appeals+PGFC
72.0 / 75.1
14.2 / 21.6
8.8 / 15.8
51.8 / 65.8
20.9 / 23.8
33.3 / 37.2
33.5 / 39.9
Appendix
Table 5: Complementarity with a stronger verifier-side baseline on Qwen2.5-Math-1.5B. The correction still improves over a stronger checker, indicating that PGFC is not merely replacing verifier engineering.
Method
MMLU
ARC-Challenge
HellaSwag
GPQA
Base
60.7
73.9
55.0
28.4
Oracle
61.1
73.9
55.9
29.1
Noise
50.0
43.9
33.9
19.4
PGBC
60.8
73.8
55.6
28.8
PGFC
61.3
74.2
56.1
29.2
Appendix
Table 6: General capability retention after RLVR training on Llama-3.2-3B-Instruct . Metrics are accuracies.
Training (GRPO)
Train batch size
128
Rollouts per question (group size)
8
Max prompt length (tokens)
512
Max response length (tokens)
3072
Sampling temperature (rollouts)
1.0
Advantage estimator
GRPO
Appendix
Table 7: Core training settings.
Method
AIME 2024
AIME 2025
AMC 2023
MATH-500
Minerva MATH
OlympiadBench
Rule
15.0±0.4
5.6±0.6
50.3±0.6
69.4±0.4
17.8±0.6
31.6±0.0
LLM-Verifier
16.4±1.0
7.2±0.9
51.5±1.0
70.8±0.7
18.5±0.8
32.0±0.6
Compass Appeals
16.7±0.8
7.4±0.8
51.8±0.8
71.0±0.6
18.9±0.7
32.3±0.5
Compass Appeals+PGBC
18.7±0.7
8.9±0.8
53.0±0.7
71.7±0.5
21.5±0.6
33.1±0.5
Compass Appeals+PGFC
19.8±0.4
9.8±0.1
54.0±0.3
72.1±0.7
20.1±0.3
33.8±0.9
Appendix
Table 8: Qwen2.5-Math-1.5B with CompassVerifier-3B and estimated channel rates. Scores are percentages.
ρ0
ρ1
c
Noise
PGBC
PGFC
Oracle
0.05
0.10
0.85
27.0±0.8
31.5±0.6
31.1±0.3
31.8±0.5
0.10
0.20
0.70
20.4±0.8
31.1±0.6
31.6±0.6
31.8±0.5
0.15
0.30
0.55
17.2±1.0
29.4±0.8
30.7±0.7
31.8±0.5
0.20
0.40
0.40
14.5±1.2
26.4±1.0
30.5±0.9
31.8±0.5
Appendix
Table 9: True-channel sweep: average accuracy (%) across six math benchmarks.
Statistic
Value
Sensitivity on true false negatives
0.93
False-positive rate on true negatives
0.03
Naive ρ1
0.785
Calibration-corrected ρ1
0.780
Appendix
Table 10: Reported deployment-time calibration on appealed rule-negative responses.
Method
Var(a∣R∗=0)
Var(a∣R∗=1)
trCov(g)
Gradient-norm CV
Noise
0.107
0.482
1.000×
0.729
PGBC
0.671
1.194
3.046×
0.873
PGFC
0.181
0.382
0.905×
0.730
Appendix
Table 11: Direct variance measurements. Gradient covariance trace is normalized to the Noise value.
Method
AIME 2024
AIME 2025
AMC 2023
MATH-500
Minerva MATH
OlympiadBench
Oracle
10.8±0.2
7.6±0.1
50.9±0.4
69.3±0.1
29.2±0.5
33.4±0.3
Noise
7.1±0.7
4.2±0.6
37.8±0.8
54.6±0.5
15.8±1.0
25.9±0.6
PGBC
10.2±0.5
9.1±0.4
47.8±0.7
67.9±0.4
30.6±0.8
31.4±0.5
PGFC
11.1±0.4
9.4±0.3
48.4±0.5
69.2±0.3
30.0±0.7
32.2±0.4
Appendix
Table 12: Additional direct-coefficient REINFORCE results on Qwen2.5-Math-1.5B.
Rollouts K
Oracle
Noise
PGBC
PGFC
8
33.5±0.3
24.2±0.7
32.8±0.6
33.4±0.5
16
34.0±0.5
26.0±0.8
33.2±0.4
33.8±0.2
32
35.2±0.3
28.7±0.4
34.4±0.5
35.0±0.4
Appendix
Table 13: Additional rollout-count comparison on Qwen2.5-Math-1.5B.
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used for post-training Large Language Models, but practical verifiers can make errors. We study how the rate and structure of reward noise affect RLVR in code generation, with a preliminary scientific-reasoning check. In multi-seed Qwen3 8B experiments on MBPP, mean validation reward over two fixed late evaluations is within 1 percentage point of the clean baseline at the tested resampled group-rollout noise rates through 20%, and within about 2 points at 30%. Confidence intervals allow larger losses; these point estimates do not establish a general tolerance threshold. We also examine full-program pass@k, four controlled noise structures, two model-based verifiers, and policy models from three families spanning 4B-9B parameters. We derive conditional advantage distributions for symmetric and asymmetric group noise, including retained format penalties, and show why clipping limits simple gradient-scaling arguments. The analysis identifies information preserved by whole-group corruption and limits on interpreting our asymmetric sweep as a precision-recall comparison. Overall, the results indicate that imperfect verification can support effective RLVR in the tested settings, while aggregate error rates alone do not characterize the learning signal.
In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.
Christian Moya, Elliott Thornley, Guang Lin
Purdue University · National University of Singapore
Reinforcement learning with verifiable rewards (RLVR) replaces human preference labels with executable reward functions such as math answer checkers, JSON tool-call validators, and code unit-test harnesses. That makes the reward partly a software artifact: if the verifier is wrong, optimization can learn the bug. We study this failure mode with a lightweight verifier-fuzzing framework that generates adversarial completions, compares buggy and stricter reference verifiers, logs paired decisions, and reports false-positive, false-negative, disagreement, exploit, and uncertainty metrics.