Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be (ε,δ)-differentially private with respect to the presence of any one training problem. Taking the group of responses to one prompt as the privacy record, our method aggregates their gradients, clips the prompt's contribution once, adds Gaussian noise, and composes the privacy loss across updates, so the budget depends on neither the number of responses per prompt nor the clipping norm; to our knowledge this is the first differential privacy guarantee for RLVR training. We train Qwen2.5-1.5B-Instruct with LoRA at a per-run budget of ε=8 and compare, on the same prompts and at the same budget, a control that removes only the reward signal and two private supervised fine-tuning recipes. The reward signal improves accuracy over the control by 2.65 points on MATH and 3.24 on GSM8K, in every seed; the improvement survives a format-robust scorer, at 1.3 points on MATH, and is not explained by response length. At the same budget the private model outperforms both supervised recipes on MATH and GSM8K by 2.3 to 3.8 points, retains 85--90% of the gain of non-private GRPO on these tasks, and on MATH the noise of an eightfold tighter budget costs at most 1.2 points. The reward effect also carries to CommonsenseQA, an exploratory non-mathematical task. Verifier feedback thus remains a usable learning signal under prompt-level privacy.
Figures & tables
Fig. 1: One round of DP-GRPO (Algorithm 1 ). Each Poisson-sampled prompt xi (line 2) generates G completions whose verifier rewards are normalized within that group only (lines 4–5). The rollout losses are combined into one prompt gradient gi and clipped to norm C (lines 6–7), so the dashed region is a single privacy record. Clipped prompt gradients are summed over the sample, perturbed with Gaussian noise, divided by the public constant qn , and applied to the LoRA parameters (lines 8–9). The next round samples rollouts from the updated policy.
MATH
GSM8K
CSQA
Role
development
confirmatory
exploratory
Training prompts n
6,988
6,961
6,961
Evaluation set
4,499 test
1,319 test
1,221 val.
Rounds T
50
100
100
Expected passes qT
0.92
1.84
1.84
Noise multiplier σ
0.484
0.520
0.520
TABLE I: Per-benchmark settings of the private runs. Shared by all three: ϵ=8 at δ=1/(10n) , qn=128 , G=8 , C=0.055 (DP-GRPO, control), learning rate 3×10−4 . Full-precision values, the supervised arms’ recipe parameters and split construction are in Appendix C ; certified budgets in Table VI .
Fig. 2: Strict accuracy of each arm at ϵ=8 . Dot: mean over seeds; ticks: individual seeds; dashed line: public base. Each bar is the 95% paired interval of the arm’s difference from the base, added to the base value, so a bar that clears the dashed line is a resolved gain and one that crosses it is not. Exact values in Appendix D-A .
Comparator
MATH (pp)
GSM8K (pp)
Zero-adv. control
+2.65 [ 1.79 , 3.51 ]
+3.24 [ 1.61 , 4.89 ]
Filtered DP-SFT
+2.32 [ 1.39 , 3.26 ]
+3.53 [ 1.74 , 5.31 ]
Gold DP-SFT
+2.99 [ 2.01 , 3.98 ]
+3.79 [ 1.97 , 5.57 ]
TABLE II: DP-GRPO accuracy gain over each comparator.
Model
Strict (%)
Format-robust (%)
Public base
49.06
50.37
Zero-adv. control
48.92(−0.14)
50.28(−0.08)
Filtered DP-SFT
49.28(+0.22)
50.47(+0.10)
Gold DP-SFT
48.61(−0.44)
49.98(−0.39)
DP-GRPO
51.57(+2.51)
51.62(+1.26)
Non-private GRPO †
52.12(+2.96)
–
TABLE III: MATH accuracy under two answer extractors.
Model
Accuracy (%)
Public base
70.74
Zero-adv. control
71.10(+0.36)
Filtered DP-SFT
70.51(−0.23)
Gold DP-SFT
70.24(−0.49)
DP-GRPO
74.34(+3.60)
Non-private GRPO †
75.36(+4.02)
TABLE IV: GSM8K accuracy.
Model
Strict (%)
Final answer (%)
Public base
46.19
66.75
Zero-adv. control
45.99(−0.20)
67.20(+0.45)
Filtered DP-SFT
52.42(+6.22)
66.71(−0.04)
Gold DP-SFT
71.70(+25.51)
71.70(+4.95)
DP-GRPO
70.84(+24.65)
71.46(+4.71)
Non-private GRPO
77.27(+31.08)
77.27(+10.52)
TABLE V: CommonsenseQA accuracy under two answer extractors.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
n
(m,σ,T)
PLD
RDP
MATH, ϵ=1 arm
6,988
(128, 1.002, 50)
1.000
1.463 (7.6)
MATH, ϵ=4.23 arm
6,988
(128, 0.612, 50)
4.235
5.322 (3.2)
MATH, reference †
6,988
(128, 0.484, 50)
8.000
9.830 (2.3)
MATH, DP-SFT T=100
6,988
(128, 0.519, 100)
8.000
9.702 (2.3)
MATH, coverage
6,988
(512, 0.702, 50)
8.000
9.323 (2.6)
MATH, hist. rank-4
6,988
(128, 0.612, 100)
5.038
6.185 (3.0)
Appendix
TABLE VI: Every accounting configuration in this paper, with δ=1/(10n) and q=m/n , under PLD (reported) and RDP accounting. Parenthesized: the minimizing RDP order.
Configuration
n
m
T
σ
MATH, reference †
6,988
128
50
0.4842404738068581
MATH, ϵ=4.23 (bf16)
6,988
128
50
0.612058477104
MATH, ϵ=1 (bf16)
6,988
128
50
1.0023417361080647
MATH, coverage ( G=2 )
6,988
512
50
0.7022244453430175
MATH, DP-SFT, 100 updates
6,988
128
100
0.519267738191411
MATH, canary (bf16)
6,988
128
100
0.612058477104
Appendix
TABLE VII: Mechanism parameters at full precision. m is the expected number of prompts per round.
Decision
Fixed when
Strict scorers (all benchmarks)
Before any fixed-protocol arm was run (Sec. VII-D ).
DP-GRPO LR, C , G , T on MATH
In the MATH design study, before the fixed protocol (Sec. VII-B ).
Supervised learning rates
Four-point pilots on MATH-500, before the fixed protocol (App. C-D ).
GSM8K horizon and cap
Before any private GSM8K run, from one earlier non-private run (Sec. VII-A ).
MATH-500 excluded from evaluation
After the earlier runs that used it for selection; applied by problem identity to saved predictions (App. C-C ).
Seeds 3–4, DP-GRPO and control, MATH and GSM8K
After the seed 1–2 results were inspected (Sec. VII-C ); four-seed contrasts are sequential.
Appendix
TABLE VIII: Protocol decisions and their timing.
Arm
Benchmark
Seeds
Reported
Stage 1: design study
ϵ∈{1,4.23,8} (bf16)
MATH
2 each
Tab. XII
C∈{21,1,2}×0.055
MATH
2 each
Tab. XII
Rollout allocation (2 alt.)
MATH
2 each
Tab. XII
Same-LR and 100-update SFT
MATH
2 each
Sec. VIII-B
σ=0 , clipped / unclipped
GSM8K
2 each
Fig. 3
Appendix
TABLE IX: Every arm, the benchmarks it ran on, its seeds, and where its results are reported. Seeds for Stage 2 are given as MATH / GSM8K / CSQA.
Arm
Seeds
Mean
MATH, strict
Public base
–
49.1
Zero-adv. control
48.9 / 48.8 / 49.1 / 48.8
48.9
DP-GRPO
51.3 / 51.9 / 51.8 / 51.3
51.6
Filtered DP-SFT
49.3 / 49.3
49.3
Gold DP-SFT
48.7 / 48.6
48.6
Appendix
TABLE X: Strict accuracy (%) of every seed plotted in Fig. 2 , and the MATH format-robust accuracies behind Table III . Seeds in run order.
Contrast
Per-seed difference
Holm p , largest
MATH
DP-GRPO − control
+2.31 / +3.16 / +2.62 / +2.51
6.4×10−5
DP-GRPO − filtered
+1.98 / +2.67
2.4×10−4
DP-GRPO − gold
+2.60 / +3.38
2.9×10−6
GSM8K
DP-GRPO − control
+3.26 / +2.81 / +2.65 / +4.25
0.013
Appendix
TABLE XI: Per-seed differences (points) behind Table II , with Holm-adjusted exact McNemar p (Holm within each contrast). The supervised contrasts pair DP-GRPO seeds 1–2 with the supervised seeds.
Fig. 3: Non-private GSM8K references at DP-GRPO’s hyperparameters: noise removed, with clipping retained or disabled. Dots show two-seed means; bars span the seeds, not confidence intervals. DP-GRPO is the mean of its two matched seeds at the final step (74.0%). The earlier non-private run of Table IV (75.4%) was trained at a different rank, learning rate and horizon and is not shown.
Alternative
Change (pp), 95% CI
MDE (pp)
Budget; reference: ϵ=8 (earlier study)
ϵ=4.23
−0.09 [ −0.66 , 0.48 ]
0.82
ϵ=1
−0.57 [ −1.19 , 0.06 ]
0.89
Clipping; reference: C=0.055
C=0.111
−0.03 [ −0.37 , 0.30 ]
0.48
C=0.028
−0.30 [ −0.72 , 0.12 ]
0.60
Appendix
TABLE XII: Supporting MATH design comparisons.
Arm
Seeds
Mean
DP-GRPO − arm
Public base
–
82.0
+1.20 [ −0.53 , 2.95 ]
Zero-adv. control
82.1 / 81.8 / 82.1 / 82.1
82.0
+1.18 [ −0.35 , 2.73 ]
Filtered DP-SFT
82.6 / 82.6
82.6
+0.25 [ −1.45 , 1.95 ]
Gold DP-SFT
82.2 / 82.7
82.5
+0.40 [ −1.40 , 2.20 ]
DP-GRPO
83.0 / 82.7 / 83.4 / 83.7
83.2
–
Appendix
TABLE XIII: SVAMP (1,000 problems): GSM8K checkpoints scored without further training. Strict accuracy per seed, and DP-GRPO’s difference from each arm (points, 95% CI).
Base response
Questions
Contribution (pp)
Strict
Format-robust
Boxed an answer
3,831
−0.24
−0.24
No boxed answer
668
+2.75
+1.49
All problems
4,499
+2.51
+1.26
Appendix
TABLE XIV: Contributions to the MATH gain over base.
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used for post-training Large Language Models, but practical verifiers can make errors. We study how the rate and structure of reward noise affect RLVR in code generation, with a preliminary scientific-reasoning check. In multi-seed Qwen3 8B experiments on MBPP, mean validation reward over two fixed late evaluations is within 1 percentage point of the clean baseline at the tested resampled group-rollout noise rates through 20%, and within about 2 points at 30%. Confidence intervals allow larger losses; these point estimates do not establish a general tolerance threshold. We also examine full-program pass@k, four controlled noise structures, two model-based verifiers, and policy models from three families spanning 4B-9B parameters. We derive conditional advantage distributions for symmetric and asymmetric group noise, including retained format penalties, and show why clipping limits simple gradient-scaling arguments. The analysis identifies information preserved by whole-group corruption and limits on interpreting our asymmetric sweep as a precision-recall comparison. Overall, the results indicate that imperfect verification can support effective RLVR in the tested settings, while aggregate error rates alone do not characterize the learning signal.
While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale. Our approach adapts PAC-Bayes compression bounds to this setting, and addresses the inherent stochasticity of token generation by applying the Gumbel-max reparameterization trick. To operationalize these bounds, we propose the Progressive RLVR framework, which integrates RLVR with on-policy distillation, TinyLoRA, and model quantization. Progressive RLVR empirically retains 84-97% performance of standard LoRA fine-tuning while producing models that are 14,796x more compressible. We show that this framework yields non-vacuous generalization bounds in four domains: mathematical problem-solving, programming, general-knowledge reasoning, and Text-to-SQL. Our bounds exceed the accuracy of the base model by 9-51% and lie within 6-11% of the accuracy of the fine-tuned models.
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training language models, but in practice, verifiers are rarely perfect. Recent theoretical work predicts that verifier noise affects the rate of learning but not its final outcome, implying that sufficient compute should close any gap induced by imperfect supervision. We test this prediction empirically by post-training Qwen2.5 (0.5B, 1.5B) with GRPO on GSM8K while injecting controlled false-positive and false-negative noise into the binary correctness signal, and varying rollouts per prompt as a compute axis. In practice, the gap in validation accuracy persists under substantial compute scaling, with returns to compute that are sharply diminishing. We further find a structural asymmetry where false negatives monotonically degrade performance more quickly than false positives. These findings suggest verifier quality and training compute are not interchangeable, and that reducing false negatives is a more effective lever than scaling compute alone.