Organizations: Beihang University · Zhongguancun Academy · Communication University of China · Nanyang Technological University · Lobachebsky University · Peking University · Hangzhou Innovation Institute of Beihang University
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts. However, existing methods overlook how reliably learning signals are extracted from sampled responses. In GRPO, a response's advantage depends on both its own outcome and the randomly sampled outcomes of its peers through group normalization. Our theoretical and experimental analyses show that uncertainty in group composition introduces composition noise, a non-vanishing variance component that imposes an irreducible lower bound on gradient estimation error and impairs downstream prompt selection. We propose MaPP (Marginalized Posterior-Predictive), a unified framework for data-efficient RLVR that denoises response-level advantage estimation and improves prompt selection using a shared Beta posterior. For each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage through closed-form Beta-Binomial marginalization. The resulting posterior-predictive estimator has an error that provably diminishes as the posterior concentrates. Using the same posterior, MaPP derives an uncertainty-aware prompt selection score to improve data efficiency without additional rollout cost. Experiments on mathematics, planning, and visual geometry across five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget and setting a new state of the art.
Figures & tables
Figure 1: On the CountDown dataset using the Qwen3-4B model, (a) MaPP achieves higher accuracy than all selection methods with the same rollout cost, while requiring 80% fewer rollouts than DS to reach comparable performance. (b) Group-composition uncertainty induces composition noise in GRPO (group size G=8 ; mid-training checkpoint of MoPPS; 100 prompts × 50 groups): under the same policy checkpoint, identical correct (red) or incorrect (blue) response receives widely varying advantages depending on group composition ( ±1 std shaded). The noise persists across all non-degenerate difficulty levels, e.g. , Δ=2.22 at γ≈0.15 and Δ=1.62 at γ≈0.8 .
Figure 2: Overview of the MaPP framework. A shared Beta posterior (αtτ,βtτ) , updated online from rollout history, serves as the foundation for two closed-form components. (A) MaPP-AD marginalizes aj⋆(γ) over a leave-one-out posterior, replacing the composition-dependent GRPO advantage with composition-invariant weights. Both components reuse the same posterior at no additional rollout cost. (B) MaPP-PS marginalizes VG(γ) over the full posterior for uncertainty-aware prompt selection, replacing the point-estimate scoring used by existing methods.
Models
Methods
CountDown
AMC
MATH500
Minerva.
Olympiad.
Avg. ↑
Rollouts ↓
GPU Hours ↓
Qwen3-4B
Random
80.95
54.22
78.63
28.31
38.25
56.07
563k
90h
MoPPS
83.68
60.24
81.04
28.31
44.28
59.51
563k
81h
DPS
83.42
59.04
80.65
28.68
43.98
59.15
563k
81h
DS
83.83
60.24
82.06
30.88
45.03
60.41
2252k
209h
Ours
85.67
62.65
82.66
31.25
47.59
61.96
563k
81h
Qwen3-8B
Random
81.29
55.42
81.05
30.51
42.17
58.09
563k
104h
Table 1: Evaluation results on CountDown3to4 and mathematical reasoning benchmarks. Bold indicates the best result. CountDown is evaluated after CountDown3to4 training, while all other benchmarks use math-trained models. Rollouts and GPU hours report math-training cost.
Figure 3: Performance comparisons of different methods and models on the MATH and Countdown tasks. Our proposed MaPP outperforms the existing SOTA methods and other baselines in both training efficiency and performance.
λ
CountDown3to4
CountDown4
Avg.
0.0
83.51
66.19
74.85
0.1
85.17
70.04
77.61
0.5
86.76
72.24
79.50
0.7
85.95
69.18
77.57
0.9
83.08
67.33
75.21
1.0
85.82
70.72
78.27
Table 2: Sensitivity to the temporal discount λ . Trained on Qwen3-4B, Countdown.
λ
CountDown3to4
CountDown4
Avg.
0.0
83.51
66.19
74.85
0.1
85.17
70.04
77.61
0.5
86.76
72.24
79.50
0.7
85.95
69.18
77.57
0.9
83.08
67.33
75.21
1.0
85.82
70.72
78.27
Table 2: Sensitivity to the temporal discount λ . Trained on Qwen3-4B, Countdown.
G=4
G=16
Method
CD-34
CD-4
#Rollouts
CD-34
CD-4
#Rollouts
Random
79.39
60.14
123k
82.54
64.01
492k
MoPPS
80.22
62.70
123k
86.39
72.79
492k
DPS
82.17
63.26
123k
86.90
70.18
492k
DS
81.85
64.11
490k
83.87
66.66
1855k
Ours
82.45
64.44
123k
88.33
73.42
492k
Table 3: Robustness to rollout group size G . Trained on Qwen3-4B, Countdown.
Method
MATH
Olympiad.
Avg.
GRPO
79.41
38.25
58.83
+ MaPP-PS
81.03
43.17
62.10
+ MaPP-AD
82.07
44.52
63.30
+ AD + PS
82.85
47.59
65.22
MoPPS
82.17
44.28
63.23
+ MaPP-AD
82.63
45.90
64.27
Table 4: Ablation on the two components. Trained on MATH, Qwen3-4B. MaPP-AD: posterior-predictive advantage (§ 4.2 ). MaPP-PS: posterior-predictive selection (§ 4.3 ).
Figure A: (a) Calibration of the intrinsic advantage: the theoretical prediction aj⋆(γ) (Proposition 4.1 ) closely matches the empirical mean advantage across 100 prompts × 50 groups ( G=8 ). The ±1 std band reflects the composition noise that MaPP eliminates via marginalization. (b) Calibration of the selection score: the posterior-predictive non-zero-variance probability VGMaPP (Eq. ( 15 )) tracks the empirical effective ratio within a ±0.05 band.
Figure B: Ablation experiments of our proposed MaPP method under different numbers of rollouts ( n=4 and n=16 , with n=8 evaluated in the main experiments). The results show that MaPP performs consistently well across different rollout group sizes.
Figure C: (a): Performance of our method trained on the Qwen2.5-VL-Instruct series models and evaluated on the Geometry test set. (b): Self-ablation of MaPP under different posterior tracking discounts. The best performance is achieved at λ=0.5 (Ours).
In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO), the KL-free reward-advantage term studied here depends on within-group reward variation. If all rollouts in a group are correct or all are wrong, their group-relative advantages are identically zero; these zero-advantage silent groups provide no reward-advantage gradient, yet uniform sampling spends 39% of a run's rollouts on them. History-based prompt selection must first spend target-policy rollouts to estimate difficulty, creating a cold start with rollout waste; ThinkPrior instead uses an external anchor in one offline pass to construct a zero-rollout difficulty prior before the first target-policy rollout. The verifier-scored anchor pass rate supplies an external-anchor initialization for a Beta posterior; ThinkPrior selects by expected learnability and then updates from training outcomes, changing neither the loss nor the optimizer. On Qwen2.5-Math-7B across sixteen seeds, ThinkPrior more than halves early silent groups and cuts wasted rollouts through step 30 by nearly a fifth, while we detect no difference in final accuracy. On this 250-prompt pool the fixed-budget result is a reallocation rather than a net saving. The measured ThinkPrior+DAPO composition reduces generated rollouts by 10.6% while both arms retain the same 3840-rollout update budget. The prior requires no target-policy rollout before the first selection, but the posterior thereafter uses target-policy outcomes.
Tommy Sha, Skylar Zhai, Siqi Zhao
1Stony Brook University · University of Minnesota Twin Cities
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation. However, existing pre-rollout methods struggle to balance exploitation and exploration: repeatedly exploiting historically informative prompts can narrow training coverage, whereas broader exploration can lower the fraction of informative prompts. To address these limitations, we introduce LEEPS, a Latent-Guided Explore--Exploit Prompt Sampler that adaptively balances the reuse of previously observed informative prompts with continued exploration of uncertain ones. LEEPS partitions candidates into exploit and explore portfolios and adaptively allocates rollout budget according to their recent non-trivial ratios. It further uses representation-space neighbors and historical rollout outcomes to prioritize uncertain prompts likely to yield non-zero reward variance, thereby making exploration more targeted without additional rollouts. Across six mathematical reasoning benchmarks, LEEPS achieves the highest average score at both model scales, with relative gains of 2.6% and 3.7% over the strongest baseline for Qwen2.5-Math-1.5B and 7B, respectively, and generally improves faster during the training process. It also achieves the highest average score across the three evaluated OOD general-reasoning benchmarks at both model scales and adds only about 2 seconds of online sampling overhead per training step. Code is available at https://github.com/ShuangLiangX/LEEPS.
Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimization methods, such as GRPO, typically allocate a fixed number of rollouts to every prompt. This uniform allocation can be inefficient: it over-allocates compute to prompts whose sampled groups are already saturated while under-exploring prompts for which additional samples may reveal useful correct trajectories. To address this limitation, we introduce hit utility, the posterior probability that at least one rollout in a proposed additional allocation for a prompt will be correct. Building on this notion, we propose Hit-Utility Optimal Rollout Allocation (HORA), a learning-free rollout allocation policy that maximizes total posterior hit utility within each allocation batch. HORA adaptively reallocates rollout budgets while leaving the downstream reward evaluation and group-based advantage estimator unchanged. Across four mathematical reasoning benchmarks and three model scales, HORA preserves comparable Pass@1 and improves Pass@K over compute-matched GRPO in ten of twelve model--benchmark configurations, with one tie and one saturated exception. It is also drop-in compatible with other group-based estimators such as RLOO. Ablation studies indicate that the uniform prior used by HORA is competitive with five prompt-conditioned learned-prior alternatives.
Tao Wang, Shuo Li, Yan Sun +2
University of Pennsylvania · New Jersey Institute of Technology · University of Tennessee, Knoxville