Organizations: Beihang University · Zhongguancun Academy · Communication University of China · Nanyang Technological University · Lobachebsky University · Peking University · Hangzhou Innovation Institute of Beihang University
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts. However, existing methods overlook how reliably learning signals are extracted from sampled responses. In GRPO, a response's advantage depends on both its own outcome and the randomly sampled outcomes of its peers through group normalization. Our theoretical and experimental analyses show that uncertainty in group composition introduces composition noise, a non-vanishing variance component that imposes an irreducible lower bound on gradient estimation error and impairs downstream prompt selection. We propose MaPP (Marginalized Posterior-Predictive), a unified framework for data-efficient RLVR that denoises response-level advantage estimation and improves prompt selection using a shared Beta posterior. For each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage through closed-form Beta-Binomial marginalization. The resulting posterior-predictive estimator has an error that provably diminishes as the posterior concentrates. Using the same posterior, MaPP derives an uncertainty-aware prompt selection score to improve data efficiency without additional rollout cost. Experiments on mathematics, planning, and visual geometry across five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget and setting a new state of the art.
Figures & tables
Figure 1: On the CountDown dataset using the Qwen3-4B model, (a) MaPP achieves higher accuracy than all selection methods with the same rollout cost, while requiring 80% fewer rollouts than DS to reach comparable performance. (b) Group-composition uncertainty induces composition noise in GRPO (group size G=8 ; mid-training checkpoint of MoPPS; 100 prompts × 50 groups): under the same policy checkpoint, identical correct (red) or incorrect (blue) response receives widely varying advantages depending on group composition ( ±1 std shaded). The noise persists across all non-degenerate difficulty levels, e.g. , Δ=2.22 at γ≈0.15 and Δ=1.62 at γ≈0.8 .
Figure 2: Overview of the MaPP framework. A shared Beta posterior (αtτ,βtτ) , updated online from rollout history, serves as the foundation for two closed-form components. (A) MaPP-AD marginalizes aj⋆(γ) over a leave-one-out posterior, replacing the composition-dependent GRPO advantage with composition-invariant weights. Both components reuse the same posterior at no additional rollout cost. (B) MaPP-PS marginalizes VG(γ) over the full posterior for uncertainty-aware prompt selection, replacing the point-estimate scoring used by existing methods.
Models
Methods
CountDown
AMC
MATH500
Minerva.
Olympiad.
Avg. ↑
Rollouts ↓
GPU Hours ↓
Qwen3-4B
Random
80.95
54.22
78.63
28.31
38.25
56.07
563k
90h
MoPPS
83.68
60.24
81.04
28.31
44.28
59.51
563k
81h
DPS
83.42
59.04
80.65
28.68
43.98
59.15
563k
81h
DS
83.83
60.24
82.06
30.88
45.03
60.41
2252k
209h
Ours
85.67
62.65
82.66
31.25
47.59
61.96
563k
81h
Qwen3-8B
Random
81.29
55.42
81.05
30.51
42.17
58.09
563k
104h
Table 1: Evaluation results on CountDown3to4 and mathematical reasoning benchmarks. Bold indicates the best result. CountDown is evaluated after CountDown3to4 training, while all other benchmarks use math-trained models. Rollouts and GPU hours report math-training cost.
Figure 3: Performance comparisons of different methods and models on the MATH and Countdown tasks. Our proposed MaPP outperforms the existing SOTA methods and other baselines in both training efficiency and performance.
λ
CountDown3to4
CountDown4
Avg.
0.0
83.51
66.19
74.85
0.1
85.17
70.04
77.61
0.5
86.76
72.24
79.50
0.7
85.95
69.18
77.57
0.9
83.08
67.33
75.21
1.0
85.82
70.72
78.27
Table 2: Sensitivity to the temporal discount λ . Trained on Qwen3-4B, Countdown.
λ
CountDown3to4
CountDown4
Avg.
0.0
83.51
66.19
74.85
0.1
85.17
70.04
77.61
0.5
86.76
72.24
79.50
0.7
85.95
69.18
77.57
0.9
83.08
67.33
75.21
1.0
85.82
70.72
78.27
Table 2: Sensitivity to the temporal discount λ . Trained on Qwen3-4B, Countdown.
G=4
G=16
Method
CD-34
CD-4
#Rollouts
CD-34
CD-4
#Rollouts
Random
79.39
60.14
123k
82.54
64.01
492k
MoPPS
80.22
62.70
123k
86.39
72.79
492k
DPS
82.17
63.26
123k
86.90
70.18
492k
DS
81.85
64.11
490k
83.87
66.66
1855k
Ours
82.45
64.44
123k
88.33
73.42
492k
Table 3: Robustness to rollout group size G . Trained on Qwen3-4B, Countdown.
Method
MATH
Olympiad.
Avg.
GRPO
79.41
38.25
58.83
+ MaPP-PS
81.03
43.17
62.10
+ MaPP-AD
82.07
44.52
63.30
+ AD + PS
82.85
47.59
65.22
MoPPS
82.17
44.28
63.23
+ MaPP-AD
82.63
45.90
64.27
Table 4: Ablation on the two components. Trained on MATH, Qwen3-4B. MaPP-AD: posterior-predictive advantage (§ 4.2 ). MaPP-PS: posterior-predictive selection (§ 4.3 ).
Figure A: (a) Calibration of the intrinsic advantage: the theoretical prediction aj⋆(γ) (Proposition 4.1 ) closely matches the empirical mean advantage across 100 prompts × 50 groups ( G=8 ). The ±1 std band reflects the composition noise that MaPP eliminates via marginalization. (b) Calibration of the selection score: the posterior-predictive non-zero-variance probability VGMaPP (Eq. ( 15 )) tracks the empirical effective ratio within a ±0.05 band.
Figure B: Ablation experiments of our proposed MaPP method under different numbers of rollouts ( n=4 and n=16 , with n=8 evaluated in the main experiments). The results show that MaPP performs consistently well across different rollout group sizes.
Figure C: (a): Performance of our method trained on the Qwen2.5-VL-Instruct series models and evaluated on the Geometry test set. (b): Self-ablation of MaPP under different posterior tracking discounts. The best performance is achieved at λ=0.5 (Ours).