Authors: Xinyu Lu, Kaiqi Zhang, Jinglin Yang, Boxi Cao, Yaojie Lu, Hongyu Lin, Min He, Xianpei Han, +1 more
Organizations: Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · National Computer Network Emergency Response Technical Team/Coordination Center of China
Reinforcement Learning with Verifiable Rewards (RLVR) enhances Large Language Model (LLM) reasoning but is suffer from advantage collapse: when all rollouts of a query receive identical rewards, the group variance vanishes, most damagingly on hard samples, where scaling rollout budgets yields little. We introduce Joint Policy and Prompt Optimization (P O) to mitigate this collapse by alternating continuous policy updates with discrete prompt evolution. P O mines hard samples with a success-rate threshold, evolves reasoning prompts for them with GEPA, and internalizes the elicited trajectories via context distillation, which optimizes each trajectory under the original query and thus removes inference-time prompting, with a Context Ratio Mask (CRM) filtering out extreme likelihood ratios. P O restores critical advantage signals and surpasses the GRPO baseline by up to 8.2 points in average accuracy on six held-out benchmarks across all training datasets and backbones, while also outperforming DAPO and other baselines. The gains are especially pronounced on hard benchmarks, reaching up to 16.3 points above GRPO on average across AIME24 and AIME25. Our findings expose the limits of standard exploration in sparse-reward environments, illuminating the potential of unifying evolutionary algorithms with reinforcement learning. This integration of discrete semantic search and continuous parameter updates provides a self-reinforcing framework that facilitates more effective LLM alignment.
Figures & tables
Figure 1: Conceptual Illustration of the P 2 O Framework. Standard policy optimization often gets trapped in local optima ( ρinit ) due to sparse rewards on hard samples. P 2 O bridges this exploration gap using optimized prompts (the red arrow) to reach high-reward regions that are inaccessible via standard exploration. Subsequently, the model consolidates these gains (the white arrows) by updating its parameters to master the new region ( ρopt ), effectively internalizing the prompt-induced capabilities.
Figure 2: Overview of the P 2 O Framework. The training process is formulated as an alternating maximization procedure between two phases: (1) Policy Optimization with Context Distillation, where the policy πθ is updated under a CRM to internalize reasoning patterns elicited by augmented inputs x~ ; and (2) Evolutionary Prompt Optimization, where the prompt template set Z is evolved using GEPA to discover successful trajectories for the remaining hard samples ( Dhard ).
Model
AIME24
AIME25
AMC
MATH500
Minerva
Olympiad
AVG
Base Models
Qwen3-4B
23.8
19.4
68.0
82.8
31.6
49.9
45.9
Qwen3-8B
28.3
22.1
69.4
81.6
31.2
49.5
47.0
DeepMath-5K
Qwen3-4B-GRPO Rollout-6
33.8
29.0
79.5
87.6
41.5
57.5
54.8
Qwen3-4B-GRPO Rollout-8
35.6
28.3
79.7
89.0
40.1
56.4
54.9
Table 1: Comparative Evaluation on Mathematical Reasoning Benchmarks. Accuracy (%) of P 2 O against the baselines; gains are largest on challenging benchmarks such as AIME24 and AIME25. Best results are highlighted in bold .
Model
AIME24
AIME25
AMC
MATH500
Minerva
Olympiad
AVG
Qwen3-4B-GRPO
46.9
37.7
88.1
90.4
41.5
58.2
60.5
Qwen3-4B-P 2 O
59.8
49.4
92.2
91.4
36.4
62.2
65.2
Qwen3-4B-P 2 O w/o-Context-Distillation
36.5
31.5
81.1
89.6
40.1
54.8
55.6
Qwen3-4B-P 2 O Same-Template-in-Group
57.3
44.6
88.9
90.8
40.4
63.0
64.2
Table 2: Ablation Study of P 2 O Components. Comparison of performance on mathematical benchmarks when excluding context distillation or group prompt diversity. Data points are derived from the Teacher-Ref w/o-CRM variant on the DeepScaler-5K dataset.
Variant
P 2 O
w/o-CRM
Δ
DeepMath-5K
Qwen3-4B Self-Ref
63.0
64.2
↓ 1.2
Qwen3-4B Teacher-Ref
58.5
57.9
↑ 0.6
Qwen3-8B Self-Ref
64.0
61.3
↑ 2.7
Qwen3-8B Teacher-Ref
63.4
61.0
↑ 2.4
DeepScaler-5K
Table 3: Effect of the Context Ratio Mask. Six-benchmark average accuracy (%) per training set, backbone, and reflection source; Δ (P 2 O − w/o-CRM) marked ↑ / ↓ . See Table 9 .
Variant
P 2 O
w/o-CRM
Δ
DeepMath-5K
Qwen3-4B Self-Ref
63.0
64.2
↓ 1.2
Qwen3-4B Teacher-Ref
58.5
57.9
↑ 0.6
Qwen3-8B Self-Ref
64.0
61.3
↑ 2.7
Qwen3-8B Teacher-Ref
63.4
61.0
↑ 2.4
DeepScaler-5K
Table 3: Effect of the Context Ratio Mask. Six-benchmark average accuracy (%) per training set, backbone, and reflection source; Δ (P 2 O − w/o-CRM) marked ↑ / ↓ . See Table 9 .
α
β
Qwen3-8B
Δ vs. w/o CRM
w/o CRM
61.3
–
0.8
1.2
58.3
↓ 3.0
0.05
5
63.1
↑ 1.8
0.01
10
64.0
↑ 2.7
0.005
50
61.5
↑ 0.2
Table 4: Hyper-parameter Ablation of the Context Ratio Mask. Six-benchmark average accuracy (%) for Qwen3-8B on the DeepMath-5K training set with Self-Ref, as the mask bounds α and β vary; “w/o CRM” removes the mask entirely. The default interval [0.01,10] reaches 64.0%, 2.7 points above w/o CRM. Best result in bold .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Policy Rollouts
GEPA Rollouts
Total Rollouts/Epoch
GRPO Rollout-6
6N
–
6N
GRPO Rollout-8
8N
–
8N
GRPO Rollout-12
12N
–
12N
P 2 O Self-Ref ( K=6 )
6N
2N
8N
P 2 O Teacher-Ref ( K=6 )
6N
2N
8N
Appendix
Table 5: Per-Epoch Rollout Budgets. All entries are rollout counts per epoch, where N is the number of training samples ( N=5,000 for DeepMath-5K and DeepScaler-5K, so that 2N=10 K). GEPA rollouts include everything charged to the budget of Algorithm 3 : mini-batch rollouts, reflection calls (including all Kimi-K2 calls), and all dev evaluations; GRPO uses none. Because the GEPA budget is a hard cap, P 2 O spends 8N rollouts per epoch at K=6 and therefore matches GRPO Rollout-8 .
Reward Variance
Reward Mean
Bucket
# Samples (%)
# Improved (%)
w/o Template
w/ Template
w/o Template
w/ Template
All hard samples
1980 (100%)
740 (37.4%)
0.2494
0.2491
0.4765
0.5295
All-incorrect ( ∑k=1Krik=0 )
559 (28.2%)
202 (36.1%)
0
0.1205
0
0.1401
Near-all-incorrect ( ∑k=1Krik/K≤1/6 )
745 (37.6%)
321 (43.1%)
0.0399
0.1617
0.0416
0.2029
Appendix
Table 6: Advantage collapse on the hard samples of DeepMath-5K at the end of epoch 0. Self-Ref w/o-CRM variant with K=6 rollouts per query. A sample is counted as Improved if its average reward over the K rollouts is higher when the template is applied than under the bare query. Variance and mean are averaged over the samples of each bucket; the two extreme buckets are nested inside the set of all hard samples.
Metric
Self-Ref
Teacher-Ref
Δ (Teacher − Self)
Mean dev-hard score
16.2
24.4
↑ 8.2
Best dev-hard score
26.7
31.7
↑ 5.0
Final downstream 6-benchmark AVG
64.2
57.9
↓ 6.3
Appendix
Table 7: Template quality versus downstream performance DeepMath-5K, Qwen3-4B, w/o-CRM. Dev-hard scores are accuracies (%) of the template set produced by GEPA at the end of epoch 0, averaged over all candidate templates ( mean ) and taken at the best template ( best ), where both variants evaluate on the same dev-hard split with the same policy weights. The last row reports the final six-benchmark average of the fully trained models (Table 9 ). Δ=Teacher-Ref−Self-Ref .
Model
AIME24
AIME25
AMC
MATH500
Minerva
Olympiad
AVG
Qwen3-4B-P 2 O τ=2/6
56.7
45.4
86.9
92.4
40.4
63.1
64.2
Qwen3-4B-P 2 O τ=1/6
50.0
42.5
87.3
90.2
40.8
61.6
62.1
Qwen3-4B-P 2 O τ=4/6
57.3
44.6
88.0
91.6
42.6
60.4
64.1
Appendix
Table 8: Hard Sample Success-rate Threshold τ Ablation. Results on DeepMath-5K using the Self-Ref w/o-CRM variant with K=6 . Thresholds are reported as rates. Our default τ=2/6 treats samples with at most one correct rollout as hard.
Model
AIME24
AIME25
AMC
MATH500
Minerva
Olympiad
AVG
DeepMath-5K
Qwen3-4B-P 2 O Self-Ref
55.3
40.0
86.2
91.2
43.0
62.4
63.0
Qwen3-4B-P 2 O Self-Ref, w/o-CRM
56.7
45.4
86.9
92.4
40.4
63.1
64.2
Qwen3-4B-P 2 O Teacher-Ref
46.5
32.5
82.3
88.6
41.5
59.3
58.5
Qwen3-4B-P 2 O Teacher-Ref, w/o-CRM
44.6
34.4
81.9
90.2
36.0
60.0
57.9
Qwen3-8B-P 2 O Self-Ref
57.1
41.7
86.7
92.2
43.0
63.4
64.0
Appendix
Table 9: Full Results of the Context Ratio Mask Ablation. Benchmark-wise accuracy (%) of P 2 O and its w/o-CRM counterpart with the Qwen3-4B and Qwen3-8B backbones, on the DeepMath-5K and DeepScaler-5K training sets and the two reflection sources. AVG averages the six benchmarks.
RL with verifiable rewards can substantially improve LLM reasoning, yet standard GRPO-style training often treats easy, hard, and learnable questions alike through uniform sampling and weighting, leading to inefficient compute allocation. We study GRPO by tracking token log-probabilities, group-normalized advantages, and the induced token-level update weights. This reveals three recurring dynamics as training proceeds: (1) confidence inflation, (2) advantage contraction, and (3) hierarchical convergence. These findings suggest that the utility of each update depends strongly on both question difficulty and the model's current competence. Motivated by this, we propose Confidence and Difficulty-adaptive Policy Optimization (CoDaPO), which assigns each question a bounded value from rollout confidence and empirical difficulty. CoDaPO then uses this value to reweight policy updates and resample high-value learnable questions within mini-batches, thereby increasing discovery within the learnable band under a fixed compute budget. Across twelve benchmarks, CoDaPO consistently improves accuracy over existing RL methods. Our code is publicly available at https://github.com/tmlr-group/CoDaPO.
Zhanke Zhou, Xiangyu Lu, Chentao Cao +4
1TMLR Group, Department of Computer Science, Hong Kong Baptist University · 2Stanford University · 3Sydney AI Centre, The University of Sydney
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising approach to improve the reasoning abilities of Large Language Models (LLMs). Among RLVR algorithms, Group Relative Policy Optimization (GRPO) and its variants have demonstrated strong performance and high training efficiency. However, GRPO-style objectives exhibit two issues on high accuracy prompts including mastered prompts (rollout accuracy =1) and majority-correct prompts (rollout accuracy in (0.5,1)). For mastered prompts, group-relative advantages vanish, yielding no training signal and unconstrained policy drift that can cause forgetting. For majority-correct prompts, the induced query weight shrinks as accuracy increases, weakening consolidation from partial correctness to mastery. To alleviate this, we propose Mastery-Consolidated Policy Optimization (MCPO), which introduces (i) a hinge-KL regularizer applied exclusively to mastered prompts to bound harmful policy drift between successive gradient steps, and (ii) a weighting mechanism that prioritizes majority-correct prompts to better allocate optimization effort. Extensive experiments across three mathematical benchmarks demonstrate that MCPO consistently improves pass@1 performance. Counter-intuitively, rather than restricting exploration, MCPO boosts pass@k metrics, indicating that mastery consolidation further catalyzes solution diversity.
Reinforcement learning with verifiable rewards, particularly Group Relative Policy Optimization (GRPO), has significantly advanced the reasoning capabilities of Large Language Models (LLMs). However, in complex tasks, GRPO frequently suffers from the ``zero-advantage problem'': when all sampled rollouts for a query fail, the relative advantage collapses to zero. Consequently, the model loses effective training signals for these questions, wasting the training data and computational budget. While simply increasing the sampling budget for these questions is a common remedy, the static sampling policy inherently constrains reasoning exploration, limiting the success rate. In this paper, we propose Lorem Perturbation for Exploration (LoPE), a simple yet effective training framework to break this exploration bottleneck. We posit that task-irrelevant prompt-space perturbations can shift the model's output distribution enough to unlock orthogonal reasoning pathways for hard questions. Specifically, LoPE prepends sequences stochastically assembled from Lorem Ipsum vocabulary (a pseudo-Latin placeholder text) to the prompts before resampling. Experiments across 1.7B, 4B, and 7B models demonstrate that LoPE significantly outperforms resampling with the original prompts. Further analysis reveals that other Latin-based random sequences with low perplexity are also effective perturbations. Our results establish LoPE as a strong baseline for broadening exploration in LLM reinforcement learning.