Authors: Xinyu Lu, Kaiqi Zhang, Jinglin Yang, Boxi Cao, Yaojie Lu, Hongyu Lin, Min He, Xianpei Han, +1 more
Organizations: Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · National Computer Network Emergency Response Technical Team/Coordination Center of China
Reinforcement Learning with Verifiable Rewards (RLVR) enhances Large Language Model (LLM) reasoning but is suffer from advantage collapse: when all rollouts of a query receive identical rewards, the group variance vanishes, most damagingly on hard samples, where scaling rollout budgets yields little. We introduce Joint Policy and Prompt Optimization (P O) to mitigate this collapse by alternating continuous policy updates with discrete prompt evolution. P O mines hard samples with a success-rate threshold, evolves reasoning prompts for them with GEPA, and internalizes the elicited trajectories via context distillation, which optimizes each trajectory under the original query and thus removes inference-time prompting, with a Context Ratio Mask (CRM) filtering out extreme likelihood ratios. P O restores critical advantage signals and surpasses the GRPO baseline by up to 8.2 points in average accuracy on six held-out benchmarks across all training datasets and backbones, while also outperforming DAPO and other baselines. The gains are especially pronounced on hard benchmarks, reaching up to 16.3 points above GRPO on average across AIME24 and AIME25. Our findings expose the limits of standard exploration in sparse-reward environments, illuminating the potential of unifying evolutionary algorithms with reinforcement learning. This integration of discrete semantic search and continuous parameter updates provides a self-reinforcing framework that facilitates more effective LLM alignment.
Figures & tables
Figure 1: Conceptual Illustration of the P 2 O Framework. Standard policy optimization often gets trapped in local optima ( ρinit ) due to sparse rewards on hard samples. P 2 O bridges this exploration gap using optimized prompts (the red arrow) to reach high-reward regions that are inaccessible via standard exploration. Subsequently, the model consolidates these gains (the white arrows) by updating its parameters to master the new region ( ρopt ), effectively internalizing the prompt-induced capabilities.
Figure 2: Overview of the P 2 O Framework. The training process is formulated as an alternating maximization procedure between two phases: (1) Policy Optimization with Context Distillation, where the policy πθ is updated under a CRM to internalize reasoning patterns elicited by augmented inputs x~ ; and (2) Evolutionary Prompt Optimization, where the prompt template set Z is evolved using GEPA to discover successful trajectories for the remaining hard samples ( Dhard ).
Model
AIME24
AIME25
AMC
MATH500
Minerva
Olympiad
AVG
Base Models
Qwen3-4B
23.8
19.4
68.0
82.8
31.6
49.9
45.9
Qwen3-8B
28.3
22.1
69.4
81.6
31.2
49.5
47.0
DeepMath-5K
Qwen3-4B-GRPO Rollout-6
33.8
29.0
79.5
87.6
41.5
57.5
54.8
Qwen3-4B-GRPO Rollout-8
35.6
28.3
79.7
89.0
40.1
56.4
54.9
Table 1: Comparative Evaluation on Mathematical Reasoning Benchmarks. Accuracy (%) of P 2 O against the baselines; gains are largest on challenging benchmarks such as AIME24 and AIME25. Best results are highlighted in bold .
Model
AIME24
AIME25
AMC
MATH500
Minerva
Olympiad
AVG
Qwen3-4B-GRPO
46.9
37.7
88.1
90.4
41.5
58.2
60.5
Qwen3-4B-P 2 O
59.8
49.4
92.2
91.4
36.4
62.2
65.2
Qwen3-4B-P 2 O w/o-Context-Distillation
36.5
31.5
81.1
89.6
40.1
54.8
55.6
Qwen3-4B-P 2 O Same-Template-in-Group
57.3
44.6
88.9
90.8
40.4
63.0
64.2
Table 2: Ablation Study of P 2 O Components. Comparison of performance on mathematical benchmarks when excluding context distillation or group prompt diversity. Data points are derived from the Teacher-Ref w/o-CRM variant on the DeepScaler-5K dataset.
Variant
P 2 O
w/o-CRM
Δ
DeepMath-5K
Qwen3-4B Self-Ref
63.0
64.2
↓ 1.2
Qwen3-4B Teacher-Ref
58.5
57.9
↑ 0.6
Qwen3-8B Self-Ref
64.0
61.3
↑ 2.7
Qwen3-8B Teacher-Ref
63.4
61.0
↑ 2.4
DeepScaler-5K
Table 3: Effect of the Context Ratio Mask. Six-benchmark average accuracy (%) per training set, backbone, and reflection source; Δ (P 2 O − w/o-CRM) marked ↑ / ↓ . See Table 9 .
Variant
P 2 O
w/o-CRM
Δ
DeepMath-5K
Qwen3-4B Self-Ref
63.0
64.2
↓ 1.2
Qwen3-4B Teacher-Ref
58.5
57.9
↑ 0.6
Qwen3-8B Self-Ref
64.0
61.3
↑ 2.7
Qwen3-8B Teacher-Ref
63.4
61.0
↑ 2.4
DeepScaler-5K
Table 3: Effect of the Context Ratio Mask. Six-benchmark average accuracy (%) per training set, backbone, and reflection source; Δ (P 2 O − w/o-CRM) marked ↑ / ↓ . See Table 9 .
α
β
Qwen3-8B
Δ vs. w/o CRM
w/o CRM
61.3
–
0.8
1.2
58.3
↓ 3.0
0.05
5
63.1
↑ 1.8
0.01
10
64.0
↑ 2.7
0.005
50
61.5
↑ 0.2
Table 4: Hyper-parameter Ablation of the Context Ratio Mask. Six-benchmark average accuracy (%) for Qwen3-8B on the DeepMath-5K training set with Self-Ref, as the mask bounds α and β vary; “w/o CRM” removes the mask entirely. The default interval [0.01,10] reaches 64.0%, 2.7 points above w/o CRM. Best result in bold .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Policy Rollouts
GEPA Rollouts
Total Rollouts/Epoch
GRPO Rollout-6
6N
–
6N
GRPO Rollout-8
8N
–
8N
GRPO Rollout-12
12N
–
12N
P 2 O Self-Ref ( K=6 )
6N
2N
8N
P 2 O Teacher-Ref ( K=6 )
6N
2N
8N
Appendix
Table 5: Per-Epoch Rollout Budgets. All entries are rollout counts per epoch, where N is the number of training samples ( N=5,000 for DeepMath-5K and DeepScaler-5K, so that 2N=10 K). GEPA rollouts include everything charged to the budget of Algorithm 3 : mini-batch rollouts, reflection calls (including all Kimi-K2 calls), and all dev evaluations; GRPO uses none. Because the GEPA budget is a hard cap, P 2 O spends 8N rollouts per epoch at K=6 and therefore matches GRPO Rollout-8 .
Reward Variance
Reward Mean
Bucket
# Samples (%)
# Improved (%)
w/o Template
w/ Template
w/o Template
w/ Template
All hard samples
1980 (100%)
740 (37.4%)
0.2494
0.2491
0.4765
0.5295
All-incorrect ( ∑k=1Krik=0 )
559 (28.2%)
202 (36.1%)
0
0.1205
0
0.1401
Near-all-incorrect ( ∑k=1Krik/K≤1/6 )
745 (37.6%)
321 (43.1%)
0.0399
0.1617
0.0416
0.2029
Appendix
Table 6: Advantage collapse on the hard samples of DeepMath-5K at the end of epoch 0. Self-Ref w/o-CRM variant with K=6 rollouts per query. A sample is counted as Improved if its average reward over the K rollouts is higher when the template is applied than under the bare query. Variance and mean are averaged over the samples of each bucket; the two extreme buckets are nested inside the set of all hard samples.
Metric
Self-Ref
Teacher-Ref
Δ (Teacher − Self)
Mean dev-hard score
16.2
24.4
↑ 8.2
Best dev-hard score
26.7
31.7
↑ 5.0
Final downstream 6-benchmark AVG
64.2
57.9
↓ 6.3
Appendix
Table 7: Template quality versus downstream performance DeepMath-5K, Qwen3-4B, w/o-CRM. Dev-hard scores are accuracies (%) of the template set produced by GEPA at the end of epoch 0, averaged over all candidate templates ( mean ) and taken at the best template ( best ), where both variants evaluate on the same dev-hard split with the same policy weights. The last row reports the final six-benchmark average of the fully trained models (Table 9 ). Δ=Teacher-Ref−Self-Ref .
Model
AIME24
AIME25
AMC
MATH500
Minerva
Olympiad
AVG
Qwen3-4B-P 2 O τ=2/6
56.7
45.4
86.9
92.4
40.4
63.1
64.2
Qwen3-4B-P 2 O τ=1/6
50.0
42.5
87.3
90.2
40.8
61.6
62.1
Qwen3-4B-P 2 O τ=4/6
57.3
44.6
88.0
91.6
42.6
60.4
64.1
Appendix
Table 8: Hard Sample Success-rate Threshold τ Ablation. Results on DeepMath-5K using the Self-Ref w/o-CRM variant with K=6 . Thresholds are reported as rates. Our default τ=2/6 treats samples with at most one correct rollout as hard.
Model
AIME24
AIME25
AMC
MATH500
Minerva
Olympiad
AVG
DeepMath-5K
Qwen3-4B-P 2 O Self-Ref
55.3
40.0
86.2
91.2
43.0
62.4
63.0
Qwen3-4B-P 2 O Self-Ref, w/o-CRM
56.7
45.4
86.9
92.4
40.4
63.1
64.2
Qwen3-4B-P 2 O Teacher-Ref
46.5
32.5
82.3
88.6
41.5
59.3
58.5
Qwen3-4B-P 2 O Teacher-Ref, w/o-CRM
44.6
34.4
81.9
90.2
36.0
60.0
57.9
Qwen3-8B-P 2 O Self-Ref
57.1
41.7
86.7
92.2
43.0
63.4
64.0
Appendix
Table 9: Full Results of the Context Ratio Mask Ablation. Benchmark-wise accuracy (%) of P 2 O and its w/o-CRM counterpart with the Qwen3-4B and Qwen3-8B backbones, on the DeepMath-5K and DeepScaler-5K training sets and the two reflection sources. AVG averages the six benchmarks.