Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO
Authors: Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen, Tianhuang Su, Haonan Lu, Quanlong Guan, +2 more
Organizations: Guangdong Institute of Smart Education, Jinan University, Guangzhou, China · OPPO AI Center, Shenzhen, China · Ragentile Intelligence Inc, Edmonton, Canada
Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07× training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://github.com/Nicozwy/FastRL.
Figures & tables
Figure 1: Comparison of training efficiency and out-of-domain performance across reinforcement learning methods using Qwen2.5-VL-7B-Instruct on Geometry3K (left) and GeoQA8K-R1V (right).
Figure 2: The proposed FastRL framework. (a) Maximizing Differential Advantage Pruning (MDAP) identifies and prunes uninformative trajectories (i.e., θij<τ ) and homogeneous trajectories (i.e., Ai=Aj ); (b) Adaptive Rollout Sampling (ARS) dynamically adjusts the rollout budget for each iteration based on the accumulated pruning distribution.
Dataset
Method
In-domain
MathVerse
MathVision
MathVista
WeMath
Avg( ↑ )
Train-Time( ↓ )
Speed( ↑ )
Qwen2.5-VL-7B-Instruct
Geometry3K
GRPO [ 15 ]
53.74
43.38
26.25
66.90
68.74
51.32
29.72
1 ×
+CPPO [ 6 ]
54.41
43.98
27.37
66.30
69.08
51.68
18.52
1.60 ×
+GRESO [ 25 ]
53.24
42.08
26.09
64.90
67.93
50.25
27.29
1.09 ×
+FastRL (Ours)
55.90
45.76
27.96
67.30
69.48
52.63
13.87
2.14 ×
DAPO [ 21 ]
53.57
41.24
26.25
67.20
70.69
51.35
28.08
1 ×
Table 1: Performance comparison (%) on Geometry3K and GeoQA8K-R1V regarding accuracy. The bold numbers denote the best results. Avg denotes the average accuracy across out-of-domain datasets, and Speed denotes the training speed regarding GRPO. ↑ indicates that the higher is better, while ↓ indicates that the lower is better.
Table 2: Ablation study (%) of FastRL. “w/o MDAP” denotes FastRL without maximizing differential advantage pruning (with ARS simulated to match its standard dynamics for a fair comparison), and “w/o ARS” denotes FastRL without adaptive rollout sampling.
Figure 3: Training efficiency and parameter sensitivity of FastRL on Geometry3K (top) and GeoQA8K-R1V (bottom). (a1)-(a2) show the the rollout sampling scale Gt ; (b1)-(b2) present the number of trajectories used for policy updates; (c1)-(c2) show the sensitivity to the hyperparameter τ .
Dataset
Method
In-domain
AIME2023
AIME2024
AIME2025
AIME2026
Avg( ↑ )
Train-Time( ↓ )
Speed( ↑ )
Llama3.1-8B-Instruct
MATH
GRPO [ 15 ]
72.32
10.00
10.00
6.67
3.33
7.50
21.16
1 ×
DAPO [ 21 ]
72.72
6.67
10.00
6.67
3.33
6.67
19.72
1.07 ×
GSPO [ 24 ]
72.56
6.67
10.00
0.00
0.00
4.17
18.89
1.12 ×
CPPO [ 6 ]
71.92
3.33
10.00
0.00
3.33
4.17
11.44
1.85 ×
GRESO [ 25 ]
71.60
3.33
6.67
3.33
0.00
3.33
17.35
1.22 ×
Table 3: Generalization across model architectures and datasets.
Dataset
Method
MathVerse
MathVision
MathVista
WeMath
avg@16( ↑ )
Geometry3K
GRPO [ 15 ]
60.53
62.96
82.50
90.34
74.08
DAPO [ 21 ]
58.98
62.57
81.50
90.57
73.41
GSPO [ 24 ]
59.21
62.30
81.90
90.11
73.38
GRESO [ 25 ]
59.44
62.47
82.00
90.23
73.54
CPPO [ 6 ]
60.81
63.19
82.20
90.69
74.22
FastRL (Ours)
63.60
67.86
83.60
93.85
77.23
Table 4: Pass@16 results (%) on out-of-domain benchmarks. Base model: Qwen2.5-VL-7B-Instruct.
Pruning Strategy
MathVerse
MathVision
MathVista
WeMath
Avg( ↑ )
Speed( ↑ )
Random-pruning 0.25
43.98
27.40
66.90
68.85
51.78
1.27 ×
Random-pruning 0.50
45.35
26.74
66.60
67.70
51.59
1.60 ×
Random-pruning 0.75
42.79
26.45
64.60
66.49
50.08
2.10 ×
Advantage-only (Ours)
44.54
27.57
67.10
69.14
52.09
2.35 ×
FastRL (Ours)
45.76
27.96
67.30
69.48
52.63
2.14 ×
Table 5: Ablation (%) of pruning strategies on Geometry3K using Qwen2.5-VL-7B-Instruct.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Multimodal
Unimodal
Dataset
Sum
Train
Test
Modality
Dataset
Sum
Train
Test
Modality
Geometry3K
2,702
2,101
601
Text + Image
MATH
12,500
11,250
1,250
Text
GeoQA8K-R1V
8,785
8,031
754
Text + Image
AIME2023
30
-
30
Text
MathVerse
3,940
-
3,940
Text + Image
AIME2024
30
-
30
Text
MathVision
3,040
-
3,040
Text + Image
AIME2025
30
-
30
Text
MathVista (mini)
1,000
-
1,000
Text + Image
AIME2026
30
-
30
Text
Appendix
Table 6: Statistics of the datasets used in our experiments.
Figure 4: Examples from the three training datasets: Geometry3K, GeoQA8k-R1V, and MATH.
Figure 5: Correlation analysis between token-level Jaccard similarity and gradient cosine similarity of trajectories with the same advantage for each question. The left plot corresponds to the Geometry3K dataset, while the right plot corresponds to the GeoQA8k-R1V dataset. The high Pearson and Spearman correlation coefficients demonstrate a strong positive correlation between the two metrics across both datasets.
Figure 6: Training dynamics and computational overhead analysis of Qwen2.5-VL-7B-Instruct on Geometry3K and GeoQA8K-R1V. We compare FastRL (ours) with GRPO, DAPO, CPPO, GSPO, and GRESO across three key metrics: total token consumption (left), policy entropy (middle), and accuracy (right).
Group Relative Policy Optimization (GRPO) has been a key driver of recent progress in reinforcement learning with verifiable rewards (RLVR) for large language models, but it is typically trained in a low-staleness, near-on-policy regime that incurs substantial system overhead. We ask a simple question: How off-policy can GRPO be? We show that GRPO-style algorithms can tolerate substantially larger rollout staleness than previously assumed, and propose Mu-GRPO, an RL training framework that organizes training into a small number (e.g., four) of large sequential generation-optimization stages. This design induces high rollout staleness while greatly reducing rollout-optimization switching overhead. To stabilize learning under stale data, Mu-GRPO combines relaxed clipping, which preserves useful stale-rollout gradients, with negative-advantage veto, which removes destabilizing post-trigger suffix updates in negative-advantage responses. Across five language models and multiple math reasoning benchmarks, Mu-GRPO matches or exceeds the performance of standard GRPO while achieving around 2x speedup in wall-clock training time, establishing a substantially improved performance-efficiency trade-off for LLM reinforcement learning.
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very different learning signals. Rollouts with low group success exhibit larger IS ratios and carry stronger gradient signal for exploration and solving new problems, yet are disproportionately suppressed by fixed clipping. To address this, we propose Group Adaptive Clipping Policy Optimization (GAPO), a plug-in modification to GRPO methods that adapts the clipping boundary to the rollout advantage. GAPO is motivated by a reverse-KL trust-region perspective, which suggests that rollouts with larger learning signal should receive proportionally greater update headroom. GAPO requires no reward shaping and preserves the standard PPO/GSPO surrogate while adapting only the clipping threshold. Across Qwen and Llama models, GAPO consistently improves both Pass@1 and Pass@k over fixed clipping and advantage-shaping baselines on math reasoning and coding benchmarks where the pass rates by the base model are relatively low.
Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieved strong results in improving the reasoning capabilities of large language models (LLMs). However, GRPO is prone to advantage collapse, a failure mode where homogeneous rewards within a group (e.g., all correct or all incorrect answers) yield near-zero advantages and vanishing gradients. To address this, we introduce the Advantage Collapse Rate (ACR), the first diagnostic metric quantifying the proportion of training batches with ineffective gradients. Across models from 0.5B to 14B parameters on mathematical reasoning benchmarks, we show that ACR strongly predicts training stagnation and final performance. We then propose Adaptive Virtual Sample Policy Optimization (AVSPO), a lightweight extension of GRPO that injects virtual reward samples, guided by real-time ACR monitoring, to enable learning from homogeneous groups without additional model rollouts. AVSPO reduces advantage collapse by 58-63% relative to GRPO and yields consistent accuracy gains of 4-6 percentage points across all model scales, while maintaining generalization on the evaluated out-of-domain task. Code and datasets are available at https://github.com/hexixiang/Advantage-Collapse-Rate.
Xixiang He, Qiyao Sun, Ao Cheng +5
National University of Defense Technology, Changsha, Hunan, China · Intelligent Game and Decision Lab, Beijing, China · The Chinese University of Hong Kong, Shenzhen, Guangdong, China