Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO
Authors: Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen, Tianhuang Su, Haonan Lu, Quanlong Guan, +2 more
Organizations: Guangdong Institute of Smart Education, Jinan University, Guangzhou, China · OPPO AI Center, Shenzhen, China · Ragentile Intelligence Inc, Edmonton, Canada
Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07× training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://github.com/Nicozwy/FastRL.
Figures & tables
Figure 1: Comparison of training efficiency and out-of-domain performance across reinforcement learning methods using Qwen2.5-VL-7B-Instruct on Geometry3K (left) and GeoQA8K-R1V (right).
Figure 2: The proposed FastRL framework. (a) Maximizing Differential Advantage Pruning (MDAP) identifies and prunes uninformative trajectories (i.e., θij<τ ) and homogeneous trajectories (i.e., Ai=Aj ); (b) Adaptive Rollout Sampling (ARS) dynamically adjusts the rollout budget for each iteration based on the accumulated pruning distribution.
Dataset
Method
In-domain
MathVerse
MathVision
MathVista
WeMath
Avg( ↑ )
Train-Time( ↓ )
Speed( ↑ )
Qwen2.5-VL-7B-Instruct
Geometry3K
GRPO [ 15 ]
53.74
43.38
26.25
66.90
68.74
51.32
29.72
1 ×
+CPPO [ 6 ]
54.41
43.98
27.37
66.30
69.08
51.68
18.52
1.60 ×
+GRESO [ 25 ]
53.24
42.08
26.09
64.90
67.93
50.25
27.29
1.09 ×
+FastRL (Ours)
55.90
45.76
27.96
67.30
69.48
52.63
13.87
2.14 ×
DAPO [ 21 ]
53.57
41.24
26.25
67.20
70.69
51.35
28.08
1 ×
Table 1: Performance comparison (%) on Geometry3K and GeoQA8K-R1V regarding accuracy. The bold numbers denote the best results. Avg denotes the average accuracy across out-of-domain datasets, and Speed denotes the training speed regarding GRPO. ↑ indicates that the higher is better, while ↓ indicates that the lower is better.
Table 2: Ablation study (%) of FastRL. “w/o MDAP” denotes FastRL without maximizing differential advantage pruning (with ARS simulated to match its standard dynamics for a fair comparison), and “w/o ARS” denotes FastRL without adaptive rollout sampling.
Figure 3: Training efficiency and parameter sensitivity of FastRL on Geometry3K (top) and GeoQA8K-R1V (bottom). (a1)-(a2) show the the rollout sampling scale Gt ; (b1)-(b2) present the number of trajectories used for policy updates; (c1)-(c2) show the sensitivity to the hyperparameter τ .
Dataset
Method
In-domain
AIME2023
AIME2024
AIME2025
AIME2026
Avg( ↑ )
Train-Time( ↓ )
Speed( ↑ )
Llama3.1-8B-Instruct
MATH
GRPO [ 15 ]
72.32
10.00
10.00
6.67
3.33
7.50
21.16
1 ×
DAPO [ 21 ]
72.72
6.67
10.00
6.67
3.33
6.67
19.72
1.07 ×
GSPO [ 24 ]
72.56
6.67
10.00
0.00
0.00
4.17
18.89
1.12 ×
CPPO [ 6 ]
71.92
3.33
10.00
0.00
3.33
4.17
11.44
1.85 ×
GRESO [ 25 ]
71.60
3.33
6.67
3.33
0.00
3.33
17.35
1.22 ×
Table 3: Generalization across model architectures and datasets.
Dataset
Method
MathVerse
MathVision
MathVista
WeMath
avg@16( ↑ )
Geometry3K
GRPO [ 15 ]
60.53
62.96
82.50
90.34
74.08
DAPO [ 21 ]
58.98
62.57
81.50
90.57
73.41
GSPO [ 24 ]
59.21
62.30
81.90
90.11
73.38
GRESO [ 25 ]
59.44
62.47
82.00
90.23
73.54
CPPO [ 6 ]
60.81
63.19
82.20
90.69
74.22
FastRL (Ours)
63.60
67.86
83.60
93.85
77.23
Table 4: Pass@16 results (%) on out-of-domain benchmarks. Base model: Qwen2.5-VL-7B-Instruct.
Pruning Strategy
MathVerse
MathVision
MathVista
WeMath
Avg( ↑ )
Speed( ↑ )
Random-pruning 0.25
43.98
27.40
66.90
68.85
51.78
1.27 ×
Random-pruning 0.50
45.35
26.74
66.60
67.70
51.59
1.60 ×
Random-pruning 0.75
42.79
26.45
64.60
66.49
50.08
2.10 ×
Advantage-only (Ours)
44.54
27.57
67.10
69.14
52.09
2.35 ×
FastRL (Ours)
45.76
27.96
67.30
69.48
52.63
2.14 ×
Table 5: Ablation (%) of pruning strategies on Geometry3K using Qwen2.5-VL-7B-Instruct.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Multimodal
Unimodal
Dataset
Sum
Train
Test
Modality
Dataset
Sum
Train
Test
Modality
Geometry3K
2,702
2,101
601
Text + Image
MATH
12,500
11,250
1,250
Text
GeoQA8K-R1V
8,785
8,031
754
Text + Image
AIME2023
30
-
30
Text
MathVerse
3,940
-
3,940
Text + Image
AIME2024
30
-
30
Text
MathVision
3,040
-
3,040
Text + Image
AIME2025
30
-
30
Text
MathVista (mini)
1,000
-
1,000
Text + Image
AIME2026
30
-
30
Text
Appendix
Table 6: Statistics of the datasets used in our experiments.
Figure 4: Examples from the three training datasets: Geometry3K, GeoQA8k-R1V, and MATH.
Figure 5: Correlation analysis between token-level Jaccard similarity and gradient cosine similarity of trajectories with the same advantage for each question. The left plot corresponds to the Geometry3K dataset, while the right plot corresponds to the GeoQA8k-R1V dataset. The high Pearson and Spearman correlation coefficients demonstrate a strong positive correlation between the two metrics across both datasets.
Figure 6: Training dynamics and computational overhead analysis of Qwen2.5-VL-7B-Instruct on Geometry3K and GeoQA8K-R1V. We compare FastRL (ours) with GRPO, DAPO, CPPO, GSPO, and GRESO across three key metrics: total token consumption (left), policy entropy (middle), and accuracy (right).
National University of Defense Technology, Changsha, Hunan, China · Intelligent Game and Decision Lab, Beijing, China · The Chinese University of Hong Kong, Shenzhen, Guangdong, China