Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at https://github.com/hexuandeng/GRPODropout/.
Figures & tables
Figure 1: GRPODropout procedure. Filter removes only selected positive-advantage rollouts; Re-center applies probability-weighted advantage recentering. Other GRPO updates are unchanged.
Method
Mathematical reasoning
General QA
Avg.
GSM8k
Math500
AMC23
Olympiad
AIME24
AIME25
HMMT25
ARC-C
MMLUP
SGPQA
Qwen3-1.7B
Base
89.92
80.60
59.92
42.96
22.08
21.04
11.46
87.37
57.01
29.28
50.16
GRPO
90.60
84.00
72.27
51.56
28.54
23.44
14.06
87.80
57.91
31.07
54.13
DAPO
88.86
85.20
72.27
52.00
33.75
26.15
15.21
87.63
55.62
29.70
54.64
Clip-Cov
82.34
77.20
53.91
36.15
11.56
12.71
5.21
87.37
55.19
28.26
44.99
KL-Cov
90.07
86.40
70.39
50.07
31.67
25.73
14.79
88.31
56.60
29.68
54.37
Table 1: Main results on three backbones. Avg. averages ten benchmarks; Base denotes the untrained backbone. Bold marks the best result within each backbone block.
Figure 2: Qwen3-4B actor entropy and rollout usage across checkpoints. (a–b) Token-mean entropy on training DAPO-17K and held-out MMLU-Pro (MMLUP); (c) mean update rollouts per question in training batches, excluding all-zero-advantage groups. The other five methods use 8 rollouts at every recorded checkpoint (gray dashed line). Curves connect raw observations without smoothing.
Deletion
Advantages
Mathematical reasoning
General QA
Avg.
Avg. Ent
Pos.
Neg.
Center
P-wt.
GSM8k
Math500
AMC23
Olympiad
AIME24
AIME25
HMMT25
ARC-C
MMLUP
SGPQA
Qwen3-1.7B
×
×
–
–
90.60
84.00
72.27
51.56
28.54
23.44
14.06
87.80
57.91
31.07
54.13
0.0503
✓
✓
✓
✓
90.30
86.80
70.47
49.04
30.94
24.90
13.75
87.71
56.53
29.77
54.02
0.2025
✓
×
×
×
77.33
39.00
10.70
8.74
0.00
0.10
0.00
87.12
44.04
21.09
28.81
0.3282
✓
×
✓
×
89.76
85.00
65.31
46.67
27.29
22.40
11.46
87.88
56.64
29.44
52.19
0.2239
Table 2: Qwen3-1.7B and Qwen3-4B ablations of positive- and negative-advantage rollouts, advantage recentering, and old-policy probability weighting of the mean. ✓ / × : enabled/disabled; –: not applicable. Avg. Ent reports benchmark-average actor entropy.
Method
Correctness
Path diversity
Entropy ↑
Overall mean rank ↓
pass@32 ↑
W≥8↓
W<8↓
Dist. ≥8 ↑
Dist. <8 ↑
Vendi ≥8 ↑
Vendi <8 ↑
GRPO
63.54
1.1846
7.6216
0.1063
0.1641
1.8311
2.4194
0.0448
5.63
DAPO
69.38
1.6154
8.0811
0.1150
0.1798
1.9080
2.6013
0.1796
3.88
Clip-Cov
76.88
0.7692
5.2162
0.1139
0.2141
1.9041
2.9851
0.0900
2.63
KL-Cov
72.71
0.4615
3.0000
0.1134
0.2014
1.8930
2.8510
0.1509
3.25
Beyond8020
73.54
1.1231
7.8378
0.1142
0.1726
1.9037
2.5252
0.1671
4.00
Table 3: Qwen3-4B solution coverage, wrong-answer dispersion, and reasoning diversity. Definitions are in Section 4.3 . Bold marks the best and underline marks the second-best in each column.
Department of Electrical and Computer Engineering, Seoul National University · Interdisciplinary Program in AI, Seoul National University · AIIS, ASRI, INMC, and ISRC, Seoul National University