Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at https://github.com/hexuandeng/GRPODropout/.
Figures & tables
Figure 1: GRPODropout procedure. Filter removes only selected positive-advantage rollouts; Re-center applies probability-weighted advantage recentering. Other GRPO updates are unchanged.
Method
Mathematical reasoning
General QA
Avg.
GSM8k
Math500
AMC23
Olympiad
AIME24
AIME25
HMMT25
ARC-C
MMLUP
SGPQA
Qwen3-1.7B
Base
89.92
80.60
59.92
42.96
22.08
21.04
11.46
87.37
57.01
29.28
50.16
GRPO
90.60
84.00
72.27
51.56
28.54
23.44
14.06
87.80
57.91
31.07
54.13
DAPO
88.86
85.20
72.27
52.00
33.75
26.15
15.21
87.63
55.62
29.70
54.64
Clip-Cov
82.34
77.20
53.91
36.15
11.56
12.71
5.21
87.37
55.19
28.26
44.99
KL-Cov
90.07
86.40
70.39
50.07
31.67
25.73
14.79
88.31
56.60
29.68
54.37
Table 1: Main results on three backbones. Avg. averages ten benchmarks; Base denotes the untrained backbone. Bold marks the best result within each backbone block.
Figure 2: Qwen3-4B actor entropy and rollout usage across checkpoints. (a–b) Token-mean entropy on training DAPO-17K and held-out MMLU-Pro (MMLUP); (c) mean update rollouts per question in training batches, excluding all-zero-advantage groups. The other five methods use 8 rollouts at every recorded checkpoint (gray dashed line). Curves connect raw observations without smoothing.
Deletion
Advantages
Mathematical reasoning
General QA
Avg.
Avg. Ent
Pos.
Neg.
Center
P-wt.
GSM8k
Math500
AMC23
Olympiad
AIME24
AIME25
HMMT25
ARC-C
MMLUP
SGPQA
Qwen3-1.7B
×
×
–
–
90.60
84.00
72.27
51.56
28.54
23.44
14.06
87.80
57.91
31.07
54.13
0.0503
✓
✓
✓
✓
90.30
86.80
70.47
49.04
30.94
24.90
13.75
87.71
56.53
29.77
54.02
0.2025
✓
×
×
×
77.33
39.00
10.70
8.74
0.00
0.10
0.00
87.12
44.04
21.09
28.81
0.3282
✓
×
✓
×
89.76
85.00
65.31
46.67
27.29
22.40
11.46
87.88
56.64
29.44
52.19
0.2239
Table 2: Qwen3-1.7B and Qwen3-4B ablations of positive- and negative-advantage rollouts, advantage recentering, and old-policy probability weighting of the mean. ✓ / × : enabled/disabled; –: not applicable. Avg. Ent reports benchmark-average actor entropy.
Method
Correctness
Path diversity
Entropy ↑
Overall mean rank ↓
pass@32 ↑
W≥8↓
W<8↓
Dist. ≥8 ↑
Dist. <8 ↑
Vendi ≥8 ↑
Vendi <8 ↑
GRPO
63.54
1.1846
7.6216
0.1063
0.1641
1.8311
2.4194
0.0448
5.63
DAPO
69.38
1.6154
8.0811
0.1150
0.1798
1.9080
2.6013
0.1796
3.88
Clip-Cov
76.88
0.7692
5.2162
0.1139
0.2141
1.9041
2.9851
0.0900
2.63
KL-Cov
72.71
0.4615
3.0000
0.1134
0.2014
1.8930
2.8510
0.1509
3.25
Beyond8020
73.54
1.1231
7.8378
0.1142
0.1726
1.9037
2.5252
0.1671
4.00
Table 3: Qwen3-4B solution coverage, wrong-answer dispersion, and reasoning diversity. Definitions are in Section 4.3 . Bold marks the best and underline marks the second-best in each column.
Reinforcement learning from verifiable rewards with GRPO is a standard approach for post-training reasoning LLMs. It remains sample inefficient. Each rollout is used for a single gradient update and then discarded. Naive replay is not well suited in this setting because LLM policies drift quickly per gradient step. Stored rollouts therefore become stale and can destabilize training. We propose a rollout-level replay buffer for GRPO that stores and samples individual rollouts rather than whole groups. The buffer bounds staleness through age eviction. Any rollout older than tau_max training steps is removed. The buffer also preserves on-policy data via fresh-anchored composition. Each batch keeps its fresh on-policy rollouts and then concatenates replay rollouts drawn separately from the buffer. We prioritize replay by per-rollout advantage magnitude and recycle individual rollouts whose advantages are large. Across three Qwen3-Base scales on five math benchmarks, our method outperforms GRPO and naive replay baselines. Gains are positive at every scale and reach +1.66 pp on the five-benchmark average at 4B. Under an AES metric that jointly measures accuracy and token efficiency, our method is the only condition with a positive margin over GRPO at every scale.
Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang +2
Department of Electrical and Computer Engineering, Seoul National University · Interdisciplinary Program in AI, Seoul National University · AIIS, ASRI, INMC, and ISRC, Seoul National University
Group-relative RL training (GRPO) samples a small group of parallel rollouts for every training prompt and uses their within-group reward spread to compute per-trajectory advantages. In agentic environments each rollout is a long multi-turn dialogue with one LLM call per step, so this multi-sample multiplier dominates the total training cost. When every rollout of a prompt ends with the same reward, the group has zero reward variance and contributes no gradient, so the extra rollouts add no information; such groups are common in practice (typically around 40% of all groups), so the wasted-compute fraction is substantial rather than marginal. Existing methods filter such groups at the prompt level, either after their rollouts are paid for or before any rollout begins, but both decide without using information that becomes available during the rollout itself. We instead ask whether the in-group divergence between the partial trajectories at an intermediate step can already predict that the group will be zero-variance: when the parallel rollouts have already converged on the same action prefix, the group is on track to produce a single reward, and we can stop early. We propose a one-parameter gate that stops a group when the mean pairwise prefix edit distance between its partial action sequences falls below a threshold. On a 60-iteration on-policy GRPO run on ALFWorld with Qwen2.5-7B, averaged over four random seeds, the gated arm finishes 10.7% faster in wall-clock (bootstrap 95% CI excludes 0) and shifts held-out success rate on 50 unseen tasks by +2.5 pp, with the held-out gain tracing to a measurable reduction in zero-advantage gradient-batch dilution. Code is available at https://github.com/zhiyuanZhai20/selective-rollout.
Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful instantiations. However, GRPO relies on repeated generation of long chain-of-thought rollouts. Training time scales with the number of rollouts, a large fraction of which are uninformative. Thus, GRPO is computationally expensive and unstable. To mitigate this, existing approaches either generate a larger pool of rollouts and filter the most informative prompts, or leverage historical signals for filtering at later stages of training. These strategies offer modest performance gains, but slow down the overall process. To address this, we propose VarIance Guided Online Rollout allocation (VIGOR) which instead of allocating a fixed rollout budget per example, begins with a small number of rollouts for all examples in a batch and iteratively allocates additional rollouts to those with the highest group reward variance until a fixed total rollout budget is reached. Theoretically, we show that under RLVR, reward variance controls the gradient magnitude, and derive VIGOR's closed-form speedup ratio over GRPO, which grows with refinement rounds under Pareto-distributed reward variance. Experiments on mathematical reasoning and coding tasks show that VIGOR reaches target accuracy with up to 2.3× fewer rollouts on math, reaches GRPO's final coding full pass rate with 1.49× fewer rollouts, and improves the coding average test pass rate by 3.4 points.