Organizations: Shanghai Jiao Tong University · Central South University · SenseTime Research · The Chinese University of Hong Kong · Beihang University
Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models' accuracy while preserving the efficiency gains of partial rollout. Experiments on mathematical and tool-assisted mathematical reasoning show that RollVerify achieves accuracy comparable to on-policy training while reducing training cost. Additional code-generation results provide preliminary evidence beyond mathematics.
Figures & tables
Figure 1: (a) Long-tail response length distributions observed during RL rollout under 16K and 32K settings. (b) Comparison between the standard on-policy and partial rollout in terms of speed and performance.
Setting
Filter
Threshold
AIME 24 ↑
AIME 25 ↑
OPS
GPU days
GRPO baseline (on-policy)
None
–
29.4
27.4
0
30.1 (1.00x)
Partial Rollout (off-policy)
None
–
19.5
19.1
0.025
19.4 ( 1.55x )
Staleness
1
21.2
20.9
0.017
26.6 ( 1.13x )
Staleness
2
19.7
19.3
0.018
21.5 ( 1.40x )
Staleness
3
18.7
18.4
0.018
19.8 ( 1.52x )
OPS
0.02
21.1
20.3
0.016
27.1 ( 1.11x )
Table 1: Analysis results of staleness-based and OPS-based filtering strategies under partial rollout on DAPO-Math. The baseline setting uses Qwen3-4B-Base, with a batch size of 128.
Figure 2: Training dynamics under different partial rollout settings using Qwen3-4B-Base model. We train the model for 500 steps in total, with the x-axis denoting the training step.
Figure 3: Overview of RollVerify. It extends the standard rollout–training loop with an additional verification phase. left: rollout phase; mid: training phase; right: proposed verification phase.
Figure 5
Model
Setting
AIME24
AIME25
AMC23
MATH500
AVG
OPS
GPU days
Qwen3-8B-Base
On-policy
40.3
29.5
81.4
78.0
57.3
0
98.1(1.0x)
Partial
23.7
19.4
67.5
73.2
45.9 ↓
0.017
51.6( 1.9x )
RollVerify
39.8
30.5
81.5
77.2
57.2
0.004
57.5( 1.7x )
Qwen3-30B-A3B-Base
On-policy
52.3
37.8
89.7
70.3
62.5
0
112 (1.0x)
Partial
36.8
28.5
77.8
69.4
53.1 ↓
0.016
56.1( 2.0x )
RollVerify
52.0
38.0
89.8
70.5
62.5
0.006
65.1( 1.7x )
Table 2: Comparison among on-policy training, partial rollout, and proposed RollVerify.
Figure 4: Comparison of training dynamics among On-Policy baseline, Partial, and RollVerify. The first row shows results on Qwen3-8B-Base, and the second row presents results on the MoE model Qwen3-30B-A3B-Base.
Setting
AIME24
AIME25
AMC23
MATH500
AVG
OPS
AC Rate
GPU Days
Naive partial
23.7
19.4
67.5
73.2
45.9
0.017
1.0
51.6
+seq-level verify
32.6
26.5
78.5
78.1
53.9
0.009
0.67
68.1
+token-level verify
39.6
30.8
81.8
77.2
57.4
0.005
0.83
61.5
+switch
39.8
30.5
81.5
77.2
57.2
0.004
0.83
57.5
Table 3: Ablation study of RollVerify under Qwen3-8B-Base setting. We progressively add sequence-level verification, token-level verification, and conditional switching on top of the naive partial rollout baseline. AC Rate indicates acceptance rate.
Figure 5: Ablation Results. (a) Effect of Two-stage verification with respect to entropy and OPS. (b) Analysis of conditional switching on generation time and acceptance rate. The dashed vertical line indicates the switching point from partial rollout to fully on-policy training.
Table 10
Setting
Training?
Rollout?
AIME24
AIME25
AMC23
MATH500
AVG
OPS
GPU Days
GRPO Shao et al. (2024)
40.3
29.5
81.4
78.0
57.3
0
98.1
Partial Team et al. (2025)
23.7
19.4
67.5
73.2
45.9
0.017
51.6
Staleness-1
✓
28.2
22.8
71.9
75.6
49.6
0.015
64.5
GSPO Zheng et al. (2025)
✓
37.3
27.0
79.8
76.5
55.2
0.007
51.8
SAPO Gao et al. (2025a)
✓
37.5
27.7
79.6
75.7
55.1
0.008
51.5
VESPO Shen et al. (2026)
✓
31.0
26.2
74.3
74.5
51.5
0.010
52.0
Table 6: Comparison with other off-policy handling methods on Qwen3-8B-Base. GRPO is a fully on-policy baseline, and other methods use partial rollout. Training? indicates that methods operate on loss level in training phase. Rollout? indicates that methods operate on rollout samples.
Max Length
Setting
AIME24
AIME25
AMC23
MATH500
AVG
GPU Days
8K
on-policy
29.7
25.7
75.9
75.5
51.7
16.6 (1.0x)
RollVerify
30.2
25.4
75.7
75.8
51.8
15.1 (1.1x)
16K
on-policy
36.1
28.8
80.6
72.6
54.5
51.8 (1.0x)
RollVerify
36.3
28.7
80.4
73.0
54.6
37.1 (1.4x)
32K
on-policy
40.3
29.5
81.4
78.0
57.3
98.1 (1.0x)
RollVerify
39.8
30.5
81.5
77.2
57.2
57.5 (1.7x)
Table 7: Performance comparison under different maximum lengths.
Model
Setting
Verification overhead
Step Time
OPS compute
Two-stage Verify
Qwen3-8B-Base
RollVerify
20s ( 6.4% )
2s ( 0.6% )
310s
Qwen3-30B-A3B-Base
RollVerify
25s ( 6.9% )
2s ( 0.5% )
362s
Table 8: Overhead under different settings. Step time indicates the average training step time.
Setting
Reward
Entropy
AC Rate
Lengths
OPS
AIME24/25
GPU Days
on-policy
0.665
0.113
–
7118
0
38.7 / 32.8
81.3 (1.00 × )
RollVerify w/o switching
0.659
0.122
0.740
7109
0.005
38.7 / 32.5
64.3 (1.26 × )
RollVerify w/ switching
0.662
0.121
0.826
7201
0.005
38.8 / 32.8
60.2 (1.35 × )
Table 9: Effect of conditional switching under the 16K setting with 700 training steps. RollVerify with switching switches to on-policy rollout at step 598. AC Rate denotes the average acceptance rate over training. Lengths mean response lengths. Speedups in parentheses are relative to on-policy training.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Setting
AIME24
AIME25
AMC23
MATH500
AVG
GPU Days
DeepScaleR
on-policy
38.9
28.4
78.9
89.8
59.0
91.2 (1.0x)
partial
25.1
20.7
67.9
84.4
49.5
47.2 (1.9x)
RollVerify
38.8
28.5
79.4
90.1
59.2
52.1 (1.8x)
ReTool
on-policy
39.4
32.0
74.0
68.9
53.6
102 (1.0x)
partial
28.9
20.0
69.0
60.8
44.7
50.5 (2.0x)
RollVerify
39.9
31.7
73.9
68.5
53.5
63.5 (1.6x)
Appendix
Table 10: Performance comparison across datasets and settings.
Setting
LiveCodeBench
HumanEval
AVG
GPU Days
on-policy
28.1
85.2
56.6
37.0 (1.0x)
partial
24.7
80.3
52.5
24.7 (1.5x)
RollVerify
27.9
85.1
56.5
26.4 (1.4x)
Appendix
Table 11: Performance comparison on coding benchmarks.
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China · University of Chinese Academy of Sciences, Beijing, China · Southern University of Science and Technology, Shenzhen, China