RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning
Organizations: Shanghai Jiao Tong University · Central South University · SenseTime Research · The Chinese University of Hong Kong · Beihang University
Abstract
Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models' accuracy while preserving the efficiency gains of partial rollout. Experiments on mathematical and tool-assisted mathematical reasoning show that RollVerify achieves accuracy comparable to on-policy training while reducing training cost. Additional code-generation results provide preliminary evidence beyond mathematics.
Figures & tables
| Setting | Filter | Threshold | AIME 24 | AIME 25 | OPS | GPU days |
|---|---|---|---|---|---|---|
| GRPO baseline (on-policy) | None | – | 29.4 | 27.4 | 0 | 30.1 (1.00x) |
| Partial Rollout (off-policy) | None | – | 19.5 | 19.1 | 0.025 | 19.4 ( 1.55x ) |
| Staleness | 1 | 21.2 | 20.9 | 0.017 | 26.6 ( 1.13x ) | |
| Staleness | 2 | 19.7 | 19.3 | 0.018 | 21.5 ( 1.40x ) | |
| Staleness | 3 | 18.7 | 18.4 | 0.018 | 19.8 ( 1.52x ) | |
| OPS | 0.02 | 21.1 | 20.3 | 0.016 | 27.1 ( 1.11x ) |
| Model | Setting | AIME24 | AIME25 | AMC23 | MATH500 | AVG | OPS | GPU days |
|---|---|---|---|---|---|---|---|---|
| Qwen3-8B-Base | On-policy | 40.3 | 29.5 | 81.4 | 78.0 | 57.3 | 0 | 98.1(1.0x) |
| Partial | 23.7 | 19.4 | 67.5 | 73.2 | 45.9 | 0.017 | 51.6( 1.9x ) | |
| RollVerify | 39.8 | 30.5 | 81.5 | 77.2 | 57.2 | 0.004 | 57.5( 1.7x ) | |
| Qwen3-30B-A3B-Base | On-policy | 52.3 | 37.8 | 89.7 | 70.3 | 62.5 | 0 | 112 (1.0x) |
| Partial | 36.8 | 28.5 | 77.8 | 69.4 | 53.1 | 0.016 | 56.1( 2.0x ) | |
| RollVerify | 52.0 | 38.0 | 89.8 | 70.5 | 62.5 | 0.006 | 65.1( 1.7x ) |
| Setting | AIME24 | AIME25 | AMC23 | MATH500 | AVG | OPS | AC Rate | GPU Days |
|---|---|---|---|---|---|---|---|---|
| Naive partial | 23.7 | 19.4 | 67.5 | 73.2 | 45.9 | 0.017 | 1.0 | 51.6 |
| +seq-level verify | 32.6 | 26.5 | 78.5 | 78.1 | 53.9 | 0.009 | 0.67 | 68.1 |
| +token-level verify | 39.6 | 30.8 | 81.8 | 77.2 | 57.4 | 0.005 | 0.83 | 61.5 |
| +switch | 39.8 | 30.5 | 81.5 | 77.2 | 57.2 | 0.004 | 0.83 | 57.5 |
| Setting | Training? | Rollout? | AIME24 | AIME25 | AMC23 | MATH500 | AVG | OPS | GPU Days |
| GRPO Shao et al. (2024) | 40.3 | 29.5 | 81.4 | 78.0 | 57.3 | 0 | 98.1 | ||
| Partial Team et al. (2025) | 23.7 | 19.4 | 67.5 | 73.2 | 45.9 | 0.017 | 51.6 | ||
| Staleness-1 | ✓ | 28.2 | 22.8 | 71.9 | 75.6 | 49.6 | 0.015 | 64.5 | |
| GSPO Zheng et al. (2025) | ✓ | 37.3 | 27.0 | 79.8 | 76.5 | 55.2 | 0.007 | 51.8 | |
| SAPO Gao et al. (2025a) | ✓ | 37.5 | 27.7 | 79.6 | 75.7 | 55.1 | 0.008 | 51.5 | |
| VESPO Shen et al. (2026) | ✓ | 31.0 | 26.2 | 74.3 | 74.5 | 51.5 | 0.010 | 52.0 |
| Max Length | Setting | AIME24 | AIME25 | AMC23 | MATH500 | AVG | GPU Days |
|---|---|---|---|---|---|---|---|
| 8K | on-policy | 29.7 | 25.7 | 75.9 | 75.5 | 51.7 | 16.6 (1.0x) |
| RollVerify | 30.2 | 25.4 | 75.7 | 75.8 | 51.8 | 15.1 (1.1x) | |
| 16K | on-policy | 36.1 | 28.8 | 80.6 | 72.6 | 54.5 | 51.8 (1.0x) |
| RollVerify | 36.3 | 28.7 | 80.4 | 73.0 | 54.6 | 37.1 (1.4x) | |
| 32K | on-policy | 40.3 | 29.5 | 81.4 | 78.0 | 57.3 | 98.1 (1.0x) |
| RollVerify | 39.8 | 30.5 | 81.5 | 77.2 | 57.2 | 57.5 (1.7x) |
| Model | Setting | Verification overhead | Step Time | |
|---|---|---|---|---|
| OPS compute | Two-stage Verify | |||
| Qwen3-8B-Base | RollVerify | 20s ( 6.4% ) | 2s ( 0.6% ) | 310s |
| Qwen3-30B-A3B-Base | RollVerify | 25s ( 6.9% ) | 2s ( 0.5% ) | 362s |
| Setting | Reward | Entropy | AC Rate | Lengths | OPS | AIME24/25 | GPU Days |
|---|---|---|---|---|---|---|---|
| on-policy | 0.665 | 0.113 | – | 7118 | 0 | 38.7 / 32.8 | 81.3 (1.00 ) |
| RollVerify w/o switching | 0.659 | 0.122 | 0.740 | 7109 | 0.005 | 38.7 / 32.5 | 64.3 (1.26 ) |
| RollVerify w/ switching | 0.662 | 0.121 | 0.826 | 7201 | 0.005 | 38.8 / 32.8 | 60.2 (1.35 ) |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Setting | AIME24 | AIME25 | AMC23 | MATH500 | AVG | GPU Days |
|---|---|---|---|---|---|---|---|
| DeepScaleR | on-policy | 38.9 | 28.4 | 78.9 | 89.8 | 59.0 | 91.2 (1.0x) |
| partial | 25.1 | 20.7 | 67.9 | 84.4 | 49.5 | 47.2 (1.9x) | |
| RollVerify | 38.8 | 28.5 | 79.4 | 90.1 | 59.2 | 52.1 (1.8x) | |
| ReTool | on-policy | 39.4 | 32.0 | 74.0 | 68.9 | 53.6 | 102 (1.0x) |
| partial | 28.9 | 20.0 | 69.0 | 60.8 | 44.7 | 50.5 (2.0x) | |
| RollVerify | 39.9 | 31.7 | 73.9 | 68.5 | 53.5 | 63.5 (1.6x) |
| Setting | LiveCodeBench | HumanEval | AVG | GPU Days |
|---|---|---|---|---|
| on-policy | 28.1 | 85.2 | 56.6 | 37.0 (1.0x) |
| partial | 24.7 | 80.3 | 52.5 | 24.7 (1.5x) |
| RollVerify | 27.9 | 85.1 | 56.5 | 26.4 (1.4x) |