Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions
Organizations: École de technologie supérieure, Univ. of Québec, Canada · Huawei Technologies, Canada · The Univ. of Melbourne, Australia · Huawei Technologies, China
Abstract
Reasoning-oriented reinforcement learning enables large language models to solve mathematical, coding, and other multi-step tasks, but shifts a substantial portion of the training cost to rollout, where trajectories are generated for policy updates. Efficient rollout mechanisms are therefore essential to reduce this cost while maintaining the freshness, consistency, and statistical validity of training data. This survey provides a systematic taxonomy of recent research on rollout efficiency for reasoning-oriented reinforcement learning, classifying existing approaches from both mechanism and bottleneck perspectives. Based on this taxonomy, we analyze how different technique families address distinct sources of rollout inefficiency, examine opportunities and potential conflicts for combining them, identify gaps in the evaluation and reporting of efficiency gains, and discuss open challenges and future research directions.
Figures & tables
| Survey | Year | Scope | Rollout-efficiency coverage | Empirical analysis | ||||
|---|---|---|---|---|---|---|---|---|
| System lever | Algorithmic lever | Bottleneck mapping | Reported gains | Compos- ability | Metric standardization | |||
| Bai et al. ( bai2401beyond ) | 2024 | Training (+inference) | ||||||
| Duan et al. ( Duan et al., 2024 ) | 2024 | Training | ||||||
| Tie et al. ( Tie et al., 2025 ) | 2025 | Training | ||||||
| Kumar et al. ( Kumar et al., 2025 ) | 2025 | Training | ||||||
| Qu et al. ( Qu et al., 2025 ) | 2025 | Inference | ||||||
| Method | System gain | Algorithmic gain | Scale | Task | Hardware | Q | Cost / requirement |
| A.1 Pipeline Decoupling (Problem A: idle resources from stage coupling — 2 of 11 shown) | |||||||
| AReaL ( Fu et al., 2026 ) | tput | — | 1.5B–32B | math, code | — | / | stale trajectories within lag bound |
| LlamaRL ( Wu et al., 2025b ) | step (405B) | — | 8B–405B | math | 1,024 H100 | 1-step policy lag; AIPO correction | |
| A.2 Resource-Aware Execution (Problem A: idle, mismatched, or over-provisioned hardware — 2 of 16 shown) | |||||||
| RLBoost ( Wu et al., 2025a ) | tput – | — | — | — | preempt. GPUs | — | needs preemptible capacity; interruption recovery |
| AReaL-DTE ( Peng et al., 2026 ) | weight sync – ( cross-cluster) | — | 8B; 30B-A3B | math, code, logic | 8–16 H200 | ✓ | weights change/step; AdamW inversion; periodic full anchors |
| Decoup- ling | Resource- aware | Sched- uling | Partial rollout | Specu- lative | Rollout selection | |
| Resource-aware | — | |||||
| Scheduling | — | |||||
| Partial rollout | — | |||||
| Speculative | — | |||||
| Rollout selection | — | |||||
| Prompt filtering |
| Mechanism family | Methods | System gain | Algorithmic gain |
|---|---|---|---|
| A.1 Pipeline decoupling | 17 | (15/17) | (2/17) |
| A.2 Resource-aware execution | 15 | ||
| B.1 Scheduling and load balancing | 7 | (1/7) | |
| B.2 Partial and early-stop rollout | 7 | (5/7) | |
| B.3 Speculative decoding | 12 | ||
| C. 1 Rollout selection | 5 | (1/5) | (4/5) |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Property | Deployment inference | RL rollout generation |
|---|---|---|
| Request arrival | Requests arrive online according to an externally determined and potentially bursty workload. Future load is generally not known exactly. | The trainer explicitly launches a rollout workload, typically a batch of tasks available at the beginning of the rollout phase. |
| Workload control | The serving system controls scheduling after requests arrive, but generally cannot choose which requests arrive or reorder them arbitrarily across long time scales. | The trainer controls task sampling, batching, ordering, and often the number of trajectories generated for each task. |
| Request recurrence | Requests from different users are generally treated as unrelated, and the system cannot assume that a particular request will appear again. | Tasks are sampled repeatedly from a training distribution or finite dataset across optimization iterations or epochs, allowing information from previous rollouts to guide later scheduling and allocation decisions. |
| Primary objective | Per-request latency, throughput, and service-level objectives are central; requests should generally begin and complete promptly. | Individual request latency is usually secondary. In synchronous training, the primary systems objective is reducing the makespan of the complete rollout workload that blocks the next policy update. |
| Scheduling freedom | Deliberately delaying an admitted request typically worsens user-visible latency or fairness. | Tasks may be reordered, delayed, assigned different resources, truncated, or occasionally dropped if doing so reduces overall rollout cost without harming the learning objective. |
| Workload predictability | Prompt lengths, output lengths, and future arrivals are only partially known before execution. | The batch composition is known before generation, and previous executions of the same or similar tasks can provide estimates of trajectory length, reward, or computational cost. |
| Framework | Control | Placement | Main idea | Reported gain |
| HybridFlow / veRL ( Sheng et al., 2025 ) | One controller between nodes, local controllers inside each node | Shared devices, hybrid engine | 3D-HybridEngine: move the actor between training and generation layouts without a second copy of the weights | – throughput (range across configurations) |
| DISTFLOW ( Wang et al., 2025a ) | Local controllers only; no central node | Either | Planner plus local data coordinator (cache and double buffer) | Up to ; near-linear to 512 GPUs |
| StreamRL ( Zhong et al., 2025 ) | One controller over two services | Separate devices; can span datacenters | Streaming generation; separates pipeline idle time from length idle time | Up to (best configuration); cost-effectiveness |
| AsyncFlow ( Han et al., 2025 ) | Service interfaces; engines can be swapped | Separate devices | TransferQueue: hand out samples as soon as they are ready | avg., up to (256 NPUs) |
| Method | Draft source | Drafter cost | Guarantee |
|---|---|---|---|
| BubbleSpec ( Xu et al., 2026c ) | The policy itself, run in idle data-parallel gaps | No additional resources; uses idle GPU capacity | Exact match; fully synchronous |
| SPEC-RL ( Liu et al., 2025a ) | Previous step’s trajectory for the same prompt | Checking only ( s/step) | Suffix re-checked under the current policy |
| DAS ( Shao et al., 2026 ) | Suffix tree over past completions | No training; runs on CPU | Exact match (verified); identical training curves |
| RhymeRL ( He et al., 2025 ) | Suffix tree (HistoSpec) | No training; runs on CPU | RL procedure unchanged |
| WAR ( Xu et al., 2026d ) | Suffix tree, used only at low load | No training; turned off when batching fills the GPUs | Synchronous execution maintained |
| Seer ( Qin et al., 2026 ) | Tokens generated by sibling rollouts in the same group (grouped SD) | No model; online bookkeeping only | Synchronous and on-policy execution maintained |
| Method | Signal | Cost of the signal | Tracking policy change | Decision |
| GRESO ( Zheng et al., 2026b ) | Per-prompt reward history | Reuses training rewards; no extra model | Repeated re-estimation | Skip |
| MoPPS ( Qu et al., 2026 ) | Bayesian estimate of prompt difficulty | Lightweight posterior updates | Posterior updated during training | Prioritize / skip |
| HIVE ( Wu et al., 2026 ) | Informativeness estimates carried across the run | Bookkeeping only | Repeated re-estimation | Skip |
| AERO ( Zhang et al., 2026b ) | Statistics already collected in earlier rollouts | No separate predictor | Statistics refresh with the run | Skip |
| VCRL ( Jiang et al., 2025 ) | Group reward variance | Free from current rollouts | Replay revisits filtered prompts | Filter + replay |
| VIP ( Nguyen et al., 2026 ) | Gaussian-process prediction of success probability | Predictor fit plus convex allocation solve | Budget reduced rather than removed | Allocate |
| Method | System gain | Algorithmic gain | Scale | Task | Hardware | Q | Cost / requirement |
|---|---|---|---|---|---|---|---|
| A.1 Pipeline Decoupling (Problem A: idle resources from stage coupling) | |||||||
| AReaL ( Fu et al., 2026 ) | tput | — | 1.5B–32B | math, code | — | / | stale trajectories within lag bound |
| LlamaRL ( Wu et al., 2025b ) | step (405B) | — | 8B–405B | math | 1,024 H100 | 1-step policy lag; AIPO correction | |
| DORA ( Hu et al., 2026a ) | E2E ; rollout | — | 560B MoE | reasoning | 4,096 acc. | — | trajectories span multiple policy versions |
| Laminar ( Sheng et al., 2026 ) | train tput | — | — | — | 1,024 GPU | — | lag ; relay-worker weight service |
| FlexMARL ( Jiang et al., 2026b ) | speedup ; util. | — | 32B+14B | multi-agent | prod. cluster | async micro-batch coordination | |