Organizations: École de technologie supérieure, Univ. of Québec, Canada · Huawei Technologies, Canada · The Univ. of Melbourne, Australia · Huawei Technologies, China
Reasoning-oriented reinforcement learning enables large language models to solve mathematical, coding, and other multi-step tasks, but shifts a substantial portion of the training cost to rollout, where trajectories are generated for policy updates. Efficient rollout mechanisms are therefore essential to reduce this cost while maintaining the freshness, consistency, and statistical validity of training data. This survey provides a systematic taxonomy of recent research on rollout efficiency for reasoning-oriented reinforcement learning, classifying existing approaches from both mechanism and bottleneck perspectives. Based on this taxonomy, we analyze how different technique families address distinct sources of rollout inefficiency, examine opportunities and potential conflicts for combining them, identify gaps in the evaluation and reporting of efficiency gains, and discuss open challenges and future research directions.
Figures & tables
Figure 1. Phase-time breakdown of a synchronous RL training step.
Survey
Year
Scope
Rollout-efficiency coverage
Empirical analysis
System lever
Algorithmic lever
Bottleneck mapping
Reported gains
Compos- ability
Metric standardization
Bai et al. ( bai2401beyond )
2024
Training (+inference)
Duan et al. ( Duan et al., 2024 )
2024
Training
Tie et al. ( Tie et al., 2025 )
2025
Training
Kumar et al. ( Kumar et al., 2025 )
2025
Training
Qu et al. ( Qu et al., 2025 )
2025
Inference
Table 1. Comparison of representative surveys and this work. Scope states each survey’s primary subject; parentheses mark a secondary subject treated in passing. full coverage; partial coverage; limited or no coverage.
Figure 2. Survey methodology. The taxonomy is constructed through four stages: search, screen, classify, and group.
Figure 3. Conceptual workflow of reasoning RL.
Method
System gain
Algorithmic gain
Scale
Task
Hardware
Q
Cost / requirement
A.1 Pipeline Decoupling (Problem A: idle resources from stage coupling — 2 of 11 shown)
AReaL ( Fu et al., 2026 )
tput 2.77×
—
1.5B–32B
math, code
—
= / +
stale trajectories within lag bound
LlamaRL ( Wu et al., 2025b )
step 10.7× (405B)
—
8B–405B
math
1,024 H100
=
1-step policy lag; AIPO correction
A.2 Resource-Aware Execution (Problem A: idle, mismatched, or over-provisioned hardware — 2 of 16 shown)
RLBoost ( Wu et al., 2025a )
tput 1.51 – 1.97×
—
—
—
preempt. GPUs
—
needs preemptible capacity; interruption recovery
AReaL-DTE ( Peng et al., 2026 )
weight sync 6.8 – 7.6× ( 19.9× cross-cluster)
—
8B; 30B-A3B
math, code, logic
8–16 H200
✓
<2% weights change/step; AdamW inversion; periodic full anchors
Table 2. Representative methods per family (2 of each; full 80-method comparison in Table 9 , Appendix F ). Q: = preserved, + improved, ≈ comparable, ✓ exact/lossless; — not reported.
Figure 4. Mechanism taxonomy of rollout-efficiency methods.
Figure 5. Bottleneck taxonomy of rollout-efficiency methods.
Decoup- ling
Resource- aware
Sched- uling
Partial rollout
Specu- lative
Rollout selection
Resource-aware
(×/∙m/∘)†
—
∙
∙
(∙/∘)†
∙
Scheduling
×s
∙
—
×
×m
∙
Partial rollout
∘s
∙
×
—
×
∘
Speculative
×
(∙/∘)†
×m
×
—
∙
Rollout selection
∘
∙
∙
∘
∙
—
Prompt filtering
∙
∙
∘
∙s
×
×
Table 3. Pairwise composability of technique families, with evidence status.
Mechanism family
Methods
System gain
Algorithmic gain
A.1 Pipeline decoupling
17
(15/17)
(2/17)
A.2 Resource-aware execution
15
B.1 Scheduling and load balancing
7
(1/7)
B.2 Partial and early-stop rollout
7
(5/7)
B.3 Speculative decoding
12
C. 1 Rollout selection
5
(1/5)
(4/5)
Table 4. Which efficiency lever each mechanism family reports , from Table 9 . every method; some (fraction shown); none. I-PPO ( Shu et al., 2026 ) reports an unquantified claim and is counted on neither lever.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Property
Deployment inference
RL rollout generation
Request arrival
Requests arrive online according to an externally determined and potentially bursty workload. Future load is generally not known exactly.
The trainer explicitly launches a rollout workload, typically a batch of B tasks available at the beginning of the rollout phase.
Workload control
The serving system controls scheduling after requests arrive, but generally cannot choose which requests arrive or reorder them arbitrarily across long time scales.
The trainer controls task sampling, batching, ordering, and often the number of trajectories generated for each task.
Request recurrence
Requests from different users are generally treated as unrelated, and the system cannot assume that a particular request will appear again.
Tasks are sampled repeatedly from a training distribution or finite dataset across optimization iterations or epochs, allowing information from previous rollouts to guide later scheduling and allocation decisions.
Primary objective
Per-request latency, throughput, and service-level objectives are central; requests should generally begin and complete promptly.
Individual request latency is usually secondary. In synchronous training, the primary systems objective is reducing the makespan of the complete rollout workload that blocks the next policy update.
Scheduling freedom
Deliberately delaying an admitted request typically worsens user-visible latency or fairness.
Tasks may be reordered, delayed, assigned different resources, truncated, or occasionally dropped if doing so reduces overall rollout cost without harming the learning objective.
Workload predictability
Prompt lengths, output lengths, and future arrivals are only partially known before execution.
The batch composition is known before generation, and previous executions of the same or similar tasks can provide estimates of trajectory length, reward, or computational cost.
Appendix
Table 5. Deployment inference versus RL rollout generation.
Framework
Control
Placement
Main idea
Reported gain
HybridFlow / veRL ( Sheng et al., 2025 )
One controller between nodes, local controllers inside each node
Shared devices, hybrid engine
3D-HybridEngine: move the actor between training and generation layouts without a second copy of the weights
1.53 – 20.57× throughput (range across configurations)
DISTFLOW ( Wang et al., 2025a )
Local controllers only; no central node
Either
Planner plus local data coordinator (cache and double buffer)
Up to 2.63× ; near-linear to 512 GPUs
StreamRL ( Zhong et al., 2025 )
One controller over two services
Separate devices; can span datacenters
Streaming generation; separates pipeline idle time from length idle time
Up to 2.66× (best configuration); 1.33× cost-effectiveness
AsyncFlow ( Han et al., 2025 )
Service interfaces; engines can be swapped
Separate devices
TransferQueue: hand out samples as soon as they are ready
1.59× avg., up to 2.03× (256 NPUs)
Appendix
Table 6. Training frameworks for RL post-training.
Method
Draft source
Drafter cost
Guarantee
BubbleSpec ( Xu et al., 2026c )
The policy itself, run in idle data-parallel gaps
No additional resources; uses idle GPU capacity
Exact match; fully synchronous
SPEC-RL ( Liu et al., 2025a )
Previous step’s trajectory for the same prompt
Checking only ( ≈20 s/step)
Suffix re-checked under the current policy
DAS ( Shao et al., 2026 )
Suffix tree over past completions
No training; runs on CPU
Exact match (verified); identical training curves
RhymeRL ( He et al., 2025 )
Suffix tree (HistoSpec)
No training; runs on CPU
RL procedure unchanged
WAR ( Xu et al., 2026d )
Suffix tree, used only at low load
No training; turned off when batching fills the GPUs
Synchronous execution maintained
Seer ( Qin et al., 2026 )
Tokens generated by sibling rollouts in the same group (grouped SD)
No model; online bookkeeping only
Synchronous and on-policy execution maintained
Appendix
Table 7. Speculative decoding for RL rollout.
Method
Signal
Cost of the signal
Tracking policy change
Decision
GRESO ( Zheng et al., 2026b )
Per-prompt reward history
Reuses training rewards; no extra model
Repeated re-estimation
Skip
MoPPS ( Qu et al., 2026 )
Bayesian estimate of prompt difficulty
Lightweight posterior updates
Posterior updated during training
Prioritize / skip
HIVE ( Wu et al., 2026 )
Informativeness estimates carried across the run
Bookkeeping only
Repeated re-estimation
Skip
AERO ( Zhang et al., 2026b )
Statistics already collected in earlier rollouts
No separate predictor
Statistics refresh with the run
Skip
VCRL ( Jiang et al., 2025 )
Group reward variance
Free from current rollouts
Replay revisits filtered prompts
Filter + replay
VIP ( Nguyen et al., 2026 )
Gaussian-process prediction of success probability
Predictor fit plus convex allocation solve
Budget reduced rather than removed
Allocate
Appendix
Table 8. Prompt filtering and selection methods.
Method
System gain
Algorithmic gain
Scale
Task
Hardware
Q
Cost / requirement
A.1 Pipeline Decoupling (Problem A: idle resources from stage coupling)
Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful instantiations. However, GRPO relies on repeated generation of long chain-of-thought rollouts. Training time scales with the number of rollouts, a large fraction of which are uninformative. Thus, GRPO is computationally expensive and unstable. To mitigate this, existing approaches either generate a larger pool of rollouts and filter the most informative prompts, or leverage historical signals for filtering at later stages of training. These strategies offer modest performance gains, but slow down the overall process. To address this, we propose VarIance Guided Online Rollout allocation (VIGOR) which instead of allocating a fixed rollout budget per example, begins with a small number of rollouts for all examples in a batch and iteratively allocates additional rollouts to those with the highest group reward variance until a fixed total rollout budget is reached. Theoretically, we show that under RLVR, reward variance controls the gradient magnitude, and derive VIGOR's closed-form speedup ratio over GRPO, which grows with refinement rounds under Pareto-distributed reward variance. Experiments on mathematical reasoning and coding tasks show that VIGOR reaches target accuracy with up to 2.3× fewer rollouts on math, reaches GRPO's final coding full pass rate with 1.49× fewer rollouts, and improves the coding average test pass rate by 3.4 points.
Existing reinforcement learning methods for LLM reasoning implicitly assume that the policy generating training trajectories should coincide with the one producing inference responses. We argue that this is a misleading inductive bias: the optimization-optimal trajectory distribution favors informative gradients, whereas the inference-optimal response distribution emphasizes accuracy and consistency. Forcing both into a single policy entangles their gradients and suppresses exploration. We propose R2PO (Residual Rollout Policy Optimization), which attaches a lightweight Residual Rollout-Head atop the policy to decouple training trajectories from inference responses, diversifying rollouts during training while keeping inference generation intact. Experiments show that R2PO consistently outperforms baselines, with average accuracy gains of 3.4% on MATH-500 and 1.3% on APPS, alongside more diverse rollouts and reduced length bias. Our code is available at https://github.com/RRPO-ARR/Code.
Jingchu Wang, Bingbing Xu, Yige Yuan +4
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS · University of Chinese Academy of Sciences · National University of Singapore
Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine-tuning (ReFT). In standard ReFT frameworks, a behavior model generates multiple completions with answers per problem, for the answer to be then scored by a reward function. While such RL post-training methods demonstrate significant performance improvements across challenging reasoning domains, the computational cost of generating completions during training with multiple inference steps makes the training cost non-trivial. To address this, we draw inspiration from off-policy RL, and speculative decoding to introduce a novel ReFT framework, dubbed Nested-ReFT, where a subset of layers of the target model acts as the behavior model to generate off-policy completions during training. The behavior model configured with dynamic layer skipping per batch during training decreases the inference cost compared to the standard ReFT frameworks. Our theoretical analysis shows that Nested-ReFT yields unbiased gradient estimates with controlled variance. Our empirical analysis demonstrates improved computational efficiency measured as tokens/sec across multiple math reasoning benchmarks and model sizes. Additionally, we explore three variants of bias mitigation to minimize the off-policyness in the gradient updates that allows for maintaining performance that matches the baseline ReFT performance.
Maxime Heuillet, Yufei Cui, Boxing Chen +2
Universit´e Laval (IID), Canada · Mila - Qu´ebec AI Institute, Canada · Huawei Noah’s Ark Lab (Montreal Research Center), Canada +1