cs.AISep 21, 2026

Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions

Authors: Niloofar Gholipour, Marcos Assuncao, Gursimran Singh, Timothy Yu, Rajkumar Buyya, Julien Gascon-Samson, Zhenan Fan, Yong Zhang, +3 more

Organizations: École de technologie supérieure, Univ. of Québec, Canada · Huawei Technologies, Canada · The Univ. of Melbourne, Australia · Huawei Technologies, China

Abstract

Reasoning-oriented reinforcement learning enables large language models to solve mathematical, coding, and other multi-step tasks, but shifts a substantial portion of the training cost to rollout, where trajectories are generated for policy updates. Efficient rollout mechanisms are therefore essential to reduce this cost while maintaining the freshness, consistency, and statistical validity of training data. This survey provides a systematic taxonomy of recent research on rollout efficiency for reasoning-oriented reinforcement learning, classifying existing approaches from both mechanism and bottleneck perspectives. Based on this taxonomy, we analyze how different technique families address distinct sources of rollout inefficiency, examine opportunities and potential conflicts for combining them, identify gaps in the evaluation and reporting of efficiency gains, and discuss open challenges and future research directions.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. R2^2PO: Decoupling Rollout and Inference Policies for LLM Reasoning

    Jan 17, 2026Jingchu Wang, Bingbing Xu, Yige Yuan +4LLM Reasoning StrategiesFrictive Policy Optimization

  2. Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts

    Aug 13, 2025Maxime Heuillet, Yufei Cui, Boxing Chen +2Mathematical Reasoning BenchmarksLarge Language Model Fine-Tuning