Parallel Rollouts

Momentum

3 papers in the last four weeks, level with the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 26

All topics
CardsList
  1. Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

    Sep 29, 2026Ziheng Huang, Yicheng Bao, Xueheng Li +8Streaming Video UnderstandingLong-Horizon Task Planning

  2. Sample What You Say: Aligning Language Models to Sample the Distributions They State

    Sep 28, 2026Kasra Arabi, Virginia Smith, Chhavi YadavDistributionsParallel Rollouts

  3. From Base Rollouts to RL Reasoning: A Budgeted Search Perspective

    Sep 1, 2026Wenhe Sun, Cunxiang Wang, Zijun Yao +1Offline Reinforcement LearningParallel Rollouts

  4. MemoryWalker: Stop Training Agents on Contexts They Never Saw

    Sep 1, 2026Zinco J, Xunjie Zhu, Shen Huang +3Autoregressive RolloutParallel Rollouts

  5. Scheduling Mixed RL Rollouts Beyond Prefix Locality

    Aug 11, 2026Zetao Hong, Song Yuan, Yuanhao Ding +4Parallel RolloutsAutoregressive Rollout

  6. Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks

    Aug 4, 2026Chenyu Wang, Yunbo Lyu, Junda He +4RetryingSwe-Bench Verified

  7. Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning

    Aug 2, 2026Wenhao Zhang, Yibo Xie, Rui Wang +9Autoregressive RolloutLarge Language Model Reinforcement Learning

  8. ADP: Adversarial Dynamics Priors for Physically Grounded Humanoid Locomotion

    Jul 3, 2026Seokju Lee, Jeongtae Lee, Jeonghyeok Lim +6Full-Body Humanoid ControlKinematic Priors

  9. Scaling Laws for Collapse in Asynchronous GRPO

    Jul 1, 2026Jingwei Song, Haofeng Xu, Jie Xiao +7Parallel RolloutsBetter Call Group Relative Policy Optimization

  10. Understanding Rollout Error in Graph World Models

    Jun 26, 2026Xinyuan Song, Zekun CaiWorld ModelsPrediction Error

  11. EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

    Jun 17, 2026Minseo Kim, Minjae Lee, Seunghyuk Oh +7Autoregressive RolloutParallel Rollouts

  12. CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward

    Jun 12, 2026Md Amirul Islam, Sumiran Thakur, Huancheng Chen +3Large Language Model AgentsAgentic Benchmarks

  13. Rollout-Level Advantage-Prioritized Experience Replay for GRPO

    Jun 3, 2026Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang +2Parallel RolloutsReplay

  14. Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR

    Jun 2, 2026Chuanyu Qin, Chenxu Yang, Qingyi Si +3Reinforcement Learning With Verifiable RewardVerifiable Rewards

  15. Libra: Efficient Resource Management for Agentic RL Post-Training

    Jun 2, 2026Kaiwen Chen, Xin Tan, Jingzong Li +6Agentic Reinforcement LearningResource-Efficient

  16. Are Full Rollouts Necessary for On-Policy Distillation?

    May 29, 2026Yaocheng Zhang, Jiajun Chai, Yuqian Fu +7Efficient On-Policy DistillationParallel Rollouts

  17. Policy Library CBF: Finite-Horizon Safety at Runtime via Parallel Rollouts

    May 15, 2026Taekyung Kim, Hideki Okamoto, Bardh Hoxha +2Safety FiltersControl Barrier Functions

  18. Multi-Rollout On-Policy Distillation via Peer Successes and Failures

    May 12, 2026Weichen Yu, Xiaomin Li, Yizhou Zhao +8Token-Level SupervisionParallel Rollouts

  19. Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

    May 12, 2026Zhong Guan, Yongjian Guo, Haoran Sun +5Off-Policy LearningParallel Rollouts

  20. Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair

    May 8, 2026Jia Li, Yuxin Su, Ting Peng +3Group-Based Reinforcement LearningAutomated Program Repair