cs.LGOct 5, 2026

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

Authors: Seil Kang, Hangoo Kang, Tarun Suresh, Youngeun Kim, Shreyas Pimpalgaonkar, Seong Jae Hwang, Azalia Mirhoseini

Organizations: Stanford University · Yonsei University · Korea University · Bespoke Labs

Abstract

Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchronous training overlaps rollout and learning across updates, but comes at the cost of policy staleness. We introduce ThunderSyncRL, which starts gradient computation as soon as all required inputs are fixed, without policy staleness. For group relative policy optimization (GRPO), ThunderSyncRL computes each trajectory's score gradient as soon as the reward for that trajectory arrives, without waiting for the group. For on-policy distillation (OPD), it computes gradients for each completed agentic turn's teacher-scored actions while tool calls run in the sandbox. We prove that gradient streaming produces the same GRPO and OPD updates as batch-synchronous training, without changing either objective. On SWE-bench Verified and Terminal Bench 4.0, we train models to the same performance up to 1.9×1.9 \times faster than synchronous training. With zero policy staleness, ThunderSyncRL also outperforms asynchronous training at a fixed budget by up to 2.472.47 percentage points.

Figures & tables

Explore similar work

CardsList
  1. Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    Jul 8, 2026Zhenyu Hou, Yujiang Li, Jie Tang +1Agentic LearningAutoregressive Rollout

  2. Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training

    Sep 29, 2026Chenliang Li, Neiwen Ling, Zijun Wei +1Gradient Staleness