Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchronous training overlaps rollout and learning across updates, but comes at the cost of policy staleness. We introduce ThunderSyncRL, which starts gradient computation as soon as all required inputs are fixed, without policy staleness. For group relative policy optimization (GRPO), ThunderSyncRL computes each trajectory's score gradient as soon as the reward for that trajectory arrives, without waiting for the group. For on-policy distillation (OPD), it computes gradients for each completed agentic turn's teacher-scored actions while tool calls run in the sandbox. We prove that gradient streaming produces the same GRPO and OPD updates as batch-synchronous training, without changing either objective. On SWE-bench Verified and Terminal Bench 4.0, we train models to the same performance up to 1.9× faster than synchronous training. With zero policy staleness, ThunderSyncRL also outperforms asynchronous training at a fixed budget by up to 2.47 percentage points.
Figures & tables
Figure 1: ThunderSyncRL matches Sync’s SWE-bench Verified pass@3 at up to 1.9× lower training cost under GRPO (left pair) and OPD (right pair). At a 128-GPU-hour budget, its zero-staleness schedule improves pass@3 by up to 2.47 percentage points over Async and Fully Async, both of which admit stale rollouts.
Figure 2: Schematic schedules for GRPO (a) and OPD (b). ∙ Sync waits for full batches; ∙ Async overlaps adjacent batches with one-step staleness; ∙ Fully Async updates rollout weights while trajectories remain in flight. ∙ ThunderSyncRL streams ready gradients with one batch-synchronous optimizer step and zero staleness.
Figure 3: Pipeline bubbles in synchronous agentic RL. GPU utilization during GRPO training on SWE-smith, with four rollout GPUs above the axis and four learner GPUs below. Rollout utilization (busy share of generation slots) drops as long-tail trajectories finish. Hatched bands mark idle time. ∙ Sync learners wait for the last reward for 49.5% of each step, and rollout GPUs wait while the learner trains. ∙ ThunderSyncRL starts each trajectory’s backward pass when its reward arrives and completes 15 optimizer steps in the GPU-hours Sync needs for 8.
Figure 4: Step wall time (left) and trainer active MFU (right) for GRPO and OPD.
Figure 5: Left: time to matched pass@3. Right: Async and Fully Async relative to ThunderSyncRL at 128 GPU-hours.
Figure 6: Terminal Bench 4.0 pass@3 vs. training cost and mean training reward vs. GPU-hours. Arrows mark where ThunderSyncRL reaches the score of a Sync checkpoint; labels give the cost ratio.
Table 1: Full fine-tuning efficiency on SWE-smith across models using NVIDIA B200 GPUs. GRPO and OPD are compared separately per configuration; for OPD, the model is the student.
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged as a more efficient alternative by updating the model as rollouts arrive. However, existing asynchronous RL systems often emphasize throughput, while leaving training stability and task effectiveness largely underexplored. For example, a key challenge is that group-wise sampling in the widely-used GRPO framework does not naturally fit asynchronous agentic training. In this paper, we present Single-rollout Asynchronous Optimization (SAO) to address the stability and off-policy challenges in asynchronous RL. To reduce off-policy effects and improve generalization, we replace group-wise sampling with single-rollout sampling, that is, using one rollout per prompt. We further improve this single-rollout strategy with practical value-model training designs. To improve optimization stability, we introduce a strict double-side token-level clipping strategy. SAO is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench. We also demonstrate that single-rollout RL is particularly effective in a simulated online learning setting, where the model must adapt to changing evolving environments. To this end, SAO is successfully deployed in the agentic RL pipeline for training the open GLM-5.2 model (750B-A40B).
Zhenyu Hou, Yujiang Li, Jie Tang +1
Tsinghua University · Work done while ZH and YL interned at Z.AI.
Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess pool occupancy into an adaptive rejection budget and ranks completed trajectories using an effective-staleness score that combines Waiting Staleness with prefix-aware Generation Staleness. This avoids penalizing long or interrupted rollouts solely because they span multiple policy versions. In single-turn mathematical reasoning, PACE improves the six-benchmark average validation accuracy by 18.7% over unfiltered asynchronous RL at the same wall-clock budget and matches synchronous RL performance with 47.1% less GPU time. PACE also improves validation performance in multi-turn tool-integrated reasoning, outperforming both synchronous and unfiltered asynchronous RL. Further experiments with the mixture-of-experts model and an alternative RL algorithm support its applicability across model architectures and training algorithms.
Chenliang Li, Neiwen Ling, Zijun Wei +1
ByteDance · Texas A&M University · University of Virginia
In large-scale reinforcement learning (RL) systems with decoupled Trainer-Rollout execution, the Trainer must regularly synchronize policy weights to the Rollout side to limit policy staleness. When inter-node bandwidth is abundant, such synchronization is usually only a small fraction of end-to-end cost. As model size grows, however, the communication demand rises rapidly. In bandwidth-constrained or network-variable deployments -- for example, cross-datacenter or cross-cluster settings, heterogeneous resource pools, and online RL -- weight synchronization can become a dominant bottleneck for throughput and tail latency. We observe that, in mainstream large-model RL training, the locations where parameters actually change are highly sparse at the element level (often 99%+ sparsity). Building on this observation, we propose and implement SparseRL-Sync, which replaces full-weight transfers with a lossless sparse update payload (indices and values) that can be exactly reconstructed on the inference side, thereby preserving 100% fidelity. Under a simplified cost model, sparse synchronization reduces the per-update communication volume from S to approximately S/X; with 99% sparsity (X ~ 100), this yields about a 100x reduction in transmitted data. Combined with appropriate bucketing, SparseRL-Sync also reduces launch and control-plane overhead, significantly improving scalability and end-to-end efficiency in bandwidth-limited and highly asynchronous RL settings.