Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchronous training overlaps rollout and learning across updates, but comes at the cost of policy staleness. We introduce ThunderSyncRL, which starts gradient computation as soon as all required inputs are fixed, without policy staleness. For group relative policy optimization (GRPO), ThunderSyncRL computes each trajectory's score gradient as soon as the reward for that trajectory arrives, without waiting for the group. For on-policy distillation (OPD), it computes gradients for each completed agentic turn's teacher-scored actions while tool calls run in the sandbox. We prove that gradient streaming produces the same GRPO and OPD updates as batch-synchronous training, without changing either objective. On SWE-bench Verified and Terminal Bench 4.0, we train models to the same performance up to 1.9× faster than synchronous training. With zero policy staleness, ThunderSyncRL also outperforms asynchronous training at a fixed budget by up to 2.47 percentage points.
Figures & tables
Figure 1: ThunderSyncRL matches Sync’s SWE-bench Verified pass@3 at up to 1.9× lower training cost under GRPO (left pair) and OPD (right pair). At a 128-GPU-hour budget, its zero-staleness schedule improves pass@3 by up to 2.47 percentage points over Async and Fully Async, both of which admit stale rollouts.
Figure 2: Schematic schedules for GRPO (a) and OPD (b). ∙ Sync waits for full batches; ∙ Async overlaps adjacent batches with one-step staleness; ∙ Fully Async updates rollout weights while trajectories remain in flight. ∙ ThunderSyncRL streams ready gradients with one batch-synchronous optimizer step and zero staleness.
Figure 3: Pipeline bubbles in synchronous agentic RL. GPU utilization during GRPO training on SWE-smith, with four rollout GPUs above the axis and four learner GPUs below. Rollout utilization (busy share of generation slots) drops as long-tail trajectories finish. Hatched bands mark idle time. ∙ Sync learners wait for the last reward for 49.5% of each step, and rollout GPUs wait while the learner trains. ∙ ThunderSyncRL starts each trajectory’s backward pass when its reward arrives and completes 15 optimizer steps in the GPU-hours Sync needs for 8.
Figure 4: Step wall time (left) and trainer active MFU (right) for GRPO and OPD.
Figure 5: Left: time to matched pass@3. Right: Async and Fully Async relative to ThunderSyncRL at 128 GPU-hours.
Figure 6: Terminal Bench 4.0 pass@3 vs. training cost and mean training reward vs. GPU-hours. Arrows mark where ThunderSyncRL reaches the score of a Sync checkpoint; labels give the cost ratio.
Table 1: Full fine-tuning efficiency on SWE-smith across models using NVIDIA B200 GPUs. GRPO and OPD are compared separately per configuration; for OPD, the model is the student.