Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emph{policy-side correction} alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term \emph{advantage staleness}. We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance. This motivates the hypothesis that policy- and advantage-side correction should be coordinated. We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation. Joint parameter sweeps across staleness levels support this hypothesis: the effect of one correction parameter depends on, and can reverse with, the other. COPC achieves the highest reported performance on tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting. It also offers a broad high-performing parameter region and improved training stability. In search, COPC remains stable throughout training, while most evaluated asynchronous baselines collapse late in training. These gains persist at 64-step policy staleness. COPC adds minimal step-time overhead over asynchronous PPO and retains a 1.7× step-time speedup over synchronous PPO.
Figures & tables
Figure 1: Mean avg@32 accuracy on AIME 2025 under joint policy- and advantage-side correction sweeps at staleness levels s∈{8,12,16} .
Method
AIME24
AIME25
AIME26
BRUMO25
BeyondAIME
Avg.
Qwen3-4B-Instruct
PPO
54.65
43.75
47.40
44.83
27.53
43.63
V-trace
55.07
46.28
50.17
46.42
28.80
45.35
DAPO
56.42
48.99
51.11
46.08
29.17
46.35
SAO
58.06
50.31
53.13
47.29
30.48
47.85
AReaL
58.40
51.25
53.47
47.01
31.19
48.26
Table 1: Performance on tool-integrated mathematical reasoning measured by avg@32. Best in bold and second-best underlined within each backbone group. Δ rows report gains of COPC over PPO.
Method
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Llama-3.2-3B-Instruct
PPO
41.00
52.20
37.20
37.60
27.00
26.00
40.80
37.40
V-trace
40.40
55.40
36.80
38.20
25.20
26.20
42.40
37.80
DAPO
44.00
54.60
36.80
39.40
39.80
28.80
45.60
41.29
SAO
39.80
48.20
39.20
37.40
40.20
27.20
46.40
39.77
AReaL
40.80
52.80
39.20
43.20
41.00
30.80
44.80
41.80
Table 2: Performance on search measured by greedy-decoding accuracy. Best in bold and second-best underlined within each backbone group. Δ reports gains of COPC over PPO.
Figure 2: Training stability and robustness. Left: training dynamics of asynchronous methods on search. Right: COPC training dynamics on AIME 2025 under different policy-staleness levels.
Step time (s)
Throughput (tokens/s/GPU)
Step-time speedup
Tool-integrated mathematical reasoning
Synchronous PPO
410.59
614.96
1.00×
Asynchronous PPO
236.80
856.93
1.73×
COPC (ours)
238.95
836.85
1.72×
Search
Synchronous PPO
213.18
65.35
1.00×
Table 3: Training efficiency on tool-integrated mathematical reasoning and search.
Table 5: Default training configurations for tool-integrated mathematical reasoning and search. Rollout generation and policy optimization are executed asynchronously on separate worker groups. Method-specific deviations are described in the accompanying text.
Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-policy correction. In heterogeneous training systems, the total importance ratio should ideally be decomposed into two semantically distinct factors: a \emph{training--inference discrepancy term} that aligns inference-side and training-side distributions at the same behavior-policy version, and a \emph{policy-staleness term} that constrains the update from the historical policy to the current policy. We show that practical asynchronous pipelines with delayed updates and partial rollouts often lose the required historical training-side logits, or old logits. This missing-old-logit problem entangles discrepancy repair with staleness correction, breaks the intended semantics of decoupled correction, and makes clipping and masking thresholds interact undesirably. To address this issue, we study both exact and approximate correction routes. We propose three exact old-logit acquisition strategies: snapshot-based version tracking, a dedicated old-logit model, and synchronization via partial rollout interruption, and compare their system trade-offs. From the perspective of approximate correction, we focus on preserving the benefits of decoupled correction through a more appropriate approximate policy when exact old logits cannot be recovered at low cost, without incurring extra system overhead. Following this analysis, we adopt a revised PPO-EWMA method, which achieves significant gains in both training speed and optimization performance.
Zhong Guan, Yongjian Guo, Haoran Sun +5
Tianjin University · Tsinghua University · JDT AI Infra +1
Asynchronous reinforcement learning has become the standard way to scale training for large language models (LLM), but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target free of the reward and a long one lets the product of importance ratios drift exponentially with the trajectory length. We propose BRACE, an anchored Bellman-residual correction for stale value models. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte-Carlo tail beyond it, which separates policy correction from reward propagation. BRACE delivers a 9.8% relative improvement in mean@1 on BrowseComp-Plus over the strongest baseline, runs 2.46× faster per step than synchronous training, and remains stable 50 updates off-policy.
Guanqun Zhao, Zijun Xie, Binbin Zheng +3
Beijing University of Posts and Telecommunications · Baidu Inc. · Peking University +1
Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess pool occupancy into an adaptive rejection budget and ranks completed trajectories using an effective-staleness score that combines Waiting Staleness with prefix-aware Generation Staleness. This avoids penalizing long or interrupted rollouts solely because they span multiple policy versions. In single-turn mathematical reasoning, PACE improves the six-benchmark average validation accuracy by 18.7% over unfiltered asynchronous RL at the same wall-clock budget and matches synchronous RL performance with 47.1% less GPU time. PACE also improves validation performance in multi-turn tool-integrated reasoning, outperforming both synchronous and unfiltered asynchronous RL. Further experiments with the mixture-of-experts model and an alternative RL algorithm support its applicability across model architectures and training algorithms.
Chenliang Li, Neiwen Ling, Zijun Wei +1
ByteDance · Texas A&M University · University of Virginia