Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emph{policy-side correction} alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term \emph{advantage staleness}. We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance. This motivates the hypothesis that policy- and advantage-side correction should be coordinated. We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation. Joint parameter sweeps across staleness levels support this hypothesis: the effect of one correction parameter depends on, and can reverse with, the other. COPC achieves the highest reported performance on tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting. It also offers a broad high-performing parameter region and improved training stability. In search, COPC remains stable throughout training, while most evaluated asynchronous baselines collapse late in training. These gains persist at 64-step policy staleness. COPC adds minimal step-time overhead over asynchronous PPO and retains a 1.7× step-time speedup over synchronous PPO.
Figures & tables
Figure 1: Mean avg@32 accuracy on AIME 2025 under joint policy- and advantage-side correction sweeps at staleness levels s∈{8,12,16} .
Method
AIME24
AIME25
AIME26
BRUMO25
BeyondAIME
Avg.
Qwen3-4B-Instruct
PPO
54.65
43.75
47.40
44.83
27.53
43.63
V-trace
55.07
46.28
50.17
46.42
28.80
45.35
DAPO
56.42
48.99
51.11
46.08
29.17
46.35
SAO
58.06
50.31
53.13
47.29
30.48
47.85
AReaL
58.40
51.25
53.47
47.01
31.19
48.26
Table 1: Performance on tool-integrated mathematical reasoning measured by avg@32. Best in bold and second-best underlined within each backbone group. Δ rows report gains of COPC over PPO.
Method
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Llama-3.2-3B-Instruct
PPO
41.00
52.20
37.20
37.60
27.00
26.00
40.80
37.40
V-trace
40.40
55.40
36.80
38.20
25.20
26.20
42.40
37.80
DAPO
44.00
54.60
36.80
39.40
39.80
28.80
45.60
41.29
SAO
39.80
48.20
39.20
37.40
40.20
27.20
46.40
39.77
AReaL
40.80
52.80
39.20
43.20
41.00
30.80
44.80
41.80
Table 2: Performance on search measured by greedy-decoding accuracy. Best in bold and second-best underlined within each backbone group. Δ reports gains of COPC over PPO.
Figure 2: Training stability and robustness. Left: training dynamics of asynchronous methods on search. Right: COPC training dynamics on AIME 2025 under different policy-staleness levels.
Step time (s)
Throughput (tokens/s/GPU)
Step-time speedup
Tool-integrated mathematical reasoning
Synchronous PPO
410.59
614.96
1.00×
Asynchronous PPO
236.80
856.93
1.73×
COPC (ours)
238.95
836.85
1.72×
Search
Synchronous PPO
213.18
65.35
1.00×
Table 3: Training efficiency on tool-integrated mathematical reasoning and search.
Table 5: Default training configurations for tool-integrated mathematical reasoning and search. Rollout generation and policy optimization are executed asynchronously on separate worker groups. Method-specific deviations are described in the accompanying text.