DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents
Authors: Hanyang Wang, Zeyuan Liu, Zhengyu Chen, Jingqing Ruan, Chaoxu Pang, Zhongda Su, Wulin Xie, Zhizhao Zeng, +2 more
Organizations: University of Chicago · Meituan LongCat Interaction Team · University of the Chinese Academy of Sciences · The Hong Kong University of Science and Technology (Guangzhou)
On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, wasting already-generated experience. To address this problem, we introduce DivOPD, a simple learner-side batch-selection method that spreads a fixed turn budget across more rollouts and, within each rollout, prioritizes turns with larger cumulative teacher-student disagreement. Turns without usable teacher feedback are excluded. The per-turn loss and optimizer remain fixed; selection only changes which student-visited turns receive training weight. For no-progress rollouts, an optional extension briefly hands control to the teacher before returning it to the student. Across six teacher-student settings on the simulated ALFWorld, ScienceWorld, and WebShop benchmarks, with 1.5B-7B students, DivOPD raises cross-setting mean peak success rate from 77.4 to 84.4 and mean success over the last five evaluations from 71.5 to 78.6. It reaches all reported setting-specific targets with geometric-mean speedups of 1.84x in training tokens and 1.87x in learner GPU time relative to vanilla OPD. Teacher intervention further raises this last-five mean to 82.4 while retaining about 1.7x learner-GPU speedup over vanilla OPD. Code will be released at https://github.com/HanyangWang0418-oss/DivOPD.
Figures & tables
Figure 1: Batch composition in asynchronous multi-turn OPD. (A) Rollout workers enqueue variable-length interactions in arrival order. (B) When production outpaces learner consumption, early or long rollouts can fill the batch, include unscored rows, and leave others to expire as stale. (C) DivOPD composes batches on the learner side through the three stages in Section 3 : validity gate, rollout-first coverage, and within-rollout disagreement focus using s(u) (Equation 2 ). (D) When rollouts repeatedly make no progress, bounded teacher recovery helps the student continue; teacher actions are excluded from the loss, and recovery is disabled at evaluation. Bottom: independent exploration and training, with periodic model synchronization and stale-row expiry.
ALFWorld
Qwen3-1.7B ( T /Init: 87.86/7.86; τ=70 )
Qwen3-4B ( T /Init: 87.86/26.43; τ=80 )
Method
Peak ↑
Final-5 ↑
Tok. ↑
GPU ↑
Peak ↑
Final-5 ↑
Tok. ↑
GPU ↑
Vanilla OPD
77.86
74.00 ± 3.18
1.00 ×
1.00 ×
86.43
84.57 ± 1.48
1.00 ×
1.00 ×
TCOD-F2B
72.86
68.43 ± 3.09
1.62 ×
1.63 ×
87.14
83.00 ± 1.28
1.49 ×
1.51 ×
TurnOPD
78.57
75.71 ± 2.72
1.77 ×
1.76 ×
87.14
86.43 ± 0.51
1.57 ×
1.59 ×
DivOPD-focus
80.00
71.00 ± 6.62
1.50 ×
1.46 ×
89.29
85.29 ± 3.10
1.69 ×
1.70 ×
DivOPD-base
74.29
72.86 ± 1.43
1.35 ×
1.35 ×
92.14
88.86 ± 2.29
1.27 ×
1.27 ×
Table 1: Main results. Peak/Final-5 are best-checkpoint/final-five mean SR (%); ± is the standard deviation over the last five checkpoints. Headers give teacher/student-initial SR. Tok./GPU are Cvanilla(τ)/Cmethod(τ) for cumulative learner training tokens and trainer GPU-hours, respectively. Larger is better; a dash means that the method did not reach τ . Red/purple mark best/second-best; DivOPD-focus is an uncontrolled reference.
Load
Reader
Rollout bs
Arrival/drain
Final SR ↑
nAUC@150 ↑
Stale rows ↓
Low
Vanilla OPD
2
0.77
52.1
25.0
240
DivOPD-base
2
0.55
52.0
25.4
53
Medium
Vanilla OPD
6
1.35
53.6
25.3
8,035
DivOPD-base
6
0.88
56.5
26.0
4,237
High
Vanilla OPD
16
2.27
49.3
26.9
30,175
DivOPD-base
16
0.89
59.3
27.8
9,766
Table 2: Stress test under increasing asynchronous producer load. ALFWorld Qwen3-1.7B over 150 learner updates. Arrival/drain is the measured queue arrival rate divided by the learner drain rate; values above one indicate a growing backlog. SR and nAUC are percentages; stale rows are raw counts. Bold marks the better value within each load level.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Input: update index q ; pending turns P ; rollout buffer E ; batch size B ; pool multiplier c ; initial per-rollout cap k ; maximum staleness δ ; maximum pending age W .
1
Set vmin←max(q−δ,0) and remove every u∈P with version(u)<vmin .
2
Read F from E with version at least vmin until ∣P∣+∣F∣=cB (or the read returns no more turns); set C←P∪F .
3
Set V←{u∈C:\textscValid(u)} and permanently discard C∖V . Stamp a fresh turn’s first-seen update as q .
4
Group V by rollout identifier r . Order rollout groups by oldest first-seen update, breaking ties by r .
5
Within every group, sort turns by descending s(u) , breaking ties by (r,t) . Initialize the selected set S←∅ and cap h←k .
6
Sweep the ordered rollout groups once. From each group append its highest ranked unselected turns until that group contributes k turns or ∣S∣=B .
Appendix
Algorithm 1: DivOPD sampler ( trajectory_first with turn_signal=raw_kl_sum ). Table 3 lists the reported numerical settings.
Table 3: Shared sampler configuration. All methods use the same learner batch size. DivOPD-base and DivOPD share these sampler values; only their within-rollout selection differs.
Setting
Value
Optimizer
AdamW
Learning rate
10−6 ; constant schedule; no warmup
Adam moments
β1=0.9,β2=0.999
Weight decay
0.01
Gradient-norm clipping
1.0
Optimizer batch
N=64 turn rows; seq-mean-token-mean (Equation 3 )
Appendix
Table 4: Shared optimizer configuration. These settings are fixed across all methods.
Environment
Student
Initial SR (%)
ALFWorld
Qwen3-1.7B
7.86
ALFWorld
Qwen3-4B
26.43
WebShop
Qwen2.5-3B-Instruct
1.56
WebShop
Qwen2.5-7B-Instruct
14.06
ScienceWorld
Qwen2.5-1.5B-Instruct
21.48
ScienceWorld
Qwen2.5-3B-Instruct
45.31
Appendix
Table 5: Student initialization, frozen teachers, and evaluation protocol. Initial SR is measured before the first learner update under the reported evaluation protocol; teacher SR uses the same fixed test episodes. Training responses are sampled at temperature 1.0 , and evaluation uses temperature 0.4 .
Table 6: ScienceWorld task-type split. Training and evaluation use disjoint task types. Within each type, we retain the first half of the official variation IDs before applying a fixed shuffle.
Setting
Method
Accuracy
Efficiency to τ
Peak ↑
Final-5 ↑
Updates ↓
Trainer GPU h ↓
GPU speedup ↑
ALFWorld Qwen3-4B τ=80
Vanilla OPD
86.43
84.57 ± 1.48
184
13.51
1.00 ×
TCOD-F2B
87.14
83.00 ± 1.28
182
8.95
1.51 ×
TurnOPD
87.14
86.43 ± 0.51
161
8.50
1.59 ×
DivOPD-focus
89.29
85.29 ± 3.10
148
7.97
1.70 ×
DivOPD-base
92.14
88.86 ± 2.29
154
10.61
1.27 ×
Appendix
Table 7: Full results across six teacher–student settings. Peak SR is the best checkpoint; Final-5 is the mean ± standard deviation over the last five evaluation checkpoints. Best values are highlighted in bold; second-best are underlined.
Design question
Component
DivOPD
How is each turn scored?
Learning signal
Unchanged (Eq. 1 )
How long can a rollout continue?
Rollout horizon
Unchanged
Who acts during the rollout?
Rollout policy
Changed only by DivOPD + R (§ 3.4 )
Which turns enter a learner batch?
Batch selection
Changed by DivOPD (§ 3 )
Appendix
Table 8: What DivOPD changes. The core method changes batch selection; only the optional recovery extension changes rollout generation.
Setting
Method
Eff. rollouts
Top-1 tok.
Turns per
KL / token
Step time
Peak SR
/ batch ↑
share ↓
rollout
↓
(s) ↓
(%) ↑
ALFWorld 4B
Vanilla OPD
6.7
0.281
8.1
0.121
94.9
86.43
TCOD-F2B
6.5
0.291
8.6
0.114
99.6
87.14
DivOPD-focus
18.6
0.110
2.4
0.174
85.2
89.29
DivOPD-base
12.5
0.136
3.9
0.111
97.8
92.14
DivOPD
12.5
0.145
3.9
0.176
82.1
91.43
Appendix
Table 9: Full batch-composition diagnostics. Values are averaged over the final 40% of training; peak SR is repeated from Appendix Table 7 . Focus-only maximizes coverage in five settings, while DivOPD has the higher peak SR in five. Bold marks the best peak SR among the five methods shown, which exclude DivOPD + R. Figure 2 averages the first three diagnostics over the two student sizes per environment.
Dropped share
Turn depth
Score/token
Setting
Turns
Grad. norm
Kept
Dropped
Kept
Dropped
ScienceWorld 3B
55%
40%
4.7
9.5
0.30
0.16
WebShop 3B
18%
12%
2.3
4.5
0.16
0.08
Appendix
Table 10: Focus retention audit ( k=5 , the initial cap used in Table 1 for these settings). Dropped gradient norm is the share of summed per-turn norms, ∑t∥gt∥ , estimated by the sketch audit. The last four columns are means for kept and dropped turns.
Never selected (%)
Gini
Composer
ALF
SciW
WS
ALF
SciW
WS
Shared buffers: first 100 updates
Vanilla OPD (arrival-order)
53.9
49.2
58.7
0.66
0.68
0.64
Top- B (no rollout limit)
–
–
–
0.44
0.45
0.27
DivOPD-focus
15.0
10.2
3.4
0.33
0.32
0.20
DivOPD-base / DivOPD
14.2
11.0
7.1
0.30
0.32
0.12
Appendix
Table 11: Cumulative rollout utilization. Never selected is the share of valid rollouts contributing no trained turn; Gini measures concentration of selected-turn counts. ALF/SciW/WS denote ALFWorld/ScienceWorld/WebShop. The separate-buffer result is grouped below the shared-buffer comparisons and is not a matched control.
Setting
max_staleness
Mean age
Peak
Final ckpt.
ALFWorld 1.7B
2
1.69
77.86
73.57
4
3.19
71.43
71.43
ScienceWorld 1.5B
2
1.68
66.02
66.02
4
2.96
65.62
60.16
Appendix
Table 12: Vanilla OPD with a wider staleness window. Mean age of consumed rows in policy versions (arrival-order replay, dead rows included), peak and final-checkpoint SR (%). The max_staleness=2 rows reuse the main vanilla runs in Table 1 ; the wider-window rows are separate auxiliary runs rather than paired reruns.
Gradient projection
Cosine similarity
Composer
ALF
SciW
WS
ALF
SciW
WS
Arrival order †
1.74
0.75
1.32
0.97
0.49
0.61
Gate only
1.11
1.02
1.29
0.93
0.53
0.62
DivOPD-base
1.36
0.97
1.31
0.96
0.54
0.62
DivOPD
2.04
2.01
1.47
0.98
0.75
0.64
DivOPD-focus
3.12
1.66
1.50
0.99
0.67
0.65
Appendix
Table 13: Selected gradients versus the candidate-pool gradient. Means over 8 audited updates per environment. Projection is ⟨G,G^⟩/∥G∥2 ; cosine similarity is cos(G,G^) . ALF/SciW/WS denote ALFWorld/ScienceWorld/WebShop. † Arrival order here draws from the finite- k composer’s pool and is not vanilla OPD. Sketch oracle selects and evaluates the B turns with the largest sketched ⟨G,gt⟩ , so it is an optimistic reference rather than a full-gradient oracle.
Depth
Outcome
Vanilla
+ Gate
+ Coverage
+ Focus
<2
Failure
3.2
3.8
6.1
10.6
Success
17.6
20.7
14.3
15.8
[2,4)
Failure
3.1
3.6
6.1
8.5
Success
17.4
20.4
14.1
14.5
[4,6)
Failure
2.9
3.4
5.9
7.4
Success
13.1
15.3
9.7
9.7
Appendix
Table 14: Where selected turns come from in the ALFWorld 1.7B DivOPD replay. Share of selected turns (%), grouped by turn depth and rollout outcome. Each method column sums to approximately 100% after rounding; stages are added one at a time from left to right.
Setting
Trained tokens
Trainer GPU
Wall time (est.)
ALFWorld 4B
1.92 ×
1.94 ×
1.36 ×
ALFWorld 1.7B
2.05 ×
2.07 ×
1.28 ×
WebShop 3B
2.22 ×
2.22 ×
0.99 ×
WebShop 7B
1.77 ×
1.92 ×
0.83 ×
ScienceWorld 3B
1.62 ×
1.63 ×
1.22 ×
ScienceWorld 1.5B
1.54 ×
1.54 ×
1.27 ×
Appendix
Table 15: Resource efficiency to τ . Vanilla-to-DivOPD ratios; larger is better. Trainer GPU time uses the two GPUs assigned to the learner. The wall estimate sums trainer-step and experience-read time. Its speedup differs from timestamp-based end-to-end measurement by at most 1.2% on the three settings whose timestamps were retained.
DivOPD ( k=3 )
DivOPD-focus
Setting
scored/trained
not selected
scored/trained
not selected
ALFWorld 1.7B
2.77 ×
63.9%
2.89 ×
65.4%
ALFWorld 4B
2.73 ×
63.4%
2.93 ×
65.9%
ScienceWorld 1.5B
3.66 ×
72.6%
3.79 ×
73.6%
ScienceWorld 3B
3.41 ×
70.6%
3.60 ×∗
72.2%
WebShop 3B
1.89 ×
47.1%
1.86 ×
46.3%
Appendix
Table 16: Teacher tokens scored per token trained. Effective tokens; “not selected” is the share never entering an optimizer batch. This audit uses the k=3 runs from Table 19 , not the per-setting caps of Table 1 . ∗ Two parallel 150-update trainer histories sharing one buffer, merged.
Vanilla reader
Rollout-first selection
Environment
Model
Peak ↑
Best-5 ↑
Final ↑
Peak ↑
Best-5 ↑
Final ↑
ALFWorld
1.7B
78.57
73.86
78.57
80.00
77.71
78.57
WebShop
3B
75.00
68.44
71.09
78.12
74.06
68.75
ScienceWorld
1.5B
73.44
72.66
73.05
77.73
74.38
74.61
Appendix
Table 17: Matched candidate access. Both arms draw from the same 256 -turn stream; vanilla otherwise reads in arrival order. Best-5 averages the five best checkpoints; Final is SR at the common learner-version cutoff. These separate runs are not numerically comparable to Table 1 .
Finite k (DivOPD-base)
k=∞ (no cap)
Environment
Model
Peak SR ↑
Final-5 SR ↑
Peak SR ↑
Final-5 SR ↑
ALFWorld
1.7B
74.29
72.86
73.57
65.14
ALFWorld
4B
92.14
88.86
90.00
87.71
WebShop
3B
78.12
73.12
75.00
65.47
WebShop
7B
88.28
78.44
87.62
81.09
ScienceWorld
1.5B
70.31
67.50
67.97
62.89
Appendix
Table 18: Matched-access cap comparison. Both arms use the same candidate-pool size, gate, ordering, pending rule, and within-rollout sampling; only the initial cap differs. SR is in percent.
Method
ALFWorld
WebShop
ScienceWorld
1.7B
4B
3B
7B
1.5B
3B
TCOD-F2B
72.86
87.14
67.19
89.06
71.88
87.89
DivOPD ( k=3 )
87.14
90.00
71.09
91.41
77.73
85.55
DivOPD ( k=4 )
79.29
91.43
73.44
90.62
78.91
86.33
DivOPD ( k=5 )
79.29
91.43
78.12
88.28
75.78
88.28
DivOPD ( k=6 )
80.71
88.57
73.44
86.72
75.78
85.55
Appendix
Table 19: Sensitivity to the initial per-rollout cap k . Peak SR (%); TCOD-F2B is a reference. Best DivOPD values are highlighted, and second-best values are underlined.
While on-policy distillation (OPD) reduces exposure bias by training student language models on their own rollouts, early student errors in long-horizon agentic scenarios can lead to contexts unfamiliar to the teacher. To improve trajectory quality, recent work on agentic OPD introduces teacher intervention into training rollouts by switching the executor between the student and the teacher. However, existing methods determine how much teacher intervention is needed---but not when. To address this limitation, we propose DASH-OPD (Discrepancy-Aware Switching with Hysteresis for OPD), the first agentic OPD method to perform adaptive, bidirectional executor switching. At each turn, DASH-OPD measures teacher--student discrepancy using a mean log-probability ratio over action tokens. Student-to-teacher ratios on student turns serve as drift signals, while teacher-to-student ratios on teacher turns serve as recovery signals. These signals are accumulated over multiple turns to form drift and recovery evidence, respectively. DASH-OPD switches executors when either type of evidence exceeds its corresponding switching threshold, introducing hysteresis that prevents rapid switching triggered by transient discrepancy fluctuations. Across three benchmarks and two student model sizes, DASH-OPD outperforms five baselines in all 14 task performance comparisons, while requiring the fewest interaction turns in nine of ten efficiency comparisons. Code, models, and training logs are available at https://github.com/Lucian1115/DASH-OPD
Yuchen Xia, Qianguo Sun, Chao Song +3
The Chinese University of Hong Kong · IDEA Research · Emdoor Research Institute
On-policy distillation (OPD) has shown strong potential for transferring reasoning ability from frontier or domain-specific models to smaller students. While effective on static single-turn tasks, its behavior in multi-turn agent settings remains underexplored. In this work, we identify a key limitation of vanilla OPD in such settings, which we term Trajectory-Level KL Instability. Specifically, we observe that KL divergence increases together with a drop in success rate, and even after convergence, the KL remains high, leading to unstable training. This instability arises from inter-turn error compounding: as errors accumulate, the student is driven beyond the teacher's effective support, rendering the supervision signal unreliable. To address this, we propose TCOD (Temporal Curriculum On-Policy Distillation), a simple yet effective framework that controls the trajectory depth exposed to the student and progressively expands it from short to long with a curriculum schedule. Experimental results across four student-teacher pairs on three multi-turn agent benchmarks (ALFWorld, WebShop, ScienceWorld) show that TCOD mitigates KL escalation and enhances KL stability throughout training, improving agent performance by up to 18 points over vanilla OPD. Further evaluations show that TCOD can even surpass the teacher's performance and generalize to tasks on which the teacher fails. Our code is available at https://github.com/kokolerk/TCOD.
Jiaqi Wang, Wenhao Zhang, Weijie Shi +2
Tongyi Lab , Alibaba Group · † The Chinese University of Hong Kong.
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to 15.8% relative to standard OPD.
Wenbo Zhang, Pengcheng Xu, Weizhi Du +2
University of California, Irvine · University of Michigan, Ann Arbor