On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
Figures & tables
Figure 1 : Failed rollouts trace back to an early pivotal mistake, and standard on-policy distillation does not repair it. (a) Pivotal turns arrive early; the remaining optimal trajectory is short, yet students waste the turns until the end. (b) Correcting the pivotal turn or guiding recovery turns both restore success effectively. (c) OPD eliminates most incomplete-trajectory failures, but pivotal-turn failures persist across training.
Figure 2 : Overview of PivotOPD . (1) Pivot detection. A teacher model reads each rollout with its outcome, selects candidate turns, and names a gold action at each. A candidate turn is pivotal when the student’s committed action differs from the gold action. After each pivotal turn, the teacher names a recovery action at each of the next few turns. (2) Preventive distillation trains the student to avoid the pivotal mistake. The frozen student hinted with the gold action serves as a privileged self-teacher that re-scores the student’s recorded response, and a reverse KL loss moves the student toward the gold action. (3) Recovery distillation trains the student to recover after the mistake. The self-teacher is instead hinted with the recovery action and writes a recovery response, on which the student is trained without the hint through a forward KL loss. Later recovery turns start from the state reached by executing its action in a copy of the environment that replays all preceding actions. Both terms are combined with group-based RL in a single PPO update.
Method
ALFWorld
Search-based QA
WebShop
Pick
Look
Clean
Heat
Cool
Pick2
Avg.
NQ
Triv
Pop
Hotp
2Wk
MuS
Bam
Avg.
Score
Succ.
Qwen3-1.7B
Base Model
6.8
60.2
0.0
0.0
3.6
4.1
12.4
16.8
50.8
37.9
26.9
21.0
10.0
15.6
25.6
45.0
4.6
OPSD
23.7
31.2
10.9
0.0
2.2
6.5
12.4
43.4
57.9
48.2
33.3
34.0
10.0
30.5
36.8
49.2
10.1
GRPO
74.0
44.1
33.3
41.9
30.4
32.5
42.7
42.4
58.3
50.2
38.2
38.5
9.1
27.4
37.7
74.5
57.0
Skill-GRPO
89.8
62.4
73.0
50.4
68.8
43.9
64.7
43.0
57.3
44.3
26.2
34.3
10.0
20.2
33.6
69.0
53.9
Table 1 : Results on ALFWorld, Search-based QA, and WebShop with Qwen3-1.7B and Qwen3-8B students. Results are averaged over three seeds, Avg. columns are unweighted means over task types or datasets, and every method with a teacher uses Qwen3-30B-A3B (1.7B) or Qwen3.5-122B-A10B (8B). QA columns abbreviate Natural Questions, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, MuSiQue, and Bamboogle. We highlight the best and second-best results.
Figure 3 : Self-distillation results on Qwen3-8B , where the student serves as its own teacher. PivotOPD is best on all three benchmarks, outperforming the strongest baseline by +3.9% on average.
Figure 4 : SWE-Bench Verified results. PivotOPD improves more than OPD.
Figure 5 : PivotOPD learns to recover from pivotal mistakes. Each policy continues from a replayed prefix that ends with the pivotal mistake ( Section E.2 ). (a) After the same pivotal mistake, the base model fails, whereas PivotOPD recovers and completes the task. (b) Top: the percentage of replays that recover within a given number of turns. Bottom: the average number of turns to recover among recovered replays, with the optimal number of turns in gray. PivotOPD recovers the most often and in the fewest turns.
Variant
Clean
Heat
Cool
Avg.
Random pivotal turns
76.2
61.1
50.0
62.6
Hints w/o gold actions
90.5
63.7
72.9
71.9
Reverse-KL recovery
85.7
50.0
61.5
64.5
Recovery only
66.7
40.9
53.9
64.7
Preventive only
93.2
70.2
72.9
72.5
PivotOPD
93.7
72.6
73.9
73.7
Table 2: Component ablations on ALFWorld with the Qwen3-1.7B student. Each variant changes one component of PivotOPD (preventive only also raises wprev ). The average is over all six task types ( Table S.4 ).
Figure 6 : Ablation on the recovery budget with the Qwen3-1.7B student. Each panel reports the validation performance trend with different numbers of recovery turns K∈{0,1,2,3} during training, where K=0 is the preventive-only variant. Overall, K=1 yields the best performance for WebShop and Search-based QA, while ALFWorld needs K=2 recovery turns.
Figure 7 : Recovery distillation quickly brings the KL divergence from the self-teacher back down after the pivotal mistake.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Teacher
Student
±1 -turn accuracy (%)
Random baseline (%)
Gain over random (%)
Qwen3-30B-A3B
Qwen3-1.7B
71.1
33.6
+37.5
Qwen3.5-122B-A10B
Qwen3-8B
84.4
28.1
+56.3
Appendix
Table S.1: Teacher–oracle agreement on ALFWorld. Each teacher analyzes failed rollouts of the student that it trains. Accuracy is the percentage of teacher samples with at least one teacher-detected pivotal turn within one turn of the first oracle-labeled pivotal turn, averaged over failed trajectories with such a turn. The random baseline selects the same number of turns uniformly at random, and the last column subtracts it from the accuracy.
ALFWorld
WebShop
Search-based QA
Tasks per step
16
16
64
Rollouts per task (group size)
8
8
8
Trajectories per step
128
128
512
PPO mini-batch size
128
64
256
Learning rate
1×10−6
1×10−6
1×10−6
KL coefficient
0.01
0.01
0.01
Appendix
Table S.2: Training hyperparameters per benchmark. Both students share this configuration.
ALFWorld
WebShop
Search-based QA
Candidate turns per trajectory m
5
5
2
Recovery turns K
2
1
1
Recovery weight wrec
1.0
0.25
0.0625 (8B) / 0.125 (1.7B)
Preventive / lesson weights wprev / wtraj
0.001 / 0.001
0.001 / 0.001
0.001 / 0.001
Recovery clip δ
5.0
5.0
5.0
Max recoveries per training step
64
64
64
Appendix
Table S.3: PivotOPD hyperparameters per benchmark. The preventive-only ablation uses wprev=0.1 ( Appendix E ).
Variant
Pick
Look
Clean
Heat
Cool
Pick2
Avg.
Random pivotal turns
71.4
78.6
76.2
61.1
50.0
38.6
62.6
Hints w/o gold actions
87.3
82.7
90.5
63.7
72.9
34.1
71.9
Reverse-KL recovery
75.0
71.4
85.7
50.0
61.5
43.4
64.5
Recovery only
87.7
63.1
66.7
40.9
53.9
76.2
64.7
Preventive only
83.8
71.4
93.2
70.2
72.9
43.4
72.5
PivotOPD
87.6
55.9
93.7
72.6
73.9
58.5
73.7
Appendix
Table S.4: Full component ablations on ALFWorld with the Qwen3-1.7B student. Each variant changes one component of PivotOPD and holds the rest of the training recipe fixed, except that preventive only also raises wprev to 0.1 . This table extends Table 2 to every task type. Results are averaged over three seeds.
Computation
K=0
K=1
K=2
K=3
Self-teacher forward pass
17.2 s ( +4.7% )
Recovery rollouts
New A
28.5 s ( +7.8% )
328.7 s ( +89.5% )
395.1 s ( +107.6% )
Total overhead
+4.7%
+12.4%
+94.2%
+112.3%
Appendix
Table S.5: Training compute overhead of PivotOPD on ALFWorld with the Qwen3-1.7B student. Each entry is the median time per training step of a computation that preventive and recovery distillation add to group-based RL, with the resulting increase over the time per training step of GRPO in parentheses. The last row sums both increases, and K=0 is the preventive-only variant. The overhead stays small with one recovery turn and grows once later recovery turns require environment replay.
Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large models whose inference cost is prohibitive in practice. On-Policy Distillation (OPD) is a natural recipe for transferring such capabilities to smaller students, but we find that it suffers a characteristic failure mode in this setting: small student errors compound across turns and push the trajectory out of the teacher's familiar state distribution, so the teacher's supervision becomes least reliable precisely where the student needs it most. We propose Guided On-Policy Distillation (Guided-OPD), a simple yet effective algorithm that mixes teacher- and student-generated turns within each rollout and schedules the teacher's intervention probability along a curriculum that decays to zero. Strong guidance keeps early trajectories close to the teacher distribution and is then gradually withdrawn to recover the purely on-policy regime used at inference. On ALFWorld, ScienceWorld, and WebShop, distilling Qwen3 students from a Qwen3-30B-A3B teacher, Guided-OPD yields average relative gains of 21.1% in Score and 25.5% in Success Rate over vanilla OPD, with larger gains on smaller students.
Gengsheng Li, Mao Zheng, Mingyang Song +8
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Large Language Model Department, Tencent +4
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes: the student acts at selected steps, while the teacher provides dense per-step supervision without executing new environment interactions. We show that multi-turn OPD introduces a prefix trap: making histories more student-on-policy improves relevance to the student, but can query the teacher on histories where its target is unreliable. This creates a two-sided distribution shift between student occupancy and teacher reliability. ReOPD addresses this by treating multi-turn OPD as a reliability-aware prefix distribution design and implements it with a simple step-decaying sampling schedule that emphasizes early, lower-shift prefixes. Across mathematical reasoning with Python and search environments over multiple teacher and student model scales, ReOPD preserves or improves OPD-level accuracy, uses zero tool calls during student training, and is at least 4× faster per rollout than OPD. ReOPD therefore turns expensive agent-environment interaction into a reusable offline resource, enabling scalable distillation across tools, tasks, and environments.
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to 15.8% relative to standard OPD.
Wenbo Zhang, Pengcheng Xu, Weizhi Du +2
University of California, Irvine · University of Michigan, Ann Arbor