PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
Organizations: Princeton University · NVIDIA · University of Maryland
Abstract
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
Figures & tables
| Method | ALFWorld | Search-based QA | WebShop | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pick | Look | Clean | Heat | Cool | Pick2 | Avg. | NQ | Triv | Pop | Hotp | 2Wk | MuS | Bam | Avg. | Score | Succ. | |
| Qwen3-1.7B | |||||||||||||||||
| Base Model | 6.8 | 60.2 | 0.0 | 0.0 | 3.6 | 4.1 | 12.4 | 16.8 | 50.8 | 37.9 | 26.9 | 21.0 | 10.0 | 15.6 | 25.6 | 45.0 | 4.6 |
| OPSD | 23.7 | 31.2 | 10.9 | 0.0 | 2.2 | 6.5 | 12.4 | 43.4 | 57.9 | 48.2 | 33.3 | 34.0 | 10.0 | 30.5 | 36.8 | 49.2 | 10.1 |
| GRPO | 74.0 | 44.1 | 33.3 | 41.9 | 30.4 | 32.5 | 42.7 | 42.4 | 58.3 | 50.2 | 38.2 | 38.5 | 9.1 | 27.4 | 37.7 | 74.5 | 57.0 |
| Skill-GRPO | 89.8 | 62.4 | 73.0 | 50.4 | 68.8 | 43.9 | 64.7 | 43.0 | 57.3 | 44.3 | 26.2 | 34.3 | 10.0 | 20.2 | 33.6 | 69.0 | 53.9 |
| Variant | Clean | Heat | Cool | Avg. |
|---|---|---|---|---|
| Random pivotal turns | 76.2 | 61.1 | 50.0 | 62.6 |
| Hints w/o gold actions | 90.5 | 63.7 | 72.9 | 71.9 |
| Reverse-KL recovery | 85.7 | 50.0 | 61.5 | 64.5 |
| Recovery only | 66.7 | 40.9 | 53.9 | 64.7 |
| Preventive only | 93.2 | 70.2 | 72.9 | 72.5 |
| PivotOPD | 93.7 | 72.6 | 73.9 | 73.7 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Teacher | Student | -turn accuracy (%) | Random baseline (%) | Gain over random (%) |
|---|---|---|---|---|
| Qwen3-30B-A3B | Qwen3-1.7B | |||
| Qwen3.5-122B-A10B | Qwen3-8B |
| ALFWorld | WebShop | Search-based QA | |
|---|---|---|---|
| Tasks per step | |||
| Rollouts per task (group size) | |||
| Trajectories per step | |||
| PPO mini-batch size | |||
| Learning rate | |||
| KL coefficient |
| ALFWorld | WebShop | Search-based QA | |
|---|---|---|---|
| Candidate turns per trajectory | |||
| Recovery turns | |||
| Recovery weight | (8B) / (1.7B) | ||
| Preventive / lesson weights / | / | / | / |
| Recovery clip | |||
| Max recoveries per training step |
| Variant | Pick | Look | Clean | Heat | Cool | Pick2 | Avg. |
|---|---|---|---|---|---|---|---|
| Random pivotal turns | 71.4 | 78.6 | 76.2 | 61.1 | 50.0 | 38.6 | 62.6 |
| Hints w/o gold actions | 87.3 | 82.7 | 90.5 | 63.7 | 72.9 | 34.1 | 71.9 |
| Reverse-KL recovery | 75.0 | 71.4 | 85.7 | 50.0 | 61.5 | 43.4 | 64.5 |
| Recovery only | 87.7 | 63.1 | 66.7 | 40.9 | 53.9 | 76.2 | 64.7 |
| Preventive only | 83.8 | 71.4 | 93.2 | 70.2 | 72.9 | 43.4 | 72.5 |
| PivotOPD | 87.6 | 55.9 | 93.7 | 72.6 | 73.9 | 58.5 | 73.7 |
| Computation | ||||
|---|---|---|---|---|
| Self-teacher forward pass | s ( ) | |||
| Recovery rollouts | New A | s ( ) | s ( ) | s ( ) |
| Total overhead | ||||