Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation
Authors: Yuhao Sun, Binrui Wu, Zhuoer Xu, Ming Wen, Haoxiang Xu, Bin Chen, Yan Lin, Qianzijing Zhang
Organizations: Ant Group · University of Science and Technology of China · Alibaba International Digital Commerce Group · Peking University · University of Electronic Science and Technology of China
On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future behavior, while consequential guidance may be beyond the current student's reach or fail to persist without privileged input. We formulate long-horizon OPD as hierarchical supervision allocation and argue that productive guidance lies at the intersection of future utility and current learnability. Crucially, this intersection evolves as the student learns. Based on this principle, we propose LENS-OPD, a coarse-to-fine framework that organizes supervision through Locate, Validate, and Refine. Locate adapts trajectory exposure to the student's evolving competence and proposes a candidate decision for intervention. Validate tests whether teacher guidance at that decision improves the same student's subsequent behavior. Refine internalizes the beneficial guided behavior into the deployable policy and concentrates token-level supervision on decisive teacher-student conflicts within the validated turn. These stages are nested: each finer allocation is conditioned on the coarser decision, rather than being optimized as an independent importance score. Experiments across multiple long-horizon agent benchmarks and student-teacher configurations show that LENS-OPD consistently improves task performance over vanilla OPD and strong curriculum- and selection-based baselines. Our results suggest that effective long-horizon distillation requires teaching at the right depth, the right decision, and the right token.
Figures & tables
Figure 1: Overview of LENS-OPD. Productive guidance combines future utility and current learnability, while Locate–Validate–Refine allocates supervision from trajectories to turns and tokens.
Figure 2: Diagnostics of hierarchical supervision allocation. (a) Student support for teacher-preferred tokens at disagreement positions changes markedly with training. (b) Turn-level KL is a weak proxy for the future benefit of guidance, with 37% of valid forks producing non-positive gain. (c) Within future-beneficial turns, one-step corrective response is largest when the student assigns lower support and the teacher expresses higher confidence. Error bars in (c) are 95% bootstrap confidence intervals clustered by fork.
Method
ALFWorld
WebShop
ScienceWorld
SR ↑
Score ↑
SR ↑
Score ↑
SR ↑
Qwen3-32B teacher → Qwen3-1.7B student
Student (zero-shot)
5.8
32.7
5.0
3.7
0.4
Teacher (zero-shot)
49.3
51.8
19.0
44.8
28.7
OPD
36.4±2.3
35.8±13.9
9.6±5.3
42.1±0.3
14.3±0.1
TCOD-F2B
42.6±1.4↑ 6.2
30.4±1.8↓ 5.4
7.4±1.3↓ 2.2
29.4±0.1↓ 12.7
15.4±0.2↑ 1.1
Table 1: Main results on ALFWorld, WebShop, and ScienceWorld under two Qwen3 student scales. SR denotes success rate; higher is better.
Method
Valid Seen
Valid Unseen
Hard
SR ↑
Rounds ↓
SR ↑
Rounds ↓
SR ↑
Rounds ↓
Qwen2.5-7B-RL teacher → Qwen2.5-3B student
Student (zero-shot)
12.9
28.1
8.2
29.1
1.7
29.7
Teacher (zero-shot)
71.4
16.6
72.4
17.0
16.5
28.4
OPD
40.4±3.6
21.7
36.3±3.4
22.9
6.3±0.9
29.1
Guided-OPD
22.3±2.8↓ 18.1
26.0↑ 4.3
18.2±2.1↓ 18.1
26.9↑ 4.0
1.0±0.7↓ 5.3
29.9↑ 0.8
Table 2: Results on ALFWorld with a Qwen2.5-7B-RL teacher and a Qwen2.5-3B student.
Figure 3: Gap-adaptive trajectory curriculum. (a) Clock pace relative to fixed F2B ( v=1 ). (b) Trajectory termination modes as the accessible horizon expands.
Method
Valid Seen
Valid Unseen
Hard
SR ↑
Rounds ↓
SR ↑
Rounds ↓
SR ↑
Rounds ↓
Fixed F2B Curriculum
37.7±1.9
23.4
40.9±1.7
23.4
16.4±2.9
28.0
Last-Turn Proposal
37.4±1.9
24.0
45.1±3.0
23.6
13.7±1.3
28.5
w/o Guidance Internalization
36.6±2.0
23.0
44.8±2.8
21.9
13.4±1.9
28.1
w/o Token Focusing
39.4±3.7
23.7
43.1±3.4
23.3
19.7±1.8
28.4
Reconsideration Prompt
39.9±2.5
23.5
44.9±3.6
22.9
19.5±1.6
28.1
Table 3: Ablation study on ALFWorld with a Qwen3-32B teacher → Qwen3-1.7B student.
Table 4: Training configuration across the three benchmarks. Batch-size field names follow TCOD.
Parameter
Value
Maximum generation tokens
4,096
Temperature
0.4
Top- p / top- k / min- p
1.0 / −1 / 0.0
History length
2 steps
Maximum interactions: ALFWorld
30
Maximum interactions: ScienceWorld
30
Appendix
Table 5: Evaluation decoding and environment execution settings.
Figure 5: WebShop trajectory curriculum. Top: Qwen3-1.7B; bottom: Qwen3-4B. Left: recorded clock pace (dots) and its moving average, relative to fixed F2B ( v=1 ). Right: trajectory-termination proportions; dashed lines mark the first rollouts using the full horizon.
Figure 6: WebShop future validation. Top: Qwen3-1.7B; bottom: Qwen3-4B. Left: guided and unguided future scores over the same complete pairs. Right: unsmoothed future-gain distributions, normalized by the number of complete pairs in each run. Green and gray denote positive and negative gains, respectively.
Figure 7: ScienceWorld trajectory curriculum. Top: Qwen3-1.7B; bottom: Qwen3-4B. Left: clock pace relative to fixed F2B ( v=1 ). Right: trajectory-termination proportions. Dashed lines mark the first rollouts using the full K=30 horizon.
Figure 8: ScienceWorld future validation. Top: Qwen3-1.7B; bottom: Qwen3-4B. Left: guided and unguided future scores. Right: future-gain distributions over valid paired forks, normalized separately for each run. Green and gray denote positive and negative gains.
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.
On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge-OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3-32B teacher to Qwen3-1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge-OPD.
Chishui Chen, Yaoyou Fan, Te Sun +11
Meituan LongCat Interaction · Peking University · Shanghai Jiao Tong University +6
On-policy distillation (OPD) leverages dense teacher rewards to enhance reasoning models. However, scaling OPD to long-horizon tasks exposes a critical flaw: as the student's generated prefix inevitably diverges from the teacher's thought process, the teacher's dense reward loses local exploitability. Continuing to generate and evaluate tokens on these ``drifted'' trajectories not only degrades reward quality but also incurs massive computational waste. To address this, we introduce \textbf{Prune-OPD}, a framework that dynamically aligns training budgets with supervision quality. By continuously monitoring the local compatibility between student and teacher predictions (e.g., via top-k overlap), Prune-OPD detects prefix-drift events in real time. Upon detecting severe drift, it monotonically down-weights subsequent unreliable rewards and triggers dynamic rollout truncation. This allows the training process to halt futile generation and reallocate compute strictly to reliable teacher supervision. Across diverse teacher-student combinations, Prune-OPD consistently aligns computation with supervision reliability. When prefix drift makes dense teacher rewards unreliable, it reduces training time by 37.6%--68.0% while preserving, and often improving, performance on challenging benchmarks (AMC, AIME, HMMT). When student-teacher compatibility remains high, it automatically preserves long-context supervision by expanding the training window. These results suggest that Prune-OPD improves OPD not by blindly shortening rollouts, but by reallocating computation toward locally exploitable teacher rewards.
Zhicheng Yang, Zhijiang Guo, Yifan Song +5
1The Hong Kong University of Science and Technology (Guangzhou) · 2The Hong Kong University of Science and Technology · 3MBZUAI +2