Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation
Authors: Yuhao Sun, Binrui Wu, Zhuoer Xu, Ming Wen, Haoxiang Xu, Bin Chen, Yan Lin, Qianzijing Zhang
Organizations: Ant Group · University of Science and Technology of China · Alibaba International Digital Commerce Group · Peking University · University of Electronic Science and Technology of China
On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future behavior, while consequential guidance may be beyond the current student's reach or fail to persist without privileged input. We formulate long-horizon OPD as hierarchical supervision allocation and argue that productive guidance lies at the intersection of future utility and current learnability. Crucially, this intersection evolves as the student learns. Based on this principle, we propose LENS-OPD, a coarse-to-fine framework that organizes supervision through Locate, Validate, and Refine. Locate adapts trajectory exposure to the student's evolving competence and proposes a candidate decision for intervention. Validate tests whether teacher guidance at that decision improves the same student's subsequent behavior. Refine internalizes the beneficial guided behavior into the deployable policy and concentrates token-level supervision on decisive teacher-student conflicts within the validated turn. These stages are nested: each finer allocation is conditioned on the coarser decision, rather than being optimized as an independent importance score. Experiments across multiple long-horizon agent benchmarks and student-teacher configurations show that LENS-OPD consistently improves task performance over vanilla OPD and strong curriculum- and selection-based baselines. Our results suggest that effective long-horizon distillation requires teaching at the right depth, the right decision, and the right token.
Figures & tables
Figure 1: Overview of LENS-OPD. Productive guidance combines future utility and current learnability, while Locate–Validate–Refine allocates supervision from trajectories to turns and tokens.
Figure 2: Diagnostics of hierarchical supervision allocation. (a) Student support for teacher-preferred tokens at disagreement positions changes markedly with training. (b) Turn-level KL is a weak proxy for the future benefit of guidance, with 37% of valid forks producing non-positive gain. (c) Within future-beneficial turns, one-step corrective response is largest when the student assigns lower support and the teacher expresses higher confidence. Error bars in (c) are 95% bootstrap confidence intervals clustered by fork.
Method
ALFWorld
WebShop
ScienceWorld
SR ↑
Score ↑
SR ↑
Score ↑
SR ↑
Qwen3-32B teacher → Qwen3-1.7B student
Student (zero-shot)
5.8
32.7
5.0
3.7
0.4
Teacher (zero-shot)
49.3
51.8
19.0
44.8
28.7
OPD
36.4±2.3
35.8±13.9
9.6±5.3
42.1±0.3
14.3±0.1
TCOD-F2B
42.6±1.4↑ 6.2
30.4±1.8↓ 5.4
7.4±1.3↓ 2.2
29.4±0.1↓ 12.7
15.4±0.2↑ 1.1
Table 1: Main results on ALFWorld, WebShop, and ScienceWorld under two Qwen3 student scales. SR denotes success rate; higher is better.
Method
Valid Seen
Valid Unseen
Hard
SR ↑
Rounds ↓
SR ↑
Rounds ↓
SR ↑
Rounds ↓
Qwen2.5-7B-RL teacher → Qwen2.5-3B student
Student (zero-shot)
12.9
28.1
8.2
29.1
1.7
29.7
Teacher (zero-shot)
71.4
16.6
72.4
17.0
16.5
28.4
OPD
40.4±3.6
21.7
36.3±3.4
22.9
6.3±0.9
29.1
Guided-OPD
22.3±2.8↓ 18.1
26.0↑ 4.3
18.2±2.1↓ 18.1
26.9↑ 4.0
1.0±0.7↓ 5.3
29.9↑ 0.8
Table 2: Results on ALFWorld with a Qwen2.5-7B-RL teacher and a Qwen2.5-3B student.
Figure 3: Gap-adaptive trajectory curriculum. (a) Clock pace relative to fixed F2B ( v=1 ). (b) Trajectory termination modes as the accessible horizon expands.
Method
Valid Seen
Valid Unseen
Hard
SR ↑
Rounds ↓
SR ↑
Rounds ↓
SR ↑
Rounds ↓
Fixed F2B Curriculum
37.7±1.9
23.4
40.9±1.7
23.4
16.4±2.9
28.0
Last-Turn Proposal
37.4±1.9
24.0
45.1±3.0
23.6
13.7±1.3
28.5
w/o Guidance Internalization
36.6±2.0
23.0
44.8±2.8
21.9
13.4±1.9
28.1
w/o Token Focusing
39.4±3.7
23.7
43.1±3.4
23.3
19.7±1.8
28.4
Reconsideration Prompt
39.9±2.5
23.5
44.9±3.6
22.9
19.5±1.6
28.1
Table 3: Ablation study on ALFWorld with a Qwen3-32B teacher → Qwen3-1.7B student.
Table 4: Training configuration across the three benchmarks. Batch-size field names follow TCOD.
Parameter
Value
Maximum generation tokens
4,096
Temperature
0.4
Top- p / top- k / min- p
1.0 / −1 / 0.0
History length
2 steps
Maximum interactions: ALFWorld
30
Maximum interactions: ScienceWorld
30
Appendix
Table 5: Evaluation decoding and environment execution settings.
Figure 5: WebShop trajectory curriculum. Top: Qwen3-1.7B; bottom: Qwen3-4B. Left: recorded clock pace (dots) and its moving average, relative to fixed F2B ( v=1 ). Right: trajectory-termination proportions; dashed lines mark the first rollouts using the full horizon.
Figure 6: WebShop future validation. Top: Qwen3-1.7B; bottom: Qwen3-4B. Left: guided and unguided future scores over the same complete pairs. Right: unsmoothed future-gain distributions, normalized by the number of complete pairs in each run. Green and gray denote positive and negative gains, respectively.
Figure 7: ScienceWorld trajectory curriculum. Top: Qwen3-1.7B; bottom: Qwen3-4B. Left: clock pace relative to fixed F2B ( v=1 ). Right: trajectory-termination proportions. Dashed lines mark the first rollouts using the full K=30 horizon.
Figure 8: ScienceWorld future validation. Top: Qwen3-1.7B; bottom: Qwen3-4B. Left: guided and unguided future scores. Right: future-gain distributions over valid paired forks, normalized separately for each run. Green and gray denote positive and negative gains.