Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early training stage where outcome rewards provide little learning signal. As the gap narrows and eventually reverses, however, continued distillation becomes less beneficial and may hinder further improvement. Motivated by this observation, we propose Gap-Adaptive Teacher Scheduling (GATS), which augments the student's RL objective with an OPD term whose weight adapts to the teacher-student performance gap. Specifically, GATS gradually reduces teacher guidance as the student approaches the teacher's reference performance and withdraws it once that reference is reached. This enables GATS to leverage task-trained teachers smaller than the student, since teacher guidance is primarily needed during early training. Across ALFWorld, WebShop, and ScienceWorld with three Qwen2.5 teacher-student configurations, GATS achieves the highest average success rate among the compared methods in all three configurations, improving over reward-only GRPO by 4.37%-11.87% under matched student rollout budgets. Code is available at https://github.com/Ricardo-H/guide-then-let-go.
Figures & tables
Figure 1: Guide early, then let go. Representative training dynamics on WebShop for the Qwen2.5-3B → 7B configuration. (a) OPD reduces all-failure rollout groups during the sparse-reward cold start. (b) After teacher withdrawal, GATS continues to improve beyond the teacher reference, whereas fixed-weight OPD plateaus near it. Dashed lines indicate the teacher reference MT , and dotted lines mark the withdrawal step.
Figure 2: Overview of GATS. The student generates trajectory groups that are used to compute both the GRPO loss and the teacher-guided OPD loss. GATS adaptively weights the OPD loss according to the teacher–student capability gap and combines it with the GRPO loss to update the student. As the gap narrows, the OPD weight decreases to zero, after which the teacher is permanently withdrawn and training continues with GRPO alone.
Figure 3: OPD gain over GRPO versus the teacher–student performance gap. The observed gain diminishes as the gap narrows and becomes negative near parity, motivating gap-adaptive supervision and eventual teacher withdrawal.
ALFWorld
WebShop
ScienceWorld
Avg. SR
Method
ID
OOD
Eval
ID
OOD
Teacher: Qwen2.5-1.5B
53.65
60.42
63.80
12.76
13.80
40.89
Student: Qwen2.5-7B
Prompt-only
14.84
13.02
0.26
10.94
7.03
9.22
GRPO
63.02
73.70
61.20
38.02
28.91
52.97
GRPO + OPD
55.21
46.09
61.20
18.49
14.58
39.11
Table 1: Final success rates (%) on ALFWorld, WebShop, and ScienceWorld. Each entry is averaged over three evaluation runs. Avg. SR denotes the unweighted mean of the five reported metrics. Within each student block, the best and second-best results in each column are shown in bold and underlined, respectively.
Figure 4: Training dynamics for Qwen2.5-3B → 7B on ALFWorld, WebShop, and ScienceWorld. Top: training success rate. Bottom: all-failure group rate. Horizontal dashed lines denote the teacher reference MT ; vertical dotted lines mark GATS teacher withdrawal at tw . Teacher guidance accelerates early learning, while GATS continues to improve after withdrawal.
Schedule
ALF
Web
Sci.
Avg.
Different decay shapes, N=80
Linear
77.86
68.75
44.01
63.54
Cosine
71.35
71.09
44.01
62.15
Step
69.79
65.10
36.72
57.20
Linear decay with varying N
Linear, N=80
77.86
68.75
44.01
63.54
Table 2: Comparison of OPD-weight decay schedules with a Qwen2.5-1.5B teacher and Qwen2.5-7B student. ALF and Sci. denote the ID splits of ALFWorld and ScienceWorld, respectively, while Web denotes the evaluation split of WebShop. All results are averaged over three evaluation runs.
Figure 5: Training success versus GPU-hours for the Qwen2.5-3B → 7B configuration, including teacher costs. Vertical dotted lines mark the GRPO 150-update budgets; horizontal dashed lines indicate the teacher reference MT .
Environment
Budget
GRPO
GRPO+OPD
GATS
MT
Δ (GATS − GRPO)
(GPU-h)
(pp)
ALFWorld
76.4
71.9
73.6
84.9
74.4
+ 13.1
WebShop
39.5
64.0
58.8
77.1
58.6
+ 13.2
ScienceWorld
46.3
27.8
39.7
46.0
32.5
+ 18.2
Table 3: Smoothed training success rate (%) at the GRPO 150-update compute budget for each environment. MT is the teacher’s training-split success rate used as the withdrawal reference. Bold indicates the best method in each environment.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Environment
Train
ID
OOD
Eval
ALFWorld
3,553
140
134
–
WebShop
6,410
–
–
500
ScienceWorld
3,322
1,661
1,684
–
Appendix
Table 4: Task-pool sizes for training and evaluation. Dashes indicate splits not used in the reported evaluation.
Setting
Value
Training updates
150
GRPO group size
8 rollouts per prompt
Prompts per batch
16
Optimizer
AdamW
Learning rate
1×10−6
Weight decay
0.01
Appendix
Table 5: Core training hyperparameters shared across student runs.
Setting
ALFWorld
WebShop
ScienceWorld
Max prompt length
2,048
4,096
6,000
Max response length
512
1,024
1,024
PPO mini-batch size (configuration)
256
64
256
Discount γ
0.95
1.0
1.0
Max turns
50
15
30
Appendix
Table 6: Environment-specific training settings.
Teacher–student gap (pp)
OPD gain over GRPO (pp)
+41.3
+14.1
+31.7
+6.0
+19.0
+5.2
+14.3
+3.4
+9.6
+2.2
−0.8
−3.0
Appendix
Table 7: Task-performance-gap diagnostic on ALFWorld. The gap is teacher SR minus student SR before the continuation. OPD gain is the final SR difference between the matched GRPO+OPD and GRPO continuations. Both quantities are reported in percentage points (pp).
Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.
Yibin Huang, Xinming Xu, Conghui Zhu
Faculty of Computing, Harbin Institute of Technology · Tsinghua University
On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our empirical analysis reveals a supervision-benefit mismatch: large gaps can be benign, while small gaps can be outcome-critical. Teacher-student gaps capture differences at the current turn, whereas the benefit of teacher guidance depends on how the current student interacts with the environment afterward. The student may still succeed despite choosing an action that differs from the teacher's, while a teacher-preferred action may lead to a state from which the student cannot complete the task. Local gaps alone are therefore not enough to determine whether teacher guidance benefits the current student. Effective supervision should instead emphasize guidance that the current student can translate into better final task outcomes. Accordingly, we propose Outcome-Guided On-Policy Distillation (OG-OPD), which applies trajectory-relative weighting to teacher supervision and calibrates these weights using final task outcomes from paired student continuations. This calibration selectively strengthens supervision on the student's original trajectories at turns where teacher guidance benefits the current student. Across ALFWorld, ScienceWorld, and WebShop, OG-OPD consistently outperforms baselines under diverse settings. It improves task success rates by 3.6-17.7 percentage points over vanilla OPD and by up to 7.0 percentage points over the strongest baseline.
Tong Zhang, Zhou Liu, Yihao Liu +8
Qwen Large Model Application Team, Alibaba · Peking University · University of Chinese Academy of Sciences +1
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from sparse reward signals and loses gradients entirely when all responses within a group receive identical rewards. On-policy distillation (OPD) offers a natural remedy by providing dense, token-level supervision from a teacher model. However, naively combining GRPO with OPD leads to degraded performance, due to three underlying causes: not all samples benefit from distillation; fitting too quickly to the teacher undermines the exploratory capacity of RL; and OPD's advantages are asymmetric, suppressing most tokens. To address these challenges, we propose RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), which applies distillation selectively and precisely where it matters most. At the sample level, OPD is restricted to negative zero-variance prompts with each sample weighted by the teacher's confidence score. At the token level, distillation targets only tokens with high student entropy or large teacher-student divergence. We further augment training with SFT on correct trajectories generated by the teacher model, injecting positive gradient signals where RL yields none. Experiments demonstrate that RSTG substantially outperforms naive GRPO+OPD by +4.02% on math and +3.05% on code.
Zhuowen Han, Jinwei Xiao, Zhengxi Lu +9
TJUNLP Lab, School of Computer Science and Technology, Tianjin University · Meituan Longcat Team