On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.
Figures & tables
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
System prompt
User prompt
Countdown-3
You are a careful arithmetic solver.
Solve the Countdown arithmetic problem. Using the provided numbers, create an equation that equals the target. You may use the operations + , − , × , and / , and each number must be used exactly once; every division must have an integer result. Show your work in <think>...</think> tags and return the solution as compact, comma-separated equations in <answer>...</answer> tags. For example, <answer>2+3=5,5*4=20</answer> . The question is appended as Numbers: […] and Target: … .
Science
You are a careful solver of multiple-choice science questions.
Given a question and four options, select the correct answer. Reason step by step in <think>...</think> tags and place only the corresponding option letter (A, B, C, or D) in <answer>...</answer> tags. The question and answer choices are appended to this instruction.
MedReason
You are a careful medical reasoning assistant.
Given a medical multiple-choice question and its answer choices, select the correct answer. Reason step by step in <think>...</think> tags and place only the letter corresponding to the correct option in <answer>...</answer> tags; do not repeat the answer text. The question and answer choices are appended to this instruction.
Appendix
Table 1: Task prompts used for teacher SFT and policy-distillation experiments. During SFT, the worked demonstration is supplied as the target completion rather than included in the input prompt.
Teacher
Dataset
Accuracy (%)
Response length (tokens)
Qwen2.5-7B
MedReason
77.5
90.95±28.58
Science
72.0
288.82±90.38
Countdown-3
95.5
168.42±316.40
Llama-3.1-8B
MedReason
85.5
93.74±30.81
Science
67.5
134.65±50.77
Countdown-3
99.0
98.82±78.54
Appendix
Table 2: In-distribution performance of the trained teachers. Results use 200 test examples per dataset, temperature 0.5 , one sampled response per prompt, and a 2,048-token completion limit. Response lengths are reported as mean ± standard deviation.
Hyperparameter
Value
Notes
Optimization steps
150
Reported primary comparisons use the checkpoint at step 150.
Optimizer
Fused AdamW
β1=0.9 , β2=0.999 , and ϵ=10−8 .
Weight decay
0
Learning-rate schedule
Constant with linear warm-up
10 warm-up steps.
Per-device batch size
1
Gradient accumulation
32 steps
Gives an effective batch size of 32 prompts per optimizer step.
Appendix
Table 3: Principal hyperparameters used for distillation training.
Setting
Value
Notes
Optimisation steps
300
Starting from each step-150 distilled checkpoint.
RLVR learning rate
2×10−6
Shared across all 24 runs.
Learning-rate schedule
Cosine
Warm-up over the first 10% of steps (30 steps).
Optimizer
Eight-bit AdamW
Bfloat16 training.
Trainable parameters
LoRA adapters
Rank 256, scaling parameter 256, dropout 0; no bias adaptation.
Backbone weights
Frozen, loaded in four-bit precision
Gradient checkpointing enabled.
Appendix
Table 4: Training settings for RLVR on Countdown-4 after Countdown-3 distillation. These settings are identical across rollout policies, KL objectives, and learning rates used in the preceding distillation stage.
Setting
Value
Notes
Student
Qwen2.5-1.5B-Instruct
Fully fine-tuned on Science.
Teacher
Qwen2.5-3B-Instruct
Frozen base instruct model; no task-specific fine-tuning.
Experimental grid
OnPD/OffPD × forward/reverse KL
Learning rates {1×10−5,5×10−5} .
Random seeds
42
One run per condition.
Teacher conditioning
Correct demonstration and Spanish instruction
Used for teacher rollouts and teacher distributions.
Student conditioning
Vanilla Science prompt
Used for student rollouts, likelihoods, and evaluation.
Appendix
Table 5: Experiment-specific settings for the Spanish-style transfer experiment. All unlisted training and evaluation settings follow the standard protocol.
Stage
From
Studied checkpoint
Training method
Model/actor LR
SFT checkpoints (Appendix C)
SFT
meta-llama/Llama-3.1-8B
allenai/Llama-3.1-Tulu-3-8B-SFT
Full-parameter SFT [ Lambert et al., 2024 ]
5×10−6
SFT
meta-llama/Llama-3.1-70B
allenai/Llama-3.1-Tulu-3-70B-SFT
Full-parameter SFT [ Lambert et al., 2024 ]
2×10−6
SFT
Qwen/Qwen2.5-Math-7B
PRIME-RL/Eurus-2-7B-SFT
Math-reasoning SFT [ Cui et al., 2025 ]
2×10−5
RL and preference-optimization checkpoints (Table 1)
DPO
allenai/Llama-3.1-Tulu-3-8B-SFT
allenai/Llama-3.1-Tulu-3-8B-DPO
Offline DPO [ Lambert et al., 2024 ]
5×10−7
Appendix
Table 6: Learning rates used to produce the checkpoints in the main RL comparison and the explicit SFT comparison of Mukherjee et al. [2025] . “From” denotes the checkpoint against which parameter-update sparsity is measured. Rates refer to the language model or policy actor, using the peak rate where a schedule is reported.
On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
Qiwei Di, Xuheng Li, Kaixuan Ji +3
Department of Computer Science, University of California, Los Angeles, CA 90095, USA
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision from teachers. In the standard KL objective of OPD, token-level losses are uniformly averaged, implying equal weights for all tokens. However, we discover that not all tokens are created equal: as student rollouts grow longer, they deviate further from the teacher's distribution, leading to degraded supervision quality at later positions. As a result, OPD using only the first 30% of tokens can perform comparably to using all tokens, whereas OPD using only the last 30% of tokens barely learns anything. In this work, we provide a principled understanding of this issue through the lens of constrained optimization. Based on these insights, we derive Importance-Weighted On-Policy Distillation (IW-OPD), in which the weight assigned to each token depends on the accumulated discrepancy between the student's and teacher's distributions, naturally upweighting earlier tokens and downweighting later ones with larger deviations. We show that IW-OPD converges significantly faster than OPD, with better learning efficiency, and achieves better final performance than standard OPD in both same-size and cross-scale settings, improving performance up to 6.9 points on AIME-2025.
Yan Xie, Sijie Zhu, Tiansheng Wen +2
Xidian University · Georgia Institute of Technology · Amazon AGI SF Lab
On-policy distillation has recently emerged as a promising alternative to standard sequence-level imitation, training a student by scoring its own rollouts with a teacher model. However, we observe ``Off-policy Teacher Decay'' problem in this paradigm: for the later tokens, with student's earlier trajectory as context that is off-policy to the teacher, the teacher's ability to produce a corrective score would decay, and may fall back to token-completion behavior learned in the pre-training stage. We empirically verify this problem, and we propose Early Stopping Rollout (ESR) to fix it: a simple yet effective distillation strategy that simply restricts the rollout generation to the first response tokens. We show that ESR both surpasses the full rollout OPD performance across model size, family, tasks and training regime, and exhibit much higher GPU efficiency and training stability, especially under cross model family scenarios. We further investigate the mechanism behind this surprising performance and discovered "Cascading Alignment" and "Sub-mode Commitment" effect of ESR that may explain why it works effectively and even sometimes exceeding the teacher model performance. Besides, we show that this position-based token selection strategy cannot be fully explainable by KL divergence and entropy signals.
Zhou Ziheng, Jiaqi Li, Huacong Tang +2
University of California, Los Angeles · Beijing Institute of General Artificial Intelligence