On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision compatibility can vary substantially across trajectories, making a single rollout horizon difficult to match their heterogeneous reliable lengths: an overly short horizon may truncate useful supervision, while an overly long one wastes computation beyond reliable regions. Our key insight is that the trajectory-specific reliability boundary need not be predicted before generation. By viewing reliability as the first-passage of accumulated low teacher--student compatibility events, the boundary is inherently unknown before sampling, yet whether it has been reached can be determined exactly from the observed prefix. Building on this insight, we propose Flash-OPD, which shifts from rollout-horizon control to adaptive trajectory-level boundary verification. Flash-OPD interleaves cached generation with teacher verification and independently stops each trajectory according to its observed compatibility events. To reduce verification overhead, the recent event rate is used only to schedule the next verification point, while the actual stopping decision always relies on the exact cumulative count. This separation prevents estimation errors from causing premature termination while enabling efficient verification during generation. Extensive experiments across diverse datasets and teacher--student settings show that Flash-OPD achieves 2.2×--7.5× speedups over standard OPD while maintaining or improving accuracy.
Figures & tables
Figure 1 : From a shared rollout horizon to adaptive verification. Existing OPD acceleration methods commit to a shared horizon before trajectory sampling, causing premature truncation or redundant generation when usable lengths vary across trajectories; Flash-OPD instead verifies reliability during generation and retires each trajectory at its own boundary, ensuring reliability and efficiency.
Length decision
Allocation
Method
Driven by
Granularity
Spectrum
All seqs. trained
Stacks on a batch ctrl.
POPD ( Zhang et al., 2026b )
training step
batch
single
✓
✗
TOPD ( Zhang et al., 2026b )
training step
batch
single
✓
✗
Prefix OPD ( Zhang et al., 2026a )
training step
batch
single
✓
✗
ADWIN ( Liang et al., 2026 )
gradient cosine
batch
single
✓
✗
Prune-OPD ( Yang et al., 2026c )
top- k overlap
batch
single
✓
✗
Table 1: Comparison of rollout-control methods for OPD. Driven by : what the length decision consumes. Granularity : the unit the decision is made per. Spectrum : the set of realized lengths within one step. All seqs. trained : whether every sampled sequence contributes to the loss. Stacks on a batch ctrl. : can run under an independent batch-level window controller instead of replacing it.
Figure 2 : Heterogeneous and evolving reliability lengths expose the mismatch of fixed and adaptive batch-shared horizons (left and middle), while our method Flash-OPD substantially reduces excess rollout budget without prematurely truncating reliable prefixes in the illustrated setting (right).
Figure 3 : A schematic diagram of our rollout case, which adaptively adjusts verification positions according to the estimated local event rate and stops once the exact reliability boundary is reached.
Table 2: Overall performance-efficiency comparison under different benchmarks and settings.
Figure 4 : Mean wall-clock seconds per training step.
Figure 5 : Wall-clock training dynamics of vanilla OPD and our Flash-OPD across four teacher-student model pairs and five benchmarks. We perform a benchmark evaluation every 20 steps.
Figure 6 : Transferability of Flash-OPD across evaluation distributions and model generations.
Figure 7 : With a fixed student, Flash-OPD incurs smaller step-time increases than OPD while maintaining stable accuracy and consistently short rollouts as teacher size grows from 4B to 32B.
Figure 8 : Sensitivity of Flash-OPD to the drift threshold η , allowance B , and step discount ρ .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Domain
Split used
#Problems
Evaluation budget
Training corpus
DAPO-Math-17K ( Yu et al., 2026 )
Competition mathematics
train (all)
17,917
—
In-domain evaluation
AMC23
Competition mathematics
test
83
31744 tok., mean@16
AIME24
Competition mathematics
test
30
31744 tok., mean@16
AIME25
Competition mathematics
test
30
31744 tok., mean@16
Appendix
Table 3: The training corpus and the five in-domain benchmarks are competition mathematics; the four out-of-domain benchmarks are not, and GPQA is not mathematics at all.
Student
Teacher
Relationship
Max resp.
k0
W
kmin
Lmax
R1-Distill-Qwen-1.5B
JustRL-DeepSeek-1.5B
Same-size RL-improved copy of the student
12288
1024
1024
512
4400
R1-Distill-Qwen-1.5B
R1-Distill-Qwen-7B
Larger same-family reasoning model
12288
128
128
16
12288
Qwen3-1.7B-Base
Qwen3-4B (Non-thinking)
Stronger same-family teacher, base student
8192
128
128
16
8192
Qwen3-4B-Base
Qwen3-4B (Non-thinking)
Same-size post-trained teacher, base student
12288
128
128
16
12288
Appendix
Table 4: The four teacher–student pairs and the controller each one runs.
Parameter
Value
Optimisation
Train batch size (prompts per step)
64
PPO mini-batch size
64 (one update per step)
Samples per prompt n
4
Learning rate
1×10−6 , constant
Micro-batch per GPU
1 , dynamic batching enabled
Appendix
Table 5: Parameter setting of our method Flash-OPD .
Method
Step-time attribution (s/step)
Tokens vs. time
Rollout
Teacher
Log-prob
Update
Step
Length
Token ↓
Step ↓
R1-1.5B ← JustRL-1.5B
OPD
70.3
11.2
10.5
31.5
132.3
6122
–
–
Flash-OPD
26.2
14.2
0.0
14.7
60.0
2708
2.3 ×
2.20 ×
R1-1.5B ← R1-7B
OPD
88.8
40.4
13.3
42.4
194.3
7961
–
–
Appendix
Table 6: Where the speedup comes from: mean wall-clock seconds per training step.
Pair
Retained length E^
Stopping rule (of 256)
Verification cost
min
mean
max
allowance
Lmax
EOS
rounds
amplif.
overshoot
R1-1.5B ← JustRL-1.5B
830
2708
3509
130
96
30
2.64
1.54 ×
231 (8.5%)
R1-1.5B ← R1-7B
335
585
1052
256
0
0
6.23
2.41 ×
26 (4.5%)
Qwen3-1.7B-Base ← Qwen3-4B
126
304
672
243
0
13
4.98
2.23 ×
14 (4.5%)
Qwen3-4B-Base ← Qwen3-4B
293
569
2117
250
0
6
6.23
2.19 ×
62 (10.9%)
Appendix
Table 7: Mechanism counters, averaged over all training steps.
On-policy distillation (OPD) provides dense teacher supervision along student-generated trajectories, but its online rollout process incurs substantial computational cost, particularly when a few long responses delay batch completion. Existing acceleration methods typically control rollout length using fixed budgets or absolute teacher--student agreement thresholds, which may not reflect learning progress across different models and training stages. We propose Adaptive FastOPD, a progress-aware strategy that expands the rollout horizon only when learning near the current boundary region has plateaued and the current horizon is sufficiently utilized. The former is determined from four teacher--student signals measured relative to their values upon entering each horizon, making expansion responsive to stage-specific progress rather than a predefined step interval or an absolute threshold on the raw agreement signals, while the latter prevents a small number of long responses from triggering increases in rollout cost. Across two teacher--student pairs, Adaptive FastOPD achieves the highest average performance while reducing training time by 49.1--71.2% relative to OPD 15K, and remains robust across a range of hyperparameter settings.
Qian Tan, Huaifei Liang, Xuanyu Zhu +2
University of Science and Technology of China · 3Shanghai Jiao Tong University · 2Shanghai Artificial Intelligence Laboratory
On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories. However, scaling OPD to long-horizon mathematical reasoning exposes a reliability and efficiency problem: standard OPD assigns every sampled candidate the same long rollout budget, even though some trajectories may quickly become weakly aligned with the teacher and provide less useful supervision. Prior analyses suggest that successful OPD depends on local teacher-student compatibility, which can be measured by top-k overlap on student-visited prefixes. When this overlap is low, continuing to generate or train on long suffixes may waste computation and introduce noisy learning signal. To address this, we introduce Prefix-Guided On-Policy Distillation (PG-OPD), a simple rollout-allocation framework that uses fixed-length prefixes to estimate trajectory value before expensive long-horizon generation. PG-OPD first decodes every sampled candidate to the same prefix length, computes teacher-student top-k overlap within an early probe window of that prefix, and selectively continues high-overlap candidates to a fixed long length. Low-overlap candidates stop at the fixed prefix, avoiding unnecessary suffix generation. Across diverse teacher-student combinations on AMC, AIME, and HMMT benchmarks, PG-OPD improves average accuracy by up to 4.80 points while reducing training time by up to 2.46x. These results suggest that prefix-level compatibility provides a practical signal for directing OPD computation toward trajectories that remain learnable from the teacher.
On-policy distillation (OPD) provides dense teacher feedback along student-generated rollouts rather than fixed teacher traces and has emerged as a promising post-training paradigm. However, standard OPD typically generates full rollouts during training, which is computationally expensive and may expose the student to unreliable teacher feedback at late rollout positions, especially during early training. We identify the rollout horizon as a key bottleneck in OPD that substantially impacts training efficiency. Unlike Reinforcement Learning with Verifiable Rewards (RLVR), OPD does not require a final answer reward to provide learning signals. Therefore, full rollouts may not always be necessary for OPD. Motivated by this insight, we propose two simple horizon-control strategies: Progressive OPD (POPD), which gradually expands the rollout horizon during training, and Truncated OPD (TOPD), which permanently performs distillation on reliable truncated rollouts. Experiments on mathematical reasoning show that POPD improves the training efficiency of OPD by up to 3×, while TOPD matches OPD performance using only 10% of the rollout horizon, leading to substantial wall-clock and memory reductions. These results demonstrate that controlling the rollout horizon offers a simple and practical path to more efficient OPD.
Yaocheng Zhang, Jiajun Chai, Yuqian Fu +7
Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Meituan +1