On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision compatibility can vary substantially across trajectories, making a single rollout horizon difficult to match their heterogeneous reliable lengths: an overly short horizon may truncate useful supervision, while an overly long one wastes computation beyond reliable regions. Our key insight is that the trajectory-specific reliability boundary need not be predicted before generation. By viewing reliability as the first-passage of accumulated low teacher--student compatibility events, the boundary is inherently unknown before sampling, yet whether it has been reached can be determined exactly from the observed prefix. Building on this insight, we propose Flash-OPD, which shifts from rollout-horizon control to adaptive trajectory-level boundary verification. Flash-OPD interleaves cached generation with teacher verification and independently stops each trajectory according to its observed compatibility events. To reduce verification overhead, the recent event rate is used only to schedule the next verification point, while the actual stopping decision always relies on the exact cumulative count. This separation prevents estimation errors from causing premature termination while enabling efficient verification during generation. Extensive experiments across diverse datasets and teacher--student settings show that Flash-OPD achieves 2.2×--7.5× speedups over standard OPD while maintaining or improving accuracy.
Figures & tables
Figure 1 : From a shared rollout horizon to adaptive verification. Existing OPD acceleration methods commit to a shared horizon before trajectory sampling, causing premature truncation or redundant generation when usable lengths vary across trajectories; Flash-OPD instead verifies reliability during generation and retires each trajectory at its own boundary, ensuring reliability and efficiency.
Length decision
Allocation
Method
Driven by
Granularity
Spectrum
All seqs. trained
Stacks on a batch ctrl.
POPD ( Zhang et al., 2026b )
training step
batch
single
✓
✗
TOPD ( Zhang et al., 2026b )
training step
batch
single
✓
✗
Prefix OPD ( Zhang et al., 2026a )
training step
batch
single
✓
✗
ADWIN ( Liang et al., 2026 )
gradient cosine
batch
single
✓
✗
Prune-OPD ( Yang et al., 2026c )
top- k overlap
batch
single
✓
✗
Table 1: Comparison of rollout-control methods for OPD. Driven by : what the length decision consumes. Granularity : the unit the decision is made per. Spectrum : the set of realized lengths within one step. All seqs. trained : whether every sampled sequence contributes to the loss. Stacks on a batch ctrl. : can run under an independent batch-level window controller instead of replacing it.
Figure 2 : Heterogeneous and evolving reliability lengths expose the mismatch of fixed and adaptive batch-shared horizons (left and middle), while our method Flash-OPD substantially reduces excess rollout budget without prematurely truncating reliable prefixes in the illustrated setting (right).
Figure 3 : A schematic diagram of our rollout case, which adaptively adjusts verification positions according to the estimated local event rate and stops once the exact reliability boundary is reached.
Table 2: Overall performance-efficiency comparison under different benchmarks and settings.
Figure 4 : Mean wall-clock seconds per training step.
Figure 5 : Wall-clock training dynamics of vanilla OPD and our Flash-OPD across four teacher-student model pairs and five benchmarks. We perform a benchmark evaluation every 20 steps.
Figure 6 : Transferability of Flash-OPD across evaluation distributions and model generations.
Figure 7 : With a fixed student, Flash-OPD incurs smaller step-time increases than OPD while maintaining stable accuracy and consistently short rollouts as teacher size grows from 4B to 32B.
Figure 8 : Sensitivity of Flash-OPD to the drift threshold η , allowance B , and step discount ρ .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Domain
Split used
#Problems
Evaluation budget
Training corpus
DAPO-Math-17K ( Yu et al., 2026 )
Competition mathematics
train (all)
17,917
—
In-domain evaluation
AMC23
Competition mathematics
test
83
31744 tok., mean@16
AIME24
Competition mathematics
test
30
31744 tok., mean@16
AIME25
Competition mathematics
test
30
31744 tok., mean@16
Appendix
Table 3: The training corpus and the five in-domain benchmarks are competition mathematics; the four out-of-domain benchmarks are not, and GPQA is not mathematics at all.
Student
Teacher
Relationship
Max resp.
k0
W
kmin
Lmax
R1-Distill-Qwen-1.5B
JustRL-DeepSeek-1.5B
Same-size RL-improved copy of the student
12288
1024
1024
512
4400
R1-Distill-Qwen-1.5B
R1-Distill-Qwen-7B
Larger same-family reasoning model
12288
128
128
16
12288
Qwen3-1.7B-Base
Qwen3-4B (Non-thinking)
Stronger same-family teacher, base student
8192
128
128
16
8192
Qwen3-4B-Base
Qwen3-4B (Non-thinking)
Same-size post-trained teacher, base student
12288
128
128
16
12288
Appendix
Table 4: The four teacher–student pairs and the controller each one runs.
Parameter
Value
Optimisation
Train batch size (prompts per step)
64
PPO mini-batch size
64 (one update per step)
Samples per prompt n
4
Learning rate
1×10−6 , constant
Micro-batch per GPU
1 , dynamic batching enabled
Appendix
Table 5: Parameter setting of our method Flash-OPD .
Method
Step-time attribution (s/step)
Tokens vs. time
Rollout
Teacher
Log-prob
Update
Step
Length
Token ↓
Step ↓
R1-1.5B ← JustRL-1.5B
OPD
70.3
11.2
10.5
31.5
132.3
6122
–
–
Flash-OPD
26.2
14.2
0.0
14.7
60.0
2708
2.3 ×
2.20 ×
R1-1.5B ← R1-7B
OPD
88.8
40.4
13.3
42.4
194.3
7961
–
–
Appendix
Table 6: Where the speedup comes from: mean wall-clock seconds per training step.
Pair
Retained length E^
Stopping rule (of 256)
Verification cost
min
mean
max
allowance
Lmax
EOS
rounds
amplif.
overshoot
R1-1.5B ← JustRL-1.5B
830
2708
3509
130
96
30
2.64
1.54 ×
231 (8.5%)
R1-1.5B ← R1-7B
335
585
1052
256
0
0
6.23
2.41 ×
26 (4.5%)
Qwen3-1.7B-Base ← Qwen3-4B
126
304
672
243
0
13
4.98
2.23 ×
14 (4.5%)
Qwen3-4B-Base ← Qwen3-4B
293
569
2117
250
0
6
6.23
2.19 ×
62 (10.9%)
Appendix
Table 7: Mechanism counters, averaged over all training steps.
Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Meituan +1