Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the student's own rollouts. This transfer rewards the policy shift at every state, yet the log-ratio measures only relative change: it can stay fixed even as the probability mass that both checkpoints assign to the student's candidate tokens vanishes. Through an exact construction, we show that the Direct-OPD reward and its update can remain unchanged while the Jensen-Shannon divergence (JSD) and both KL directions between the checkpoints vanish with this mass, and we note that a small JSD bounds how much the teacher's behavior changed. Motivated by this analysis, we propose Selective Supervision for Direct-OPD (S2D-OPD), which ranks student-sampled states by their teacher-reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% of states per response. Across two teacher pairs and four student models ranging from 1.7B to 8B parameters, S2D-OPD improves held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings and matches it in the eighth, without extra forward passes. Our code is available at https://anonymous.4open.science/r/S2D-OPD-8868.
Figures & tables
Figure 1: Overview of S 2 D-OPD. At each student-sampled state, teacher–reference JSD is computed over the student’s top- K candidates, with the remaining probability mass grouped into a residual token vo . The highest-JSD states within each response are retained for the Direct-OPD objective.
Figure 2: Validation accuracy (Avg@32) under the JustRL teacher pair across four students. Dashed lines mark the initial student; dotted lines mark the JustRL-1.5B teacher.
Figure 3: Which states are retained, not how many, determines transfer (Qwen3-1.7B, JustRL teacher pair). (a) Validation accuracy when training on each JSD percentile bin or on a random 10% of states. (b) Gain in peak validation accuracy over the initial student versus the mean training JSD of the retained states; the line is a log-linear fit.
Figure 4: Teacher-pair comparison with top-10% retention across four students. (a) Mean teacher–reference JSD over all valid response states (solid) and retained states (dashed) during S 2 D-OPD training. (b) Held-out Test Avg. gains of S 2 D-OPD over dense Direct-OPD under each teacher pair, in percentage points.
Figure 5: JSD and token log-ratio rank states differently. The state with the largest sampled-token log-ratio in the response is masked, while a larger redistribution of probability mass is retained.
Figure 6: Robustness of selective masking across design choices. (a) Top-10% selection by JSD, reverse KL, and forward KL on Qwen3-1.7B, 4B, and 8B under the JustRL teacher pair. (b) Response-level versus batch-level selection on Qwen3-1.7B under both teacher pairs.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Top-16 token overlap over training for Qwen3-1.7B under the JustRL teacher pair, using JSD deciles or random 10% selection. Panels compare the student with the post-RL teacher (left) and pre-RL reference (right). Overlap is the fraction of shared tokens between top-16 sets.
Figure 8: Sensitivity to the per-response retention ratio for Qwen3-1.7B under the JustRL teacher pair. The four settings retain the top 20%, 15%, 10%, and 5% of positions ranked by teacher–reference JSD. Dashed lines indicate base-model accuracy.
Figure 9: Same sampled token, different mask decisions. The mask filters a calculation on which the teacher and reference agree but retains supervision for a boundary correction.
Figure 10: Retained shifts encode continuation preferences. At two retained states, the teacher shifts probability away from Alternatively ; the later state favors Let , Given , and Since .
Figure 11: A later reconsideration cue does not undo an earlier adverse shift. The mask retains a shift toward an incorrect parity judgment, followed by a teacher preference for reconsideration.
Setting
Training
Evaluation
Framework
verl
–
Hardware
8× NVIDIA H200
–
Global batch size
128
–
Mini-batch size
128
–
Rollout n
4
–
Max. prompt length
1,024
–
Appendix
Table 2: Default training and evaluation configuration.