On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by 9.7 points over vanilla OPD, and enables the smaller student to surpass its larger teacher.
Figures & tables
Figure 1: Illustration of Dr. OPD. Left: vanilla OPD assigns the same weight to every teacher signal, while Dr. OPD adapts the weights according to their benefit to the student. Right: student rollouts receive teacher and outcome-reward signals, whose gradient alignment determines token credits. These credits are converted into weights for distillation, where λ controls the weighting strength.
Figure 2: Empirical gains. Across both distillation setups, Dr. OPD consistently outperforms existing baselines and even produces a student that surpasses its teacher (Full results in Section 5.2 ).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Learning rate
1×10−5
LR schedule
constant
Gradient clipping
1.0
Batch / Mini-batch size
128
Rollouts per prompt K
8
Rollouts per step
1024
Appendix
Table 4: Hyperparameters shared by all methods and both tasks.
Method
Hyperparameters
OPD
None
ExOPD
λ=1.25 ; the initial student as reference policy
OPD+GRPO
Equal weighting (1:1) of the distillation and outcome terms
GRPD
None (default settings of the released method)
Dr. OPD
λ=0.4 ; wmin=0.001 ; wmax=3
Appendix
Table 5: Method-specific hyperparameters. All other settings follow Table 4 .