An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Organizations: University of North Carolina at Chapel Hill · Brigham Young University · NVIDIA
Abstract
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.
Figures & tables
| Model | Method | MATH-500 | Minerva | Olympiad | AMC23 | AIME24 | AIME25 | ||||||
| Avg@16 | Pass@16 | Avg@16 | Pass@16 | Avg@16 | Pass@16 | Avg@16 | Pass@16 | Avg@16 | Pass@16 | Avg@16 | Pass@16 | ||
| KD | 79.73 | 94.00 | 38.51 | 58.09 | 45.97 | 67.85 | 51.20 | 77.11 | 17.71 | 33.33 | 17.08 | 36.67 | |
| OPD | 79.31 | 95.20 | 35.73 | 58.46 | 44.64 | 67.85 | 48.49 | 79.97 | 17.92 | 30.00 | 15.83 | 40.00 | |
| EOPD | 80.50 | 95.00 | 37.78 | 59.56 | 46.36 | 68.00 | 51.05 | 80.72 | 17.92 | 36.67 | 17.92 | 40.00 | |
| LSPD | 81.61 | 95.40 | 38.99 | 59.56 | 47.22 | 69.04 | 52.48 | 81.93 | 18.96 | 36.67 | 18.13 | 36.67 | |
| 8B 4B | LSPD-RB | 81.55 | 95.20 | 38.49 | 58.09 | 47.41 | 69.33 | 52.11 | 84.34 | 21.25 | 50.00 | 18.75 | 36.67 |
| Benchmark | Method | Pass@1 | Pass@2 | Pass@4 | Pass@8 | Pass@16 | Pass@32 | Pass@64 |
| AMC23 | KD | 51.81 | 56.63 | 66.27 | 71.08 | 77.11 | 85.54 | 87.95 |
| OPD | 51.81 | 55.42 | 63.86 | 73.49 | 79.97 | 83.13 | 86.75 | |
| EOPD | 51.81 | 60.24 | 72.29 | 75.90 | 80.72 | 84.34 | 87.95 | |
| LSPD | 54.22 | 62.65 | 71.08 | 79.52 | 81.93 | 85.54 | 89.16 | |
| AIME24 | KD | 20.00 | 26.67 | 30.00 | 33.33 | 33.33 | 36.67 | 40.00 |
| OPD | 16.67 | 20.00 | 23.33 | 26.67 | 30.00 | 36.67 | 40.00 |
| Reg. | Method | MATH-500 | Minerva | Olympiad | AMC23 | AIME24 | AIME25 | ||||||
| Avg@16 | Pass@16 | Avg@16 | Pass@16 | Avg@16 | Pass@16 | Avg@16 | Pass@16 | Avg@16 | Pass@16 | Avg@16 | Pass@16 | ||
| w/ Ent. | LSPD | 70.86 | 91.80 | 29.32 | 55.15 | 33.47 | 59.70 | 36.97 | 67.47 | 11.88 | 30.00 | 8.33 | 23.33 |
| LSPD-RB | 70.69 | 91.00 | 29.57 | 54.41 | 33.50 | 61.04 | 36.60 | 67.47 | 11.67 | 36.67 | 8.33 | 23.33 | |
| w/o Ent. | LSPD | 71.14 | 90.60 | 28.65 | 51.84 | 32.91 | 58.52 | 37.12 | 66.27 | 11.46 | 26.67 | 8.54 | 20.00 |
| LSPD-RB | 69.39 | 90.60 | 28.84 | 53.31 | 33.65 | 60.59 | 36.67 | 66.27 | 10.00 | 33.33 | 7.29 | 20.00 | |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | w/o Replay Buffer | w/ Replay Buffer |
| Entropy regularization coefficient | ||
| Squared-loss robustification | Huber | Huber |
| Huber transition threshold | ||
| Optimizer | AdamW | AdamW |
| Learning rate | ||
| Learning-rate schedule | Constant | Constant |
| Method | GPQA | MMLU | WinoGrande | HumanEval |
| KD | 30.13 | 61.24 | 63.85 | 64.63 |
| OPD | 27.68 | 61.14 | 63.61 | 67.07 |
| EOPD | 27.90 | 61.47 | 64.80 | 65.85 |
| LSPD | 29.91 | 61.61 | 64.88 | 67.68 |
| LSPD-RB | 30.58 | 62.03 | 64.64 | 70.12 |