cs.LGSep 28, 2026

An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

Authors: Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang

Organizations: University of North Carolina at Chapel Hill · Brigham Young University · NVIDIA

Abstract

We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp O~(log⁡K)\tilde{\mathcal O}(\log K) regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ββ-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

    Jul 30, 2026Jiawei Xu, Minghui Liu, Juzheng Zhang +2Unsupervised On-Policy Self-DistillationRp-Opsd

  2. Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence

    May 13, 2026Xinyu Liu, Kechen Jiao, Chunyang Xiao +10Frictive Policy OptimizationTeacher

  3. On the Geometry of On-Policy Distillation

    Jun 5, 2026Zhennan Shen, Yanshu Li, Qingyu Yin +6SubspaceLLM Reasoning Strategies