cs.LGSep 30, 2026

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

Authors: Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu

Organizations: Zhejiang University · Ant Healthcare, Ant Group

Abstract

Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-generated trajectories, OPW targets states induced by the student's own decisions, including imperfect actions and recovery situations. We provide a theoretical explanation by connecting on-policy reverse-KL distillation to trajectory-level distribution matching. Under a competent teacher and sufficiently small population distillation loss, this connection yields a lower bound on initial verifier success and a corresponding bound on reward-discovery complexity. For group-relative RLVR, we further characterize when increased success probability produces more reward-informative groups. Together, our findings support on-policy distillation as an effective warmup for agentic RLVR and identify initial reward discovery as a mechanism that can contribute to the observed acceleration.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On-policy Distillation with Verifiable Reward

    Aug 25, 2026Wenze Lin, Jiale Zhao, Xitai Jiang +5Reinforcement Learning With Verifiable RewardOn-Policy

  2. Weak-to-Strong Generalization via Direct On-Policy Distillation

    Jul 6, 2026Shiyuan Feng, Huan-ang Gao, Haohan Chi +7Offline Reinforcement LearningVerifiable Rewards

  3. Group-Reflective Self-Distillation for Agentic Reinforcement Learning

    Jul 30, 2026Binbin Zheng, Zijun Xie, Guanqun Zhao +4Agentic Reinforcement LearningUnsupervised On-Policy Self-Distillation