cs.LGSep 28, 2026

When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

Authors: Xinke Jiang, Tao Feng, Zhibang Yang, Zhixin Zhang, Weixuan Xu, Haoyu Zhang, Xu Chu

Organizations: National Engineering Research Center of Software Engineering, Peking University, Beijing, China · School of Computer Science, Peking University, Beijing, China · Key Laboratory of High Confidence Software Technologies, Ministry of Education, Beijing, China · Center on Frontiers of Computing Studies, Peking University, Beijing, China

Abstract

Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing dense token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward--distillation training through a neural tangent kernel (NTK) analysis. We introduce the cross-signal NTK KDR(n)K_{DR}(n), a token-level statistic that measures the alignment between reward and distillation gradients at position n. Through this analysis, we identify two failure modes: 1 Magnitude drowning, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to rise despite its explicit inclusion in the training objective; and 2 Localized directional conflict, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token (KDR(n) ⁣< ⁣0K_{DR}(n)\!<\!0). The severity of these effects depends on the optimization regime: the gradient-norm ratio κ ⁣= ⁣∥∇LR∥/∥∇LD∥κ\!=\!\|\nabla\mathcal{L}_R\|/\|\nabla\mathcal{L}_D\| varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the M3 family, which combines magnitude normalization with three strategies...

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

    Sep 21, 2026Huanxin Sheng, Zhiling Ye, Haonan Wang +3Sparse SupervisionStance

  2. On-policy Distillation with Verifiable Reward

    Aug 25, 2026Wenze Lin, Jiale Zhao, Xitai Jiang +5Reinforcement Learning With Verifiable RewardOn-Policy