cs.CLSep 30, 2026

OPSRD: On-Policy Self-Role Distillation

Authors: Weijie Ren, Yanwen Zhang, Hao Li, Zhuolin Qi, Hengyi Zhang, Naibo Wang

Organizations: Zhejiang University · University of Electronic Science and Technology of China · University of Science and Technology of China

Abstract

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at https://github.com/zhansan114514/OPSRD.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On-Policy Self-Distillation without Any Supervision

    Aug 6, 2026Yijiang Li, Bingyang Wang, Yijun Liang +3Unsupervised On-Policy Self-DistillationSelf-Distillation Framework

  2. RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

    Jun 10, 2026Leyi Pan, Shuchang Tao, Yunpeng Zhai +5Unsupervised On-Policy Self-DistillationToken-Level Supervision