cs.CLSep 29, 2026

Learning from Think-Mode Advantage via On-Policy Distillation

Authors: Wanqi Ren, Jianxiang Wang, Danxuan Liu, Linyi Ding, Huaixiao Tou

Organizations: ByteDance, China

Abstract

Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reasoning is used during distillation rather than student inference. Uniform ThinkOPD, a natural think-enabled OPD baseline, conditions a fixed teacher on one shared think trace and uniformly distills every sibling student response. Although its prefixes are on-policy, the trace need not follow a route compatible with every complete response: the same privileged trace can induce different teacher-student discrepancies even when responses reach the same outcome. We summarize this interaction with trace-response divergence (TRD) and introduce ThinkOPD, which routes supervision at the response level by combining group-relative reward gain with a TRD-based compatibility proxy. Final response weights are normalized within each rollout group. Across mathematical reasoning and code generation, ThinkOPD outperforms Uniform ThinkOPD in both same-model settings and both cross-model teacher-student pairs, and it exceeds representative rationale and self-distillation baselines in a controlled comparison. Controlled interventions show that outcome benefit and the TRD-based proxy provide complementary routing signals in this setting. Think-enabled OPD provides a controlled setting for studying how teacher advantage becomes transferable along student responses.

Figures & tables

Explore similar work

CardsList
  1. What Does Privileged Information Add to On-Policy Self-Distillation?

    Sep 17, 2026XiuYu Zhang, Wei Chow, Junfeng Fang +2Unsupervised On-Policy Self-DistillationTutors

  2. Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance

    May 29, 2026Yuxuan Jiang, Francis FerraroLLM Reasoning Strategies

  3. Rethinking On-Policy Self-Distillation for Thinking Models

    Jul 6, 2026Simran Kaur, Narutatsu Ri, Yinghui He +2Unsupervised On-Policy Self-DistillationLarge Reasoning Models