cs.AISep 28, 2026

Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation

Authors: Yugu Li, Zehong Cao, Peizhen Li, Yang Zhang, Siyi Hu, Jianglin Qiao

Organizations: School of CSIT, Adelaide University, Adelaide, SA 5000, Australia · CAIAA, University of North Texas, Denton, TX 76203, USA · School of EECMS, Curtin University, Bentley, WA 6102, Australia · ACFR, The University of Sydney, Camperdown, NSW 2050, Australia

Abstract

RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6% of tokens and yielding a 1.5×\times reduction in token credit magnitude.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

    Aug 13, 2026Zechuan Wang, Siyuan Lu, Hongxuan Zhang +3Credit AssignmentReinforcement Learning With Verifiable Reward

  2. Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards

    Jun 17, 2026Yingyu Shan, Yuhang Guo, Zihao Cheng +7Reinforcement Learning With Verifiable RewardCredit Assignment

  3. Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

    Jul 30, 2026Qiangqiang He, Zhongheng Wu, ZiJian WangCredit AssignmentReinforcement Learning With Verifiable Reward