cs.LGSep 27, 2026

DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation

Authors: Ao Yu, Weibo Gao, Heng Zhou, Linan Yue, Rui Li, Suyi Liu, Yu Yan, Yizhong Zhang, +1 more

Organizations: University of Science and Technology of China · The Hong Kong Polytechnic University · The University of Hong Kong · Southeast University

Abstract

On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reward-Aligned Reweighting for On-Policy Distillation

    Sep 28, 2026Haofeng Xu, Junwei Su, Lansong Diao +2ReweightingProcess-Level Supervision

  2. Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

    Sep 29, 2026Zhenyu Wang, Tianze Wang, Linjun Zhang +1Efficient On-Policy DistillationMean