Outcome-Guided On-Policy Self-Distillation
Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · Huawei Noah’s Ark Lab · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence
Abstract
On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-token statistics and introduce extra hyperparameters and trade-offs. Based on the advantage formulation in RLVR, we analyze the OPSD objective from the same perspective, incorporating outcome correctness signals. We find that vanilla OPSD imposes insufficient penalties and excessive rewards on incorrect trajectories because it applies a fixed divergence objective regardless of outcome correctness. Furthermore, the reliability of teacher supervision is associated with both trajectory outcome and the cumulative average teacher entropy along the rollout. Based on these observations, we propose Outcome-Guided On-Policy Self-Distillation (OG-OPSD), which dynamically adapts both the divergence objective and distillation position according to binary outcome rewards and the cumulative average teacher entropy. Extensive experiments show that OG-OPSD consistently improves the performance of vanilla OPSD and multiple strong baselines in mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3 models at 1.7B, 4B, and 8B scales, as well as Qwen3-VL-2B.
Figures & tables
| Qwen3-1.7B | Qwen3-4B | Qwen3-8B | ||||||||||
| Method | AIME24 | AIME25 | HMMT25 | Avg. | AIME24 | AIME25 | HMMT25 | Avg. | AIME24 | AIME25 | HMMT25 | Avg. |
| Base | 51.5 | 36.7 | 23.1 | 37.1 | 74.9 | 66.4 | 42.2 | 61.1 | 75.8 | 65.6 | 43.9 | 61.7 |
| SFT | 48.4 | 36.3 | 22.7 | 35.8 | 70.2 | 62.3 | 43.4 | 58.6 | 72.3 | 64.2 | 42.9 | 59.8 |
| GRPO | 51.1 | 38.3 | 23.7 | 37.7 | 75.6 | 68.1 | 44.4 | 62.7 | 76.4 | 68.9 | 46.7 | 64.0 |
| OPSD | 57.2 | 41.1 | 28.8 | 42.3 | 75.6 | 67.9 | 44.5 | 62.6 | 77.8 | 67.5 | 45.8 | 63.7 |
| EOPD | 52.2 | 39.1 | 26.9 | 39.4 | 75.7 | 66.5 | 43.9 | 62.0 | 77.5 | 70.0 | 46.6 | 64.7 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| View | User-message template |
| Student | Problem: {problem} |
| Please reason step by step, and put your final answer within \boxed{}. | |
| Teacher | Problem: {problem} |
| Here is a reference solution to this problem: | |
| === Reference Solution Begin === | |
| {reference completion} |