Self-distillation can improve reasoning without a separately trained, more capable teacher, but its effectiveness depends on how the self-teacher gains an advantage over the student. Conditioning the teacher on reference answers or solutions can provide such an advantage, but this information may be unavailable. Reflection offers a way to derive explicit error diagnoses and revision guidance from self-generated attempts, yet existing reflection-based methods often combine it with reference information, rich task feedback, or persistent memory. We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification. Starting from an unsuccessful student rollout, the teacher alternates explicit reflection with renewed attempts until success or the retry budget is exhausted, without reference answers or solutions, external diagnostic feedback, or cross-example memory. Each failed retry informs subsequent reflection, while successful correction provides outcome-level evidence for the potential utility of the resulting teacher context. An outcome-aware selection and weighting strategy distinguishes initially correct, reflection-corrected, and unresolved examples, assigning separate weights to their category-normalized distillation losses. Through on-policy distillation, the student matches the teacher's context-conditioned token-level predictive distributions at prefixes of its own rollouts, transferring the benefits of iterative correction while retaining single-pass inference. Across six benchmarks spanning mathematical reasoning, science question answering, and tool use, ReTeach improves average accuracy over GRPO by 1.39 percentage points.
Figures & tables
Figure 1: Comparison across existing self-improvement methods. Our ReTeach realizes denser supervision without any privileged information from original datasets.
Figure 2: An overview of our ReTeach method.
Method
Qwen3-8B
GRPO
OPSD
RLSD
RESD
CEPO
ReTeach
Within
SciKnowEval-Physics
59.84 ±0.32
61.96 ±1.25
{\color[rgb]{0.4,0.4,0.4}--}
61.90 ±0.83
61.09 ±0.54
63.39 ±1.45
64.24 ±1.05
SciKnowEval-Material
64.89 ±1.50
68.91 ±2.03
{\color[rgb]{0.4,0.4,0.4}--}
68.35 ±0.39
68.02 ±0.97
70.83 ±1.35
69.17 ±0.97
ToolUse-ToolAlpaca
57.90 ±0.21
60.85 ±0.30
{\color[rgb]{0.4,0.4,0.4}--}
60.08 ±0.77
60.94 ±0.37
61.86 ±1.62
62.72 ±0.20
Cross
Math-AIME24
76.94 ±0.57
76.48 ±0.47
76.95 ±0.61
76.94 ±0.99
75.83 ±0.27
76.11 ±1.64
78.61 ±0.99
Math-AIME25
66.94 ±0.67
70.37 ±0.69
67.67 ±2.21
69.72 ±0.68
66.11 ±0.70
68.52 ±0.34
70.46 ±1.02
Math-HMMT25
44.72 ±0.41
46.11 ±0.60
46.67 ±0.22
46.39 ±0.68
44.35 ±0.37
44.17 ±1.71
47.78 ±1.42
Table 1: Performance comparison across cross-dataset and within-dataset evaluation. We report the mean ±std over three random seeds. The cross-dataset block reports Avg@12 on three competition-level math benchmarks under the same metric recommended in ( Zhao et al., 2026a ) . The within-dataset block reports Avg@16 on the science and tool-use tasks under the configuration of ( Hübotter et al., 2026 ) . Bold marks the best mean in each row, and underline marks the second best. The within-dataset results of OPSD are marked as {\color[rgb]{0.4,0.4,0.4}--} , because these datasets provide no reference CoT; its Total Average is therefore the mean over the three math benchmarks only.
Figure 3: Ablation on the weight wC of unresolved rollouts for ReTeach on Qwen3-8B, with wA=wB=0.5 fixed. Lines and shaded bands show the mean and standard deviation over three seeds.
Figure 4: Ablation of EMA and reflection designs on SciKnowEval-Physics (avg@16). Lines and bands show the mean and standard deviation over three seeds; the dashed line marks the base model Qwen3-8B.
Figure 5: Teacher entropy (left) and gradient norm (right) on SciKnowEval-Physics during training under tracking, frozen, and EMA teacher updates. Lines and bands show the mean and standard deviation over three seeds.
Figure 6: Reflection-round analysis for ReTeach on SciKnowEval-Physics.
On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedback), offers a promising direction for advancing reasoning capability without a stronger external teacher. Yet in math reasoning the gains are inconsistent, even when the same approach succeeds elsewhere. A pointwise mutual information analysis traces the failure to the privileged context itself: it inflates the teacher's confidence on tokens already implied by the solution (structural connectives, verifiable claims) and deflates it on deliberation tokens ("Wait", "Let", "Maybe") that drive multi-step search. We propose Anti-Self-Distillation (AntiSD), which ascends a divergence between student and teacher rather than descending it: this reverses the per-token sign and yields a naturally bounded advantage in one step. An entropy-triggered gate disables the term once the teacher entropy collapses, completing a drop-in replacement for default self-distillation. Across five models from 4B to 30B parameters on math reasoning benchmarks, AntiSD reaches the GRPO baseline's accuracy in 2 to 10x fewer training steps and improves final accuracy by up to 11.5 points. AntiSD opens a path to scalable self-improvement, where a language model bootstraps its own reasoning through its training signal.
Guobin Shen, Xiang Cheng, Chenxiao Zhao +4
1Xiaohongshu Inc. · Institute of Automation, Chinese Academy of Sciences
Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution. However, because this supervision is generated via uncontrolled sampling, it provides no diagnostic insight into the model's specific errors or corrective guidance for its individual failure patterns. Consequently, the model learns to imitate a privileged distribution rather than receiving fine-grained corrections that pinpoint where and why its reasoning fails. In this paper, we propose Trajectory-Augmented Policy Optimization (TAPO), which advances self-distillation from implicit distributional alignment to explicit trajectory construction. During RL training, the model produces both correct and incorrect rollouts to the same query, and TAPO leverages this contrastive structure to construct micro-reflective corrections, new training trajectories that retain the model's erroneous reasoning up to the point of failure, then insert a natural-language diagnosis and corrected reasoning guided by a correct reference from the same sampling group. Since each trajectory is anchored in the learner's own prefix and solutions, the corrective signal preserves the model's on-policy distribution to a greater extent than the position-wise alignment imposed by KL-based methods. To integrate these trajectories, TAPO introduces difficulty-aware candidate selection at the model's capability boundary and decoupled advantage estimation to prevent gradient contamination. Experiments on AIME 2024, AIME 2025, and HMMT 2025 show that TAPO achieves consistent improvements over GRPO under the same number of training steps. Further analysis demonstrates that TAPO strengthens both first-pass reasoning and error-correction effectiveness.
Zhilin Huang, Hang Gao, Ziqiang Dong +6
Qwen Business Unit of Alibaba · Tsinghua University · Peking University
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil +3
ACI PLC, Bangladesh · University of Dhaka, Bangladesh · BRAC University, Bangladesh +2