Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.
Figures & tables
Figure 1: Reinforcement learning versus on-policy distillation. Reinforcement learning learns from sparse, task-grounded environment rewards, whereas on-policy distillation scores the student’s own rollouts with a teacher, providing denser supervision but less direct environment feedback.
Figure 2: Overview of SIPO. SIPO combines sparse environment rewards with dense, contrastive self-teacher credit for fine-grained credit assignment. Left: the agent learns from outcome-level rewards and token-level scores from a self-teacher that contrasts a correct and an incorrect context. Right: unlike GRPO, which assigns a uniform advantage to all tokens, SIPO aims to credit the reasoning steps (in blue) that lead to the correct answer.
Competition Math
Standard Math
Avg.
AIME24
AIME25
AMC23
MATH500
Minerva
Olympiad
Qwen3-8B
26.3
23.4
59.7
73.8
20.2
41.7
40.8
+ GRPO
54.9
40.3
84.7
84.6
29.4
50.1
57.3
+ SDPO
33.4
25.5
63.8
77.6
25.4
43.5
44.9
+ RLSD
49.5
38.2
85.6
84.8
30.9
51.2
56.7
+ SIPO
56.8
41.4
87.8
85.8
32.0
53.0
59.4
Table 1: Comparison of SIPO and baselines on math reasoning benchmarks. All methods train Qwen3-8B on DAPO-Math-17k; we report the best checkpoint with avg@32 on AIME24 and AIME25, avg@8 on AMC23, and avg@1 on MATH500, Minerva Math (Minerva), and OlympiadBench (Olympiad). Best results are in bold and second-best results are underlined .
Chemistry
Physics
Biology
Materials
Tool use
Avg.
1h
5h
1h
5h
1h
5h
1h
5h
1h
5h
Qwen3-8B
41.2
59.2
30.8
58.9
57.5
49.5
+ GRPO
63.3
63.4
63.6
63.6
49.8
49.8
73.9
74.1
60.2
65.7
62.7
+ SDPO
73.2
80.9
66.6
75.6
50.6
56.8
72.1
78.4
68.0
68.5
69.1
+ RLSD
73.2
74.6
67.5
70.3
44.0
44.0
66.7
69.0
62.4
62.5
63.4
+ SIPO
76.9
79.6
74.6
79.2
46.8
58.7
77.1
79.5
64.6
67.4
70.4
Table 2: Comparison of SIPO and baselines on science QA and tool use. We report the best avg@16 within 1h and 5h of wall-clock training. See Table 6 for off-policy GRPO and Table 5 for results with Olmo3-7B. Best results are in bold and second-best results are underlined .
Figure 3: Reward curves of SIPO, SDPO, and GRPO on Chemistry and Materials. SIPO converges faster and reaches a higher final reward than both baselines by combining sparse outcome rewards with dense, token-level credit from the contrastive self-teacher.
Task
Holdout tasks
LCBv6
IFEval
ArenaHard-v2 (hard prompt)
ArenaHard-v2 (creative writing)
MMLU-Pro
Avg. (holdout)
Qwen3-8B
27.9
83.9
14.0
13.7
62.5
43.5
Self-teacher SFT
42.7
83.7
11.2
8.9
61.9
41.4
GRPO
41.2
82.2
12.0
10.8
62.3
41.8
SDPO
48.8
83.2
12.3
11.1
62.9
42.4
RLSD
55.4
83.0
12.3
11.8
62.5
42.4
Table 3: Comparison of SIPO and baselines on code generation and holdout tasks. Holdout benchmarks only measure generalization beyond LCBv6, and self-teacher SFT trains on responses generated by the self-teacher. SIPO achieves the highest LCBv6 accuracy while maintaining strong generalization. Best results are in bold and second-best results are underlined .
Figure 4: Training dynamics of SDPO on DAPO-Math-17k with Qwen3-8B. Left: training reward and entropy. Right: gradient norm and response length. Reward, entropy, and response length decrease steadily, while the gradient norm grows.
Figure 5: Training reward of SIPO, GRPO, and RLSD on Math with Qwen3-8B. SIPO improves from the start of training and stays above GRPO throughout, while RLSD stalls near its initial reward until its teacher weight decays to zero at step 50.
Competition Math
Standard Math
Avg.
AIME24
AIME25
AMC23
MATH500
Minerva
Olympiad
Qwen3-8B
26.3
23.4
59.7
73.8
20.2
41.7
40.8
+ GRPO
54.9
40.3
84.7
84.6
29.4
50.1
57.3
+ GRPO w/ analytic KL
25.6
18.2
66.5
76.2
27.5
41.6
42.6
+ RLSD
49.5
38.2
85.6
84.8
30.9
51.2
56.7
+ RLSD w/ contrastive
56.7
39.8
86.8
83.2
31.2
51.9
58.3
Table 4: Ablation of how the self-teacher is used on math with Qwen3-8B. GRPO w/ analytic KL adds an analytic reverse KL toward a privileged self-teacher. RLSD rescales the advantage with token-level evidence from a self-teacher conditioned on the reference answer. RLSD w/ contrastive replaces this evidence with the contrast between a correct and an incorrect teacher context. SIPO adds the contrastive evidence to the advantage of every rollout, including uniformly failed groups. Best results are in bold and second-best results are underlined .
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Prompt template for the SIPO self-teacher. The positive prompt xi+ contains the reference answer and the negative prompt xi− the incorrect answer ai− . Without a reference answer (code generation), the slot holds a successful or a failed sibling rollout ( Section A.1 ).
Chemistry
Physics
Biology
Materials
Tool use
Avg.
1h
5h
1h
5h
1h
5h
1h
5h
1h
5h
Olmo3-7B
22.8
37.7
16.2
36.7
39.3
30.5
+ GRPO
51.4
57.5
62.7
62.7
49.8
49.8
73.3
73.5
56.8
60.6
59.8
+ SDPO
68.0
80.0
59.9
66.1
48.0
52.8
73.7
79.1
60.8
62.1
65.1
+ SIPO
71.2
75.9
66.0
66.4
53.1
55.0
76.9
77.0
58.8
64.7
66.5
Appendix
Table 5: Comparison of SIPO and baselines on science QA and tool use with Olmo3-7B. The setting follows Table 2 . Best results are in bold and second-best results are underlined .
Chemistry
Physics
Biology
Materials
Tool use
Avg.
1h
5h
1h
5h
1h
5h
1h
5h
1h
5h
Qwen3-8B
41.2
59.2
30.8
58.9
57.5
49.5
+ GRPO (off-policy)
65.9
74.5
63.8
72.7
35.1
59.9
74.3
77.1
64.9
67.7
65.6
+ GRPO (on-policy)
63.3
63.4
63.6
63.6
49.8
49.8
73.9
74.1
60.2
65.7
62.7
Olmo3-7B
22.8
37.7
16.2
36.7
39.3
30.5
+ GRPO (off-policy)
39.7
56.7
55.3
63.3
35.6
55.8
70.9
75.0
56.4
65.0
57.4
Appendix
Table 6: GRPO with off-policy and on-policy updates. GRPO (off-policy) takes four mini-batch gradient steps per generation batch, while GRPO (on-policy) takes one step, matching the setting of all methods in Table 2 . We report the best avg@16 within 1h and 5h of wall-clock training. The better result for each base model is in bold .
Figure 7: Reward curves of SIPO and SDPO on Physics and Biology with Qwen3-8B. SIPO reaches higher reward earlier in training and maintains a more stable upward trajectory.
Chemistry
Physics
Biology
Materials
Tool use
Avg.
1h
5h
1h
5h
1h
5h
1h
5h
1h
5h
Qwen3-8B
SIPO (analytic KL)
78.8
79.2
73.8
81.5
55.3
56.0
74.0
76.4
68.3
68.3
71.2
SIPO (contrastive)
76.9
79.6
74.6
79.2
46.8
58.7
77.1
79.5
64.6
67.4
70.4
Olmo3-7B
SIPO (analytic KL)
79.1
81.1
68.1
73.3
53.5
54.0
77.6
79.3
64.0
67.3
69.7
Appendix
Table 7: Analytic KL versus contrastive token-level advantage on science QA and tool use. SIPO (analytic KL) adds an analytic KL divergence toward a self-teacher conditioned on successful sibling rollouts to the GRPO objective; SIPO (contrastive) is the final method. The setting follows Table 2 . The better result for each base model is in bold .
Task
Holdout tasks
LCBv6
IFEval
ArenaHard-v2 (hard prompt)
ArenaHard-v2 (creative writing)
MMLU-Pro
Avg. (holdout)
SIPO (analytic KL)
58.5
83.8
13.5
13.9
63.7
43.7
SIPO (contrastive)
56.3
83.5
12.7
12.0
62.6
42.7
Appendix
Table 8: Analytic KL versus contrastive token-level advantage on LCBv6 and holdout tasks with Qwen3-8B. The setting follows Table 3 . The better result is in bold .
Figure 8: Top- k token agreement and entropy of the analytic KL variant of SIPO versus SDPO on Physics and Biology. The analytic KL variant exhibits slightly lower top- k agreement, suggesting less over-optimization and greater diversity, while maintaining comparable generation entropy.
Figure 9: Reward and entropy of the analytic KL variant of SIPO and two ablations on Chemistry. SIPO-M conditions the self-teacher on an incorrect instead of a successful rollout, and SIPO-OP uses the actor’s current weights as the teacher. SIPO-M plateaus at a lower reward, and SIPO-OP suffers from entropy collapse and fails to learn.
Figure 10: Qualitative example of the earlier self-teacher on Chemistry.
Figure 11: Qualitative example of the earlier self-teacher on Physics.
Figure 12: Qualitative example of the earlier self-teacher on Biology.
Figure 13: Qualitative example of the earlier self-teacher on Materials. The self-teacher is given a mistake instead of a correct demonstration as privileged information.
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.
Wenze Lin, Jiale Zhao, Xitai Jiang +5
LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University +2
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on student-generated trajectories, complementing reinforcement learning with sparse outcome rewards. Its self-teacher, derived from the student's current or exponentially averaged parameters and conditioned on additional context, evolves alongside the student and its rollout context distribution. The benefit of modifying this moving target depends on how target--student probability mismatches translate into updates. We propose \emph{Self-Referenced On-Policy Self-Distillation (SR-OPSD)}, which constructs a normalized geometric target from the self-teacher and a frozen initial policy, then minimizes the forward Rényi divergence from this target to the student. The interpolation coefficient controls the self-teacher's contribution, while the Rényi order controls the power weighting of target-to-student probability ratios in the gradient. For fixed contexts and target components, we establish a conditional variational characterization and derive the exact token-logit gradient, revealing how anchoring and projection jointly shape the effective update target. Experiments across scientific reasoning, tool use, mathematical reasoning, and code generation demonstrate strong performance across multiple model families and scales. Ablations further show that reference anchoring can improve or degrade performance depending on the projection objective, supporting the joint design of target construction and projection geometry.
Zhuo Sun, Entong Li, Yanlong Zhao +9
Shanghai University of Finance and Economics · Imperial College London · Independent Researcher +6
GRPO-style RLVR trains reasoning models from multiple on-policy attempts per prompt, but typically uses these attempts only through terminal rewards. We show that a mixed group contains a richer process signal: a correct completion is a self-generated witness of how the current policy can solve the problem, while a wrong completion provides on-policy prefixes where the policy needs correction. We introduce \emph{Self-Supervised On-Policy Distillation} (SSOPD), which distills a teacher distribution conditioned on the shortest correct completion into prefixes of the longest wrong completion. This converts intra-group correct--wrong contrast into dense process supervision without external solution traces. A stopping-time view motivates the shortest-correct / longest-wrong rule as a finite-group approximation to editing persistent failures toward fast-success actions, and a prompt-level frontier weight concentrates the auxiliary loss where correct and wrong branches coexist. Across AIME 2024, AIME 2025, and HMMT 2025, SSOPD improves over GRPO in all nine model-benchmark settings. On Qwen3-8B, it reaches a macro Avg@12 of 65.6, outperforming GRPO by 1.6 points and the solution-conditioned OPSD baseline by 0.8 points. Code will be released at https://github.com/tzq1999/SSOPD.