Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.
Figures & tables
Figure 1: Reinforcement learning versus on-policy distillation. Reinforcement learning learns from sparse, task-grounded environment rewards, whereas on-policy distillation scores the student’s own rollouts with a teacher, providing denser supervision but less direct environment feedback.
Figure 2: Overview of SIPO. SIPO combines sparse environment rewards with dense, contrastive self-teacher credit for fine-grained credit assignment. Left: the agent learns from outcome-level rewards and token-level scores from a self-teacher that contrasts a correct and an incorrect context. Right: unlike GRPO, which assigns a uniform advantage to all tokens, SIPO aims to credit the reasoning steps (in blue) that lead to the correct answer.
Competition Math
Standard Math
Avg.
AIME24
AIME25
AMC23
MATH500
Minerva
Olympiad
Qwen3-8B
26.3
23.4
59.7
73.8
20.2
41.7
40.8
+ GRPO
54.9
40.3
84.7
84.6
29.4
50.1
57.3
+ SDPO
33.4
25.5
63.8
77.6
25.4
43.5
44.9
+ RLSD
49.5
38.2
85.6
84.8
30.9
51.2
56.7
+ SIPO
56.8
41.4
87.8
85.8
32.0
53.0
59.4
Table 1: Comparison of SIPO and baselines on math reasoning benchmarks. All methods train Qwen3-8B on DAPO-Math-17k; we report the best checkpoint with avg@32 on AIME24 and AIME25, avg@8 on AMC23, and avg@1 on MATH500, Minerva Math (Minerva), and OlympiadBench (Olympiad). Best results are in bold and second-best results are underlined .
Chemistry
Physics
Biology
Materials
Tool use
Avg.
1h
5h
1h
5h
1h
5h
1h
5h
1h
5h
Qwen3-8B
41.2
59.2
30.8
58.9
57.5
49.5
+ GRPO
63.3
63.4
63.6
63.6
49.8
49.8
73.9
74.1
60.2
65.7
62.7
+ SDPO
73.2
80.9
66.6
75.6
50.6
56.8
72.1
78.4
68.0
68.5
69.1
+ RLSD
73.2
74.6
67.5
70.3
44.0
44.0
66.7
69.0
62.4
62.5
63.4
+ SIPO
76.9
79.6
74.6
79.2
46.8
58.7
77.1
79.5
64.6
67.4
70.4
Table 2: Comparison of SIPO and baselines on science QA and tool use. We report the best avg@16 within 1h and 5h of wall-clock training. See Table 6 for off-policy GRPO and Table 5 for results with Olmo3-7B. Best results are in bold and second-best results are underlined .
Figure 3: Reward curves of SIPO, SDPO, and GRPO on Chemistry and Materials. SIPO converges faster and reaches a higher final reward than both baselines by combining sparse outcome rewards with dense, token-level credit from the contrastive self-teacher.
Task
Holdout tasks
LCBv6
IFEval
ArenaHard-v2 (hard prompt)
ArenaHard-v2 (creative writing)
MMLU-Pro
Avg. (holdout)
Qwen3-8B
27.9
83.9
14.0
13.7
62.5
43.5
Self-teacher SFT
42.7
83.7
11.2
8.9
61.9
41.4
GRPO
41.2
82.2
12.0
10.8
62.3
41.8
SDPO
48.8
83.2
12.3
11.1
62.9
42.4
RLSD
55.4
83.0
12.3
11.8
62.5
42.4
Table 3: Comparison of SIPO and baselines on code generation and holdout tasks. Holdout benchmarks only measure generalization beyond LCBv6, and self-teacher SFT trains on responses generated by the self-teacher. SIPO achieves the highest LCBv6 accuracy while maintaining strong generalization. Best results are in bold and second-best results are underlined .
Figure 4: Training dynamics of SDPO on DAPO-Math-17k with Qwen3-8B. Left: training reward and entropy. Right: gradient norm and response length. Reward, entropy, and response length decrease steadily, while the gradient norm grows.
Figure 5: Training reward of SIPO, GRPO, and RLSD on Math with Qwen3-8B. SIPO improves from the start of training and stays above GRPO throughout, while RLSD stalls near its initial reward until its teacher weight decays to zero at step 50.
Competition Math
Standard Math
Avg.
AIME24
AIME25
AMC23
MATH500
Minerva
Olympiad
Qwen3-8B
26.3
23.4
59.7
73.8
20.2
41.7
40.8
+ GRPO
54.9
40.3
84.7
84.6
29.4
50.1
57.3
+ GRPO w/ analytic KL
25.6
18.2
66.5
76.2
27.5
41.6
42.6
+ RLSD
49.5
38.2
85.6
84.8
30.9
51.2
56.7
+ RLSD w/ contrastive
56.7
39.8
86.8
83.2
31.2
51.9
58.3
Table 4: Ablation of how the self-teacher is used on math with Qwen3-8B. GRPO w/ analytic KL adds an analytic reverse KL toward a privileged self-teacher. RLSD rescales the advantage with token-level evidence from a self-teacher conditioned on the reference answer. RLSD w/ contrastive replaces this evidence with the contrast between a correct and an incorrect teacher context. SIPO adds the contrastive evidence to the advantage of every rollout, including uniformly failed groups. Best results are in bold and second-best results are underlined .
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Prompt template for the SIPO self-teacher. The positive prompt xi+ contains the reference answer and the negative prompt xi− the incorrect answer ai− . Without a reference answer (code generation), the slot holds a successful or a failed sibling rollout ( Section A.1 ).
Chemistry
Physics
Biology
Materials
Tool use
Avg.
1h
5h
1h
5h
1h
5h
1h
5h
1h
5h
Olmo3-7B
22.8
37.7
16.2
36.7
39.3
30.5
+ GRPO
51.4
57.5
62.7
62.7
49.8
49.8
73.3
73.5
56.8
60.6
59.8
+ SDPO
68.0
80.0
59.9
66.1
48.0
52.8
73.7
79.1
60.8
62.1
65.1
+ SIPO
71.2
75.9
66.0
66.4
53.1
55.0
76.9
77.0
58.8
64.7
66.5
Appendix
Table 5: Comparison of SIPO and baselines on science QA and tool use with Olmo3-7B. The setting follows Table 2 . Best results are in bold and second-best results are underlined .
Chemistry
Physics
Biology
Materials
Tool use
Avg.
1h
5h
1h
5h
1h
5h
1h
5h
1h
5h
Qwen3-8B
41.2
59.2
30.8
58.9
57.5
49.5
+ GRPO (off-policy)
65.9
74.5
63.8
72.7
35.1
59.9
74.3
77.1
64.9
67.7
65.6
+ GRPO (on-policy)
63.3
63.4
63.6
63.6
49.8
49.8
73.9
74.1
60.2
65.7
62.7
Olmo3-7B
22.8
37.7
16.2
36.7
39.3
30.5
+ GRPO (off-policy)
39.7
56.7
55.3
63.3
35.6
55.8
70.9
75.0
56.4
65.0
57.4
Appendix
Table 6: GRPO with off-policy and on-policy updates. GRPO (off-policy) takes four mini-batch gradient steps per generation batch, while GRPO (on-policy) takes one step, matching the setting of all methods in Table 2 . We report the best avg@16 within 1h and 5h of wall-clock training. The better result for each base model is in bold .
Figure 7: Reward curves of SIPO and SDPO on Physics and Biology with Qwen3-8B. SIPO reaches higher reward earlier in training and maintains a more stable upward trajectory.
Chemistry
Physics
Biology
Materials
Tool use
Avg.
1h
5h
1h
5h
1h
5h
1h
5h
1h
5h
Qwen3-8B
SIPO (analytic KL)
78.8
79.2
73.8
81.5
55.3
56.0
74.0
76.4
68.3
68.3
71.2
SIPO (contrastive)
76.9
79.6
74.6
79.2
46.8
58.7
77.1
79.5
64.6
67.4
70.4
Olmo3-7B
SIPO (analytic KL)
79.1
81.1
68.1
73.3
53.5
54.0
77.6
79.3
64.0
67.3
69.7
Appendix
Table 7: Analytic KL versus contrastive token-level advantage on science QA and tool use. SIPO (analytic KL) adds an analytic KL divergence toward a self-teacher conditioned on successful sibling rollouts to the GRPO objective; SIPO (contrastive) is the final method. The setting follows Table 2 . The better result for each base model is in bold .
Task
Holdout tasks
LCBv6
IFEval
ArenaHard-v2 (hard prompt)
ArenaHard-v2 (creative writing)
MMLU-Pro
Avg. (holdout)
SIPO (analytic KL)
58.5
83.8
13.5
13.9
63.7
43.7
SIPO (contrastive)
56.3
83.5
12.7
12.0
62.6
42.7
Appendix
Table 8: Analytic KL versus contrastive token-level advantage on LCBv6 and holdout tasks with Qwen3-8B. The setting follows Table 3 . The better result is in bold .
Figure 8: Top- k token agreement and entropy of the analytic KL variant of SIPO versus SDPO on Physics and Biology. The analytic KL variant exhibits slightly lower top- k agreement, suggesting less over-optimization and greater diversity, while maintaining comparable generation entropy.
Figure 9: Reward and entropy of the analytic KL variant of SIPO and two ablations on Chemistry. SIPO-M conditions the self-teacher on an incorrect instead of a successful rollout, and SIPO-OP uses the actor’s current weights as the teacher. SIPO-M plateaus at a lower reward, and SIPO-OP suffers from entropy collapse and fails to learn.
Figure 10: Qualitative example of the earlier self-teacher on Chemistry.
Figure 11: Qualitative example of the earlier self-teacher on Physics.
Figure 12: Qualitative example of the earlier self-teacher on Biology.
Figure 13: Qualitative example of the earlier self-teacher on Materials. The self-teacher is given a mistake instead of a correct demonstration as privileged information.