Organizations: LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University · SMS, Peking University · NLPLab, Tsinghua University
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.
Figures & tables
Figure 1: Left: Overview of OPDVR. The ReLU gating mechanism ensures correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, while preserving the teacher’s distributional guidance. Right: Results on AIME24, AIME25, and AMC under the same-architecture setting (Qwen3-4B ← Qwen3-4B-RL), reported as avg@16 accuracy.
Figure 2: Training pipeline of OPDVR.
Method
AIME24
AIME25
AMC
MATH500
Minerva
OlympiadBench
Avg.
Student (Qwen3-4B)
24.0
15.8
60.8
80.9
27.6
42.9
42.0
Teacher (Qwen3-4B-RL)
36.0
29.0
65.9
87.0
35.4
49.3
50.4
Sampled-Token OPD
34.2
26.0
63.1
85.5
31.6
46.5
47.8
Top-64 OPD
34.6
23.5
62.0
85.0
32.2
46.8
47.4
OPD+GRPO
33.1
24.2
64.4
85.2
32.9
46.5
47.7
RLSD
28.7
20.0
62.1
82.8
31.2
44.0
44.8
Table 1: Results on same-architecture distillation (Qwen3-4B ← Qwen3-4B-RL). All models are evaluated with avg@16 accuracy.
Method
AIME24
AIME25
AMC
MATH500
Minerva
OlympiadBench
Avg.
Student (Qwen3-1.7B-Base)
4.1
1.7
23.2
48.9
8.9
17.1
17.3
Teacher (Qwen3-4B-Base-RL)
10.6
13.1
40.3
74.2
17.2
30.0
30.9
Sampled-Token OPD
6.5
2.1
24.8
59.1
11.5
21.6
20.9
Top-64 OPD
8.5
3.3
26.4
60.1
10.7
21.4
21.7
OPD+GRPO
8.3
1.7
23.2
58.0
10.2
21.8
20.5
RLSD
6.6
2.9
28.7
59.3
10.4
21.6
21.6
Table 2: Results on cross-architecture distillation (Qwen3-1.7B-Base ← Qwen3-4B-Base-RL). All models are evaluated with avg@16 accuracy.
Method
AIME24
AIME25
AMC
MATH500
Minerva
OlympiadBench
Avg.
Student (Qwen3-4B)
24.0
15.8
60.8
80.9
27.6
42.9
42.0
Teacher (Qwen3-4B-RL)
36.0
29.0
65.9
87.0
35.4
49.3
50.4
GRPO
28.3
20.8
62.3
83.9
28.9
44.6
44.8
OPD
32.0
31.7
65.6
85.4
28.9
46.6
48.4
Distilled RL
33.1
28.9
66.6
86.3
30.3
46.8
48.7
GRPD (Ours)
34.8
31.7
67.0
85.6
30.5
47.0
49.4
Table 3: Results on Group Relative Policy Distillation (Qwen3-4B ← Qwen3-4B-RL). All models are evaluated with avg@16 accuracy.
Figure 3: Ablation study on the gating mechanism. Left : training-time accuracy reward of OPD, OPDVR, and the inverse-gated variant. Right : average accuracy over six benchmarks.
Method
AIME24
AIME25
AMC
MATH500
Minerva
OlympiadBench
Avg.
Student (Qwen3-4B)
24.0
15.8
60.8
80.9
27.6
42.9
42.0
Teacher (Qwen3-4B-RL)
36.0
29.0
65.9
87.0
35.4
49.3
50.4
OPD
34.2
26.0
63.1
85.5
31.6
46.5
47.8
OPDVR (Ours)
36.9
28.1
64.8
84.7
33.2
47.0
49.1
Inverse-Gated
30.3
21.2
62.3
83.6
27.7
42.8
44.6
Table 4: Ablation study on the gating mechanism. We compare OPD, OPDVR, and an inverse-gated variant where the ReLU gate is applied in the opposite direction.
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure OPD, pure RLVR, and all such joint baselines across logic and math reasoning benchmarks. Beyond the empirical results, we further provide a systematic understanding of this through pass@k behavior, learning dynamics, and parameter updates, yielding a consistent explanation: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while jointly optimizing the two signals causes them to interfere.To provide a practical recipe, we find that the OPD validation score is the key signal for when to switch to RL, and that OPD is a better cold start for RL than SFT. Together, our results establish OPD-then-RL as a simple yet strong way to combine the two methods, turning two entangled signals into complementary stages.
Boyan Li, Bingsen Chen, Chenghao Yang +3
University of Alberta · New York University · NYU Shanghai +3
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.
Zhenrui Yue, Huimin Zeng, Yueqi Wang +8
University of Illinois · UC San Diego · Rochester Institute of Technology +2
On-policy distillation is a powerful way to transfer reasoning ability from a strong teacher to a smaller student: the student samples trajectories from its own policy, and the teacher provides dense token-level supervision on the states the student actually visits. However, this supervision is not always reliable: a teacher can assign high likelihood to plausible but incorrect solutions, or low likelihood to correct student solutions that follow different reasoning paths. Unconditionally distilling the teacher can therefore reinforce bad modes or erase useful student behavior. To address these limitations, we introduce RG-OPD: Reward-Gated On-Policy Distillation that uses verifier feedback to decide when teacher logits should be trusted. RG-OPD bridges sparse verifier rewards and dense teacher logits, preserving token-level supervision while filtering misleading teacher signals. Across reasoning and coding benchmarks, RG-OPD produces stronger distilled students, outperforming both vanilla reverse-KL distillation and the recent TSD-KD baseline. At 1K generation length, RG-OPD improves over reverse-KL by 2.9 points and over TSD-KD by 4.9 points; in the long-generation setting, it improves over the untuned student by 8.2 points. Our code is available at https://github.com/UoC-tail/RG-OPD.
Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi +3