Organizations: LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University · SMS, Peking University · NLPLab, Tsinghua University
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.
Figures & tables
Figure 1: Left: Overview of OPDVR. The ReLU gating mechanism ensures correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, while preserving the teacher’s distributional guidance. Right: Results on AIME24, AIME25, and AMC under the same-architecture setting (Qwen3-4B ← Qwen3-4B-RL), reported as avg@16 accuracy.
Figure 2: Training pipeline of OPDVR.
Method
AIME24
AIME25
AMC
MATH500
Minerva
OlympiadBench
Avg.
Student (Qwen3-4B)
24.0
15.8
60.8
80.9
27.6
42.9
42.0
Teacher (Qwen3-4B-RL)
36.0
29.0
65.9
87.0
35.4
49.3
50.4
Sampled-Token OPD
34.2
26.0
63.1
85.5
31.6
46.5
47.8
Top-64 OPD
34.6
23.5
62.0
85.0
32.2
46.8
47.4
OPD+GRPO
33.1
24.2
64.4
85.2
32.9
46.5
47.7
RLSD
28.7
20.0
62.1
82.8
31.2
44.0
44.8
Table 1: Results on same-architecture distillation (Qwen3-4B ← Qwen3-4B-RL). All models are evaluated with avg@16 accuracy.
Method
AIME24
AIME25
AMC
MATH500
Minerva
OlympiadBench
Avg.
Student (Qwen3-1.7B-Base)
4.1
1.7
23.2
48.9
8.9
17.1
17.3
Teacher (Qwen3-4B-Base-RL)
10.6
13.1
40.3
74.2
17.2
30.0
30.9
Sampled-Token OPD
6.5
2.1
24.8
59.1
11.5
21.6
20.9
Top-64 OPD
8.5
3.3
26.4
60.1
10.7
21.4
21.7
OPD+GRPO
8.3
1.7
23.2
58.0
10.2
21.8
20.5
RLSD
6.6
2.9
28.7
59.3
10.4
21.6
21.6
Table 2: Results on cross-architecture distillation (Qwen3-1.7B-Base ← Qwen3-4B-Base-RL). All models are evaluated with avg@16 accuracy.
Method
AIME24
AIME25
AMC
MATH500
Minerva
OlympiadBench
Avg.
Student (Qwen3-4B)
24.0
15.8
60.8
80.9
27.6
42.9
42.0
Teacher (Qwen3-4B-RL)
36.0
29.0
65.9
87.0
35.4
49.3
50.4
GRPO
28.3
20.8
62.3
83.9
28.9
44.6
44.8
OPD
32.0
31.7
65.6
85.4
28.9
46.6
48.4
Distilled RL
33.1
28.9
66.6
86.3
30.3
46.8
48.7
GRPD (Ours)
34.8
31.7
67.0
85.6
30.5
47.0
49.4
Table 3: Results on Group Relative Policy Distillation (Qwen3-4B ← Qwen3-4B-RL). All models are evaluated with avg@16 accuracy.
Figure 3: Ablation study on the gating mechanism. Left : training-time accuracy reward of OPD, OPDVR, and the inverse-gated variant. Right : average accuracy over six benchmarks.
Method
AIME24
AIME25
AMC
MATH500
Minerva
OlympiadBench
Avg.
Student (Qwen3-4B)
24.0
15.8
60.8
80.9
27.6
42.9
42.0
Teacher (Qwen3-4B-RL)
36.0
29.0
65.9
87.0
35.4
49.3
50.4
OPD
34.2
26.0
63.1
85.5
31.6
46.5
47.8
OPDVR (Ours)
36.9
28.1
64.8
84.7
33.2
47.0
49.1
Inverse-Gated
30.3
21.2
62.3
83.6
27.7
42.8
44.6
Table 4: Ablation study on the gating mechanism. We compare OPD, OPDVR, and an inverse-gated variant where the ReLU gate is applied in the opposite direction.