Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-generated trajectories, OPW targets states induced by the student's own decisions, including imperfect actions and recovery situations. We provide a theoretical explanation by connecting on-policy reverse-KL distillation to trajectory-level distribution matching. Under a competent teacher and sufficiently small population distillation loss, this connection yields a lower bound on initial verifier success and a corresponding bound on reward-discovery complexity. For group-relative RLVR, we further characterize when increased success probability produces more reward-informative groups. Together, our findings support on-policy distillation as an effective warmup for agentic RLVR and identify initial reward discovery as a mechanism that can contribute to the observed acceleration.
Figures & tables
Figure 1: OPW accelerates GRPO and improves repeated task success in ALFWorld. Left: Qwen3-4B Seen learning curves for GRPO with and without OPW. GPU-hours include teacher scoring and exclude evaluation. Right: Final Unseen repeated task success (pass^k).
Setting
Evidence
Training study
Work
On-policy warmup
Environment interaction
Theoretical analysis
Action-level interventions
Training cost
Warmup duration
Sequential Beats Joint ( Li et al., 2026a )
✔
✗
✗
✗
✗
✔
RL Starts before RL ( Dong et al., 2026 )
✔
✗
✗
✗
✗
✗
OPDSearch+ ( Ye et al., 2026 )
✔
✔
✔
✗
✔
✗
PRISM ( Wang et al., 2026a )
✔
✗
✗
✗
✗
✗
PEAR ( Zhang et al., 2026 )
✗
✗
✔
✗
✗
✗
Table 1: Related work on warmup for RLVR.
Table 2: Final task performance across three training paths after 100 steps (Qwen3-4B).
Table 3: Final task performance after 40 warmup and 60 GRPO steps (Qwen3-4B).
Table 4: Final repeated task success (pass^8, %) for three training paths at 100 steps (Qwen3-4B).
Policy Split / k
Seen
Unseen
k=8
k=16
k=32
k=64
k=8
k=16
k=32
k=64
ALFWorld
Teacher (32B)
70.9 /27.5
76.4 /21.5
81.0 /15.2
85.7 /10.0
88.3 / 24.2
93.8 / 17.3
97.1 / 11.5
99.3 / 6.7
Base (4B)
51.5/15.8
58.4/13.2
64.6/10.9
69.3/9.3
55.2/10.9
63.1/8.0
70.7/5.9
78.4/4.5
Off-policy: SFT
52.9/12.6
60.0/9.3
66.1/6.7
71.4/4.3
58.7/7.7
67.8/5.2
75.8/3.8
82.8/3.0
Off-policy: SFT-RS
53.2/12.6
60.7/9.1
67.5/6.4
73.6/4.3
59.8/7.0
69.0/3.7
76.8/1.9
84.3/1.5
Table 5: Success coverage and repeated task success after warmup (pass@ k /pass^k, %).
Figure 2: Task performance during warmup and GRPO.
Table 8
Figure 3: Distinct successful action sequences among eight rollouts per task during GRPO.
Figure 4: Repeated task success (pass^8) across OPW durations (Qwen3-1.7B).
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Models and training schedule
Student
Qwen3-1.7B / Qwen3-4B
Teacher
Qwen3-32B
Environments
ALFWorld / ScienceWorld
Main training length
40 warmup + 60 GRPO steps
Shorter warmup comparison
20 warmup + 80 GRPO steps
Appendix
Table 8: Training and evaluation settings for the main warmup comparisons.
Collection
Rollouts/task
Measurements
Original evaluations
8
Learning curves, final task performance and pass^8
New warmup evaluations
64
Success coverage and repeated task success before GRPO
First eight of the new collection
8
Warmup action composition and standard OPW’s initial score in supervision removal
ALFWorld warmup diagnostics
32
Validity and repeated task success on 64 reserved tasks
ScienceWorld warmup diagnostics
8
Complete success and partial-credit reward contrast on 64 Seen tasks
Duration comparisons
8
Final performance under different OPW + GRPO allocations
Appendix
Table 9: Trajectory collections used for evaluation.
Table 10: Final repeated task success (pass^8, %) after 40 warmup and 60 GRPO steps (Qwen3-4B).
Table 11: Early gains and average task performance during 60 GRPO steps after warmup.
Warmup
Seen
Unseen
pass@1
pass^8
pass@1
pass^8
SFT
91.2
72.1
83.3
63.4
SFT-RS
89.1
68.6
83.5
59.0
GRPO
72.3
35.0
69.9
26.1
OPW (Ours)
95.7
85.0
93.3
76.9
Appendix
Table 12: Final task performance and repeated task success for 20-step warmup (ALFWorld 4B).
Figure 5: Task performance during 20 warmup and 80 GRPO steps.
Figure 6: Task performance (a,b) and training dynamics (c–f) during extended training.
Table 13: Average and final task performance over 100 GRPO steps after warmup (ALFWorld 4B).
Seed
Seen
Unseen
R0
ΔR20
Rˉ100
R0
ΔR20
Rˉ100
1
-9.2
+30.4
+13.0
+7.3
+19.5
+20.1
2
-0.8
+17.2
+6.6
+6.4
+9.0
+9.5
Appendix
Table 14: OPW minus GRPO warmup under two training seeds (percentage points).
Table 15: Training cost and cost to reach target Seen success in ALFWorld (GPU-hours).
Objective
All disagreements
All-failure groups
Lowest-probability 20%
SFT
+0.1279
+0.128
+0.290
GRPO
−0.0004
+0.004
+0.017
OPD
+0.1812
+0.183
+0.392
Appendix
Table 16: Teacher-target log-probability changes after one optimizer step (ALFWorld 4B).
Task set
Trajectories
GRPO steps after warmup
0
20
40
60
Seen
Fixed
41.1
79.6
94.4
97.4
OPW (Ours)
46.8
88.6
94.1
96.1
Unseen
Fixed
50.5
79.7
91.7
95.1
OPW (Ours)
56.7
88.2
95.0
96.7
Appendix
Table 17: Task performance after warmup on fixed or resampled trajectories (ALFWorld 4B).
Table 18: GRPO learning after targeted and random supervision removal (ALFWorld 4B).
Table 19: Teacher-target probabilities during teacher-free GRPO (ALFWorld 4B).
Table 20: Complete-action validity and continuation success under Base (ALFWorld 4B).
Warmup
Qwen3-1.7B
Qwen3-4B
Action validity
Repeated actions
Task performance
Action validity
Repeated actions
Task performance
V0
ΔV
L0
ΔL
R0
ΔR
V0
ΔV
L0
ΔL
R0
ΔR
ALFWorld, Seen
Off-policy: SFT
76.7
+6.5
60.1
−35.9
6.1
+39.2
87.4
+5.4
29.3
−9.6
29.6
+54.4
Off-policy: SFT-RS
77.1
+4.0
61.7
−29.5
6.2
+37.6
88.0
+7.5
29.7
−16.1
31.0
+58.0
On-policy: GRPO
71.1
+12.7
38.5
−10.6
20.1
+45.4
90.1
−0.5
22.7
−5.1
56.0
+26.7
Appendix
Table 21: Execution behavior and task performance before and after 60 GRPO steps.
Figure 7: All distinct action sequences among eight rollouts per task during GRPO.
Environment
Size
m
N
Off-policy: SFT
Off-policy: SFT-RS
On-policy: GRPO
OPW (Ours)
ALFWorld
1.7B
2
10
1.72→1.64
1.81→1.54
1.52→1.48
1.51→1.34
4
1
3.00→3.21
2.00→2.60
1.97→2.41
1.00→1.93
4B
2
49
1.69→1.60
1.70→1.50
1.62→1.58
1.67→1.52
4
39
2.53→2.46
2.65→2.08
2.48→2.36
2.54→2.14
ScienceWorld
1.7B
2
1
1.67→1.00
1.73→1.46
1.68→1.00
1.29→1.00
4
1
2.43→1.00
2.60→2.00
2.41→1.00
1.57→1.00
Appendix
Table 22: Distinct action sequences among equal numbers of successful trajectories.
OPW steps
Valid actions
pass@8
pass^8
10
72.9
39.4
0.4
15
85.1
52.4
1.3
20
89.4
69.7
10.8
25
90.4
71.4
13.4
30
91.1
71.2
14.5
35
91.6
73.7
20.7
Appendix
Table 23: Action validity and repeated task success during OPW (ALFWorld 1.7B, n=32 , %).
Figure 8: Success coverage and partial-credit reward contrast during OPW (ScienceWorld 1.7B).
Environment / task set
OPW steps
pass^1
pass^2
pass^4
pass^8
ALFWorld / Seen
20
88.3
83.5
77.2
69.3
40
94.9
91.8
86.9
80.7
ScienceWorld / Unseen
40
36.6
20.1
8.7
4.5
60
32.2
17.8
8.5
3.5
ScienceWorld / Seen
40
51.8
37.4
24.4
10.9
60
56.1
42.7
30.7
20.3
Appendix
Table 24: Repeated task success after 100 total steps (Qwen3-1.7B, n=8 , %).
Task: “clean some pan and put it in countertop.” ALFWorld Unseen, Qwen3-4B.
Base
OPW (Ours)
1
go to cabinet 1
1
go to cabinet 1
2
go to cabinet 2
2
go to cabinet 2
3
open cabinet 2
3
open cabinet 2
4
go to cabinet 3
4
go to cabinet 3
5
go to cabinet 4
5
go to cabinet 4
Appendix
Table 25: Complete action sequences before and after OPW on the pan-cleaning task.
Task: “put a clean egg in microwave.” ALFWorld Unseen, Qwen3-4B.
Off-policy: SFT
OPW (Ours)
1
go to fridge 1
1
go to fridge 1
2
open fridge 1
2
open fridge 1
3
go to cabinet 1
3
go to cabinet 1
4
go to cabinet 2
4
go to cabinet 2
5
open cabinet 2
5
open cabinet 2
Appendix
Table 26: Complete action sequences after SFT and OPW on the egg-cleaning task.
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.
Wenze Lin, Jiale Zhao, Xitai Jiang +5
LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University +2
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post-training itself becomes a bottleneck. We study a weak-to-strong alternative: run RL on a smaller model where rollouts are cheaper, then reuse what that RL run learned to improve a stronger target model. Directly distilling the post-RL weak teacher is not enough, because the teacher's final policy mixes useful RL gains with the limitations of the smaller model. We propose Direct On-Policy Distillation (Direct-OPD), which transfers the teacher's RL-induced policy shift instead. Direct-OPD compares the post-RL teacher with its own pre-RL reference and treats their log-ratio as a dense implicit reward for the student. In plain terms, the checkpoint pair tells us which actions RL made the weak model more or less likely to take, and Direct-OPD applies that signal on the stronger student's own on-policy states. This directly reuses the weak model's RL supervision signal without running sparse-reward RL on the target model. Empirically, Direct-OPD consistently leverages weaker teachers to improve stronger target models; notably, it boosts Qwen3-1.7B from 48.3% to 58.3% on AIME 2024 in just 4 hours on 8 A100 GPUs. It outperforms step-matched direct RL and enables the sequential composition of multiple policy shifts. Our results show that RL outcomes can be reused across model scales as implicit reward signals, not merely as final models to imitate.
Shiyuan Feng, Huan-ang Gao, Haohan Chi +7
1SIA-Lab of Tsinghua AIR and ByteDance Seed · Institute for AI Industry Research (AIR), Tsinghua University · 4Peking University +1
Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only coarse trajectory-level supervision, leaving successful behaviors, recurring mistakes, and incidental choices entangled in the same outcome signal. Existing agentic self-distillation methods enrich sparse supervision with natural-language skills, but skills retrieved externally or extracted from a single trajectory by stronger models may mismatch current experience, exceed the policy's capability, or remain path-specific. We propose Group-Reflective Self-Distillation (GRSD), which derives capability-aligned and outcome-discriminative guidance from the policy's own verified rollouts. For each prompt, the policy reflects on each verified trajectory in an on-policy group, and a stop-gradient snapshot contrasts the resulting reflections from successful and failed rollouts to construct group-level privileged guidance. Conditioned on this guidance, a self-teacher refines turn-level credit assignment by modulating outcome-based advantages while preserving the verifier-determined learning direction. Experiments across multiple agentic environments and model scales demonstrate that GRSD consistently outperforms competitive baselines and generalizes more effectively to unseen tasks.
Binbin Zheng, Zijun Xie, Guanqun Zhao +4
University of Science and Technology of China · 2Baidu Inc. · 3Peking University +3