Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-generated trajectories, OPW targets states induced by the student's own decisions, including imperfect actions and recovery situations. We provide a theoretical explanation by connecting on-policy reverse-KL distillation to trajectory-level distribution matching. Under a competent teacher and sufficiently small population distillation loss, this connection yields a lower bound on initial verifier success and a corresponding bound on reward-discovery complexity. For group-relative RLVR, we further characterize when increased success probability produces more reward-informative groups. Together, our findings support on-policy distillation as an effective warmup for agentic RLVR and identify initial reward discovery as a mechanism that can contribute to the observed acceleration.
Figures & tables
Figure 1: OPW accelerates GRPO and improves repeated task success in ALFWorld. Left: Qwen3-4B Seen learning curves for GRPO with and without OPW. GPU-hours include teacher scoring and exclude evaluation. Right: Final Unseen repeated task success (pass^k).
Setting
Evidence
Training study
Work
On-policy warmup
Environment interaction
Theoretical analysis
Action-level interventions
Training cost
Warmup duration
Sequential Beats Joint ( Li et al., 2026a )
✔
✗
✗
✗
✗
✔
RL Starts before RL ( Dong et al., 2026 )
✔
✗
✗
✗
✗
✗
OPDSearch+ ( Ye et al., 2026 )
✔
✔
✔
✗
✔
✗
PRISM ( Wang et al., 2026a )
✔
✗
✗
✗
✗
✗
PEAR ( Zhang et al., 2026 )
✗
✗
✔
✗
✗
✗
Table 1: Related work on warmup for RLVR.
Table 2: Final task performance across three training paths after 100 steps (Qwen3-4B).
Table 3: Final task performance after 40 warmup and 60 GRPO steps (Qwen3-4B).
Table 4: Final repeated task success (pass^8, %) for three training paths at 100 steps (Qwen3-4B).
Policy Split / k
Seen
Unseen
k=8
k=16
k=32
k=64
k=8
k=16
k=32
k=64
ALFWorld
Teacher (32B)
70.9 /27.5
76.4 /21.5
81.0 /15.2
85.7 /10.0
88.3 / 24.2
93.8 / 17.3
97.1 / 11.5
99.3 / 6.7
Base (4B)
51.5/15.8
58.4/13.2
64.6/10.9
69.3/9.3
55.2/10.9
63.1/8.0
70.7/5.9
78.4/4.5
Off-policy: SFT
52.9/12.6
60.0/9.3
66.1/6.7
71.4/4.3
58.7/7.7
67.8/5.2
75.8/3.8
82.8/3.0
Off-policy: SFT-RS
53.2/12.6
60.7/9.1
67.5/6.4
73.6/4.3
59.8/7.0
69.0/3.7
76.8/1.9
84.3/1.5
Table 5: Success coverage and repeated task success after warmup (pass@ k /pass^k, %).
Figure 2: Task performance during warmup and GRPO.
Table 8
Figure 3: Distinct successful action sequences among eight rollouts per task during GRPO.
Figure 4: Repeated task success (pass^8) across OPW durations (Qwen3-1.7B).
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Models and training schedule
Student
Qwen3-1.7B / Qwen3-4B
Teacher
Qwen3-32B
Environments
ALFWorld / ScienceWorld
Main training length
40 warmup + 60 GRPO steps
Shorter warmup comparison
20 warmup + 80 GRPO steps
Appendix
Table 8: Training and evaluation settings for the main warmup comparisons.
Collection
Rollouts/task
Measurements
Original evaluations
8
Learning curves, final task performance and pass^8
New warmup evaluations
64
Success coverage and repeated task success before GRPO
First eight of the new collection
8
Warmup action composition and standard OPW’s initial score in supervision removal
ALFWorld warmup diagnostics
32
Validity and repeated task success on 64 reserved tasks
ScienceWorld warmup diagnostics
8
Complete success and partial-credit reward contrast on 64 Seen tasks
Duration comparisons
8
Final performance under different OPW + GRPO allocations
Appendix
Table 9: Trajectory collections used for evaluation.
Table 10: Final repeated task success (pass^8, %) after 40 warmup and 60 GRPO steps (Qwen3-4B).
Table 11: Early gains and average task performance during 60 GRPO steps after warmup.
Warmup
Seen
Unseen
pass@1
pass^8
pass@1
pass^8
SFT
91.2
72.1
83.3
63.4
SFT-RS
89.1
68.6
83.5
59.0
GRPO
72.3
35.0
69.9
26.1
OPW (Ours)
95.7
85.0
93.3
76.9
Appendix
Table 12: Final task performance and repeated task success for 20-step warmup (ALFWorld 4B).
Figure 5: Task performance during 20 warmup and 80 GRPO steps.
Figure 6: Task performance (a,b) and training dynamics (c–f) during extended training.
Table 13: Average and final task performance over 100 GRPO steps after warmup (ALFWorld 4B).
Seed
Seen
Unseen
R0
ΔR20
Rˉ100
R0
ΔR20
Rˉ100
1
-9.2
+30.4
+13.0
+7.3
+19.5
+20.1
2
-0.8
+17.2
+6.6
+6.4
+9.0
+9.5
Appendix
Table 14: OPW minus GRPO warmup under two training seeds (percentage points).
Table 15: Training cost and cost to reach target Seen success in ALFWorld (GPU-hours).
Objective
All disagreements
All-failure groups
Lowest-probability 20%
SFT
+0.1279
+0.128
+0.290
GRPO
−0.0004
+0.004
+0.017
OPD
+0.1812
+0.183
+0.392
Appendix
Table 16: Teacher-target log-probability changes after one optimizer step (ALFWorld 4B).
Task set
Trajectories
GRPO steps after warmup
0
20
40
60
Seen
Fixed
41.1
79.6
94.4
97.4
OPW (Ours)
46.8
88.6
94.1
96.1
Unseen
Fixed
50.5
79.7
91.7
95.1
OPW (Ours)
56.7
88.2
95.0
96.7
Appendix
Table 17: Task performance after warmup on fixed or resampled trajectories (ALFWorld 4B).
Table 18: GRPO learning after targeted and random supervision removal (ALFWorld 4B).
Table 19: Teacher-target probabilities during teacher-free GRPO (ALFWorld 4B).
Table 20: Complete-action validity and continuation success under Base (ALFWorld 4B).
Warmup
Qwen3-1.7B
Qwen3-4B
Action validity
Repeated actions
Task performance
Action validity
Repeated actions
Task performance
V0
ΔV
L0
ΔL
R0
ΔR
V0
ΔV
L0
ΔL
R0
ΔR
ALFWorld, Seen
Off-policy: SFT
76.7
+6.5
60.1
−35.9
6.1
+39.2
87.4
+5.4
29.3
−9.6
29.6
+54.4
Off-policy: SFT-RS
77.1
+4.0
61.7
−29.5
6.2
+37.6
88.0
+7.5
29.7
−16.1
31.0
+58.0
On-policy: GRPO
71.1
+12.7
38.5
−10.6
20.1
+45.4
90.1
−0.5
22.7
−5.1
56.0
+26.7
Appendix
Table 21: Execution behavior and task performance before and after 60 GRPO steps.
Figure 7: All distinct action sequences among eight rollouts per task during GRPO.
Environment
Size
m
N
Off-policy: SFT
Off-policy: SFT-RS
On-policy: GRPO
OPW (Ours)
ALFWorld
1.7B
2
10
1.72→1.64
1.81→1.54
1.52→1.48
1.51→1.34
4
1
3.00→3.21
2.00→2.60
1.97→2.41
1.00→1.93
4B
2
49
1.69→1.60
1.70→1.50
1.62→1.58
1.67→1.52
4
39
2.53→2.46
2.65→2.08
2.48→2.36
2.54→2.14
ScienceWorld
1.7B
2
1
1.67→1.00
1.73→1.46
1.68→1.00
1.29→1.00
4
1
2.43→1.00
2.60→2.00
2.41→1.00
1.57→1.00
Appendix
Table 22: Distinct action sequences among equal numbers of successful trajectories.
OPW steps
Valid actions
pass@8
pass^8
10
72.9
39.4
0.4
15
85.1
52.4
1.3
20
89.4
69.7
10.8
25
90.4
71.4
13.4
30
91.1
71.2
14.5
35
91.6
73.7
20.7
Appendix
Table 23: Action validity and repeated task success during OPW (ALFWorld 1.7B, n=32 , %).
Figure 8: Success coverage and partial-credit reward contrast during OPW (ScienceWorld 1.7B).
Environment / task set
OPW steps
pass^1
pass^2
pass^4
pass^8
ALFWorld / Seen
20
88.3
83.5
77.2
69.3
40
94.9
91.8
86.9
80.7
ScienceWorld / Unseen
40
36.6
20.1
8.7
4.5
60
32.2
17.8
8.5
3.5
ScienceWorld / Seen
40
51.8
37.4
24.4
10.9
60
56.1
42.7
30.7
20.3
Appendix
Table 24: Repeated task success after 100 total steps (Qwen3-1.7B, n=8 , %).
Task: “clean some pan and put it in countertop.” ALFWorld Unseen, Qwen3-4B.
Base
OPW (Ours)
1
go to cabinet 1
1
go to cabinet 1
2
go to cabinet 2
2
go to cabinet 2
3
open cabinet 2
3
open cabinet 2
4
go to cabinet 3
4
go to cabinet 3
5
go to cabinet 4
5
go to cabinet 4
Appendix
Table 25: Complete action sequences before and after OPW on the pan-cleaning task.
Task: “put a clean egg in microwave.” ALFWorld Unseen, Qwen3-4B.
Off-policy: SFT
OPW (Ours)
1
go to fridge 1
1
go to fridge 1
2
open fridge 1
2
open fridge 1
3
go to cabinet 1
3
go to cabinet 1
4
go to cabinet 2
4
go to cabinet 2
5
open cabinet 2
5
open cabinet 2
Appendix
Table 26: Complete action sequences after SFT and OPW on the egg-cleaning task.