Pretrained world action models (WAMs) provide generalist capabilities across diverse robotic manipulation tasks, yet improving target-task performance to an expert level without degrading pretrained skills remains challenging. We explore on-policy distillation (OPD) for WAMs and introduce WAM-OPD. WAM-OPD inherits the advantage of OPD methods that transfer task-specific teacher knowledge under the student's own induced distribution, rather than directly fitting the student to a narrow task-specific data distribution. However, in closed-loop manipulation, the observation histories change as the student policy evolves, requiring fresh environment rollouts to remain on-policy. Applying OPD to WAMs entails repeated data collection, which is costly even in simulation and often impractical on real robots. To avoid repeated environment rollouts during distillation, we introduce prefix-weighted trajectory replay (PWTR). PWTR uses a fixed trajectory pool composed primarily of initial-student rollouts, supplemented with task-specific teacher rollouts to broaden trajectory coverage. For each trajectory replayed from this pool, PWTR conditions the current policy on successive stored histories to generate fresh denoising paths, along which the task-specific teacher provides supervision. Although these denoising paths are refreshed as the policy evolves, the replayed environment trajectories remain fixed. PWTR therefore reweights per-decision distillation losses using proxy importance weights derived from path scores accumulated over the trajectory prefix preceding each decision to mitigate the resulting shift in the history distribution. Simulated and real-world experiments demonstrate task adaptation without additional environment interaction during distillation. In both settings, WAM-OPD improves target-task performance while retaining near-initial performance on tasks excluded from adaptation.
Figures & tables
Figure 1: Task adaptation strategies for pretrained WAMs. Left: SFT specializes pretrained WAMs for target tasks but risks forgetting pretrained skills. Middle: Online RL couples policy optimization with repeated environment interaction, while sparse rewards can further reduce sample efficiency. Right: WAM-OPD uses PWTR to replay observation histories from a fixed trajectory pool. A frozen teacher provides dense velocity supervision along fresh student denoising paths, enabling adaptation without additional environment interaction during distillation.
Figure 2: Overview of WAM-OPD. Left: Prefix-weighted trajectory replay (PWTR). A fixed pool D contains primarily initial-student rollouts, supplemented by task-specific teacher rollouts. Stored denoising paths are scored by a student snapshot and the fixed initial-student and teacher references. The cumulative prefix compatibility scores Φθˉ(i) , Φθ0(i) , and ΦT(j)(i) yield proxy weights wi . Right: On-policy distillation. Given a stored history hi , the current student generates a fresh denoising path. A frozen teacher supervises the student’s velocity at a detached low-noise query state xˉs on this path. Prefix-weighted losses are averaged over the sampled decision points of one trajectory, with gradients accumulated for one student update. Here sg denotes stop-gradient and B indexes the trajectory’s sampled decision points.
Method
Target tasks
Retention tasks
Average
Stapler
Microwave
Switch
Mug
Laptop
Pillbottle
Dual Bottles
Handover
Target
Retention
Initial student
67.5
45.5
68.5
68.5
99.0
99.0
98.0
90.0
56.50
87.17
Stapler teacher
89.5
33.5
52.5
66.0
87.5
97.0
95.0
56.5
61.50
75.75
Microwave teacher
31.0
60.5
62.0
26.0
98.5
92.5
97.5
2.5
45.75
63.17
π RL
0.5
96.0
25.0
22.0
17.5
22.0
89.0
1.0
48.25 ( −8.25 )
29.42 ( −57.75 )
STEAM
76.0
39.0
61.5
42.5
97.5
95.5
91.5
67.5
57.50 ( +1.00 )
76.00 ( −11.17 )
Table 1: Success rates on RoboTwin. Both teachers supervise WAM-OPD; underlining marks teaching tasks. Parentheses show changes in group averages from the initial student. Bold denotes column maxima excluding Online OPD, which uses additional environment rollouts during distillation.
Task
Student Pool
Teacher Pool
Mixed Pool
w/o prefix
w/ prefix
w/o prefix
w/ prefix
w/o prefix
M-Shuffle
w/ prefix (Ours)
Stapler
75.0
77.5
74.0
72.0
77.0
74.0
80.5
Microwave
57.0
58.0
54.0
55.5
57.5
57.0
59.0
Avg.
66.00
67.75
64.00
63.75
67.25
65.50
69.75
Table 2: Ablation study of PWTR. Success rates are reported over 200 cases per task; M-Shuffle permutes weights among decision points at the same position k in the sampled sequence.
Figure 3: History weight dynamics during distillation. Mean history weights at sampled decision points are shown separately for initial-student and teacher rollouts in the mixed trajectory pool. The dashed line indicates uniform weighting.
Figure 4: Real-world evaluation on Corn, Block, and Button. Left: Experimental scenes for the three manipulation tasks. Right: Success rates of the base student, the corresponding task-specific teachers, and WAM-OPD. WAM-OPD adapts a single student using all three teachers and a fixed trajectory pool, without additional robot rollouts during distillation.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Source
Clean
Randomized
Decision records
Stapler
Initial student
8
32
443
Stapler
Teacher
4
16
166
Microwave
Initial student
8
32
1,802
Microwave
Teacher
4
16
780
Total
24
96
3,191
Appendix
Table 3: Composition of the fixed simulation trajectory pool.
Parameter
Value
Trainable parameters
All student MoT parameters
Optimizer
AdamW
Learning rate
10−6
Adam coefficients (β1,β2)
(0.9,0.95)
Weight decay
0
Maximum gradient norm
1.0
Appendix
Table 4: Hyperparameters for simulation distillation.