Pretrained world action models (WAMs) provide generalist capabilities across diverse robotic manipulation tasks, yet improving target-task performance to an expert level without degrading pretrained skills remains challenging. We explore on-policy distillation (OPD) for WAMs and introduce WAM-OPD. WAM-OPD inherits the advantage of OPD methods that transfer task-specific teacher knowledge under the student's own induced distribution, rather than directly fitting the student to a narrow task-specific data distribution. However, in closed-loop manipulation, the observation histories change as the student policy evolves, requiring fresh environment rollouts to remain on-policy. Applying OPD to WAMs entails repeated data collection, which is costly even in simulation and often impractical on real robots. To avoid repeated environment rollouts during distillation, we introduce prefix-weighted trajectory replay (PWTR). PWTR uses a fixed trajectory pool composed primarily of initial-student rollouts, supplemented with task-specific teacher rollouts to broaden trajectory coverage. For each trajectory replayed from this pool, PWTR conditions the current policy on successive stored histories to generate fresh denoising paths, along which the task-specific teacher provides supervision. Although these denoising paths are refreshed as the policy evolves, the replayed environment trajectories remain fixed. PWTR therefore reweights per-decision distillation losses using proxy importance weights derived from path scores accumulated over the trajectory prefix preceding each decision to mitigate the resulting shift in the history distribution. Simulated and real-world experiments demonstrate task adaptation without additional environment interaction during distillation. In both settings, WAM-OPD improves target-task performance while retaining near-initial performance on tasks excluded from adaptation.
Figures & tables
Figure 1: Task adaptation strategies for pretrained WAMs. Left: SFT specializes pretrained WAMs for target tasks but risks forgetting pretrained skills. Middle: Online RL couples policy optimization with repeated environment interaction, while sparse rewards can further reduce sample efficiency. Right: WAM-OPD uses PWTR to replay observation histories from a fixed trajectory pool. A frozen teacher provides dense velocity supervision along fresh student denoising paths, enabling adaptation without additional environment interaction during distillation.
Figure 2: Overview of WAM-OPD. Left: Prefix-weighted trajectory replay (PWTR). A fixed pool D contains primarily initial-student rollouts, supplemented by task-specific teacher rollouts. Stored denoising paths are scored by a student snapshot and the fixed initial-student and teacher references. The cumulative prefix compatibility scores Φθˉ(i) , Φθ0(i) , and ΦT(j)(i) yield proxy weights wi . Right: On-policy distillation. Given a stored history hi , the current student generates a fresh denoising path. A frozen teacher supervises the student’s velocity at a detached low-noise query state xˉs on this path. Prefix-weighted losses are averaged over the sampled decision points of one trajectory, with gradients accumulated for one student update. Here sg denotes stop-gradient and B indexes the trajectory’s sampled decision points.
Method
Target tasks
Retention tasks
Average
Stapler
Microwave
Switch
Mug
Laptop
Pillbottle
Dual Bottles
Handover
Target
Retention
Initial student
67.5
45.5
68.5
68.5
99.0
99.0
98.0
90.0
56.50
87.17
Stapler teacher
89.5
33.5
52.5
66.0
87.5
97.0
95.0
56.5
61.50
75.75
Microwave teacher
31.0
60.5
62.0
26.0
98.5
92.5
97.5
2.5
45.75
63.17
π RL
0.5
96.0
25.0
22.0
17.5
22.0
89.0
1.0
48.25 ( −8.25 )
29.42 ( −57.75 )
STEAM
76.0
39.0
61.5
42.5
97.5
95.5
91.5
67.5
57.50 ( +1.00 )
76.00 ( −11.17 )
Table 1: Success rates on RoboTwin. Both teachers supervise WAM-OPD; underlining marks teaching tasks. Parentheses show changes in group averages from the initial student. Bold denotes column maxima excluding Online OPD, which uses additional environment rollouts during distillation.
Task
Student Pool
Teacher Pool
Mixed Pool
w/o prefix
w/ prefix
w/o prefix
w/ prefix
w/o prefix
M-Shuffle
w/ prefix (Ours)
Stapler
75.0
77.5
74.0
72.0
77.0
74.0
80.5
Microwave
57.0
58.0
54.0
55.5
57.5
57.0
59.0
Avg.
66.00
67.75
64.00
63.75
67.25
65.50
69.75
Table 2: Ablation study of PWTR. Success rates are reported over 200 cases per task; M-Shuffle permutes weights among decision points at the same position k in the sampled sequence.
Figure 3: History weight dynamics during distillation. Mean history weights at sampled decision points are shown separately for initial-student and teacher rollouts in the mixed trajectory pool. The dashed line indicates uniform weighting.
Figure 4: Real-world evaluation on Corn, Block, and Button. Left: Experimental scenes for the three manipulation tasks. Right: Success rates of the base student, the corresponding task-specific teachers, and WAM-OPD. WAM-OPD adapts a single student using all three teachers and a fixed trajectory pool, without additional robot rollouts during distillation.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Source
Clean
Randomized
Decision records
Stapler
Initial student
8
32
443
Stapler
Teacher
4
16
166
Microwave
Initial student
8
32
1,802
Microwave
Teacher
4
16
780
Total
24
96
3,191
Appendix
Table 3: Composition of the fixed simulation trajectory pool.
Parameter
Value
Trainable parameters
All student MoT parameters
Optimizer
AdamW
Learning rate
10−6
Adam coefficients (β1,β2)
(0.9,0.95)
Weight decay
0
Maximum gradient norm
1.0
Appendix
Table 4: Hyperparameters for simulation distillation.
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by their reliance on costly expert demonstrations. We challenge this by asking whether future supervision for WAMs must originate from target-task expert trajectories. In this paper, we propose Vid2WAM, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student. Given an observation and language instruction, Vid2WAM distills supervision through two complementary channels: task-conditioned future rollouts directly supervise the student's future prediction branch, while an inverse dynamics model recovers embodiment-specific pseudo-actions for action learning. To robustly integrate synthetic and real supervision, we introduce source-aware residual action adaptation that learns source-specific corrections around a shared action backbone and mitigates interference from noisy pseudo-actions. During inference, both the video teacher and inverse dynamics model are discarded, leaving only the WAM student for efficient deployment. Simulation and real-world experiments demonstrate that Vid2WAM improves novel-task generalization and data efficiency under limited expert demonstrations while preserving low-latency inference.
Chenhao Qiu, Ruixiang Wang, Runyi Zhao +7
Fudan University · The Chinese University of Hong Kong, Shenzhen · Shanghai Innovation Institute
World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.
Xiangcheng Zhan, Zirui Chen, Yicheng Zhao +2
Harbin Institute of Technology · Dalian University of Technology · Southern University of Science and Technology
World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.
Rui Wang, Xiangyu Wang, Donglin Yang +4
Southern University of Science and Technology · Shenzhen Loop Area Institute · The University of Hong Kong