ΔWAM: Distilling Action Tangent Fields into World Action Models
Authors: Ke Wu, Hanwen Huang, Bo Gu, Kaizhao Zhang, Xiangting Meng, Yupeng Zheng, Zijun Xu, Jieru Zhao, +1 more
Organizations: Fudan University · TARS Robotics · ShanghaiTech University · Institute of Automation, Chinese Academy of Sciences · Shanghai Jiao Tong University
World Action Models (WAM) improve robot policies by augmenting sparse action supervision with dense future prediction. However, much of the predictable future is dominated by appearance and scene persistence rather than action-dependent dynamics. We observe that several recent WAM designs, including optical flow, motion-centric representations, and latent actions, can be understood from a common perspective in which world supervision becomes more efficient as it contains a higher proportion of action-relevant variation. Based on this insight, we introduce Action Tangent Fields, which reformulate world supervision through a local Taylor expansion of how actions induce changes in future dynamics. We represent future dynamics in Residual-VAE space, where the future latent remains recoverable from the current latent and its residual, and use a strong action-conditioned world model (ACWM) to probe the local correspondence between action variations and residual-world variations. This local first-order structure is distilled into the WAM to guide its denoising supervision toward dynamics that are more tightly coupled to action, rather than merely predictable from appearance. Across LIBERO-Plus, RoboTwin, and RoboTwin2.0-Plus, our method consistently improves robustness to lighting, background, camera, layout, and other environmental perturbations. Despite using no large-scale embodied pretraining, it achieves stronger robustness under several distribution shifts than pretrained policies. We further distill multi-step VideoDiT denoising into a single step for efficient inference. Our results suggest that effective WAM supervision should remain information-rich while concentrating its predictive capacity on the directions along which actions change the future.
Figures & tables
Figure 1: Δ WAM under distribution shift . Left: success rate by perturbation type on LIBERO-Plus and RoboTwin 2.0-Plus, with the average over perturbations shown for each method. Right: the corresponding shifts, camera noise, novel viewpoints and layouts, language, lighting, and unseen backgrounds.
Figure 2: Overview of Δ WAM . A Wan VideoDiT-5B is mid-trained to predict the Residual-VAE future Rt+1:t+H . An embodiment-specific ACWM is then queried at K perturbed actions, and its normalized response differences are orthonormalized into a cached basis Gt of locally action-sensitive directions. Post-training keeps the real residual as the target and penalizes its projection onto Gt , one VideoDiT pass at tv=0 conditions four ActionDiT Euler steps at inference.
Model
Original
Camera
Robot
Lang.
Light
BG
Noise
Layout
Total
VLAs
π0 ( 2025b )
94.2
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
π0 rerun ( 2025b )
91.3
61.0
40.8
63.5
89.3
84.1
80.1
76.4
69.4
π0 -FAST ( 2025 )
85.5
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
π0.5 ( 2025a )
96.9
75.4
77.5
85.6
96.9
94.6
89.7
85.7
85.7
OpenVLA-OFT m ( 2025 )
97.6
55.6
21.7
81.0
92.7
91.0
78.6
68.7
67.9
Table 1: LIBERO-Plus evaluation results. We report success rates (%) under the original setting and seven perturbation dimensions Zhang et al. (2026b) . The best result in each column is shown in bold , while the second-best result is underlined .
RoboTwin
RoboTwin 2.0-Plus
Method
Embodied PT.
Clean
Rand.
Avg.
Original
Camera
Robot
Lang.
Light
BG
Noise
Layout
Total
π0 ( 2025b )
✓
65.92
58.40
62.2
–
–
–
–
–
–
–
–
–
π0.5 ( 2025a )
✓
82.74
76.76
79.8
78.4
45.6
27.6
74.4
49.6
71.7
64.9
56.8
58.6
X-VLA ( 2025 )
–
–
–
–
65.6
23.2
65.2
64.4
63.1
58.6
49.7
34.8
53.1
MOTUS ( 2025 )
✓
88.66
87.02
87.8
87.0
21.6
85.0
83.2
84.6
84.4
43.1
82.8
71.5
LingBot-VA ( 2026c )
✓
92.90
91.50
92.2
92.1
28.9
36.2
87.3
89.0
91.3
80.9
87.9
74.2
Table 2: Results on RoboTwin and RoboTwin 2.0-Plus. We report success rates (%) on RoboTwin under the Clean and Randomized settings, and on RoboTwin 2.0-Plus under the original setting and seven perturbation dimensions. For RoboTwin 2.0-Plus, all methods use a single unified model across all tasks. The best result in each column is highlighted in bold , while the second-best result is underlined . Embodied PT.: embodied pretraining. Lang.: Language. BG: Background.
Figure 3: Real-world rollouts across embodiments. Left: Tianji performs bimanual cooking with dexterous hands under a new background. Right: NERO inserts test tubes with a parallel gripper under new objects layout, numbered frames show the progression of each task.
7-DoF NERO Arm & Gripper
Task
Method
Standard
Layout 1
Layout 2
Red & Green Light
Average
Stack Bowls
π0.5 ( 2025a )
0.33
0.33
0.17
0.50
0.33
FastWAM ( 2026 )
0.50
0.50
0.17
0.17
0.33
Δ WAM (Ours)
0.67
0.50
0.50
0.50
0.54
Insert Tubes
π0.5 ( 2025a )
0.83
0.50
0.50
0.50
0.58
FastWAM ( 2026 )
0.50
0.50
0.33
0.50
0.46
Table 3: Real-world evaluation on NERO and Tianji . Scores are averaged over three trials for each task and condition. The Average column is the mean across the four conditions.
Figure 4: Ablations on LIBERO-Plus. (a) Average success and inference time for different VideoDiT denoising steps. (b) Average success for K=0,4,8 local action perturbations.
World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet explicitly generating future videos at inference introduces substantial computational overhead. Removing future generation improves efficiency, but leaves future dynamics only implicitly encoded in observation features, which can limit robustness under distribution shifts. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Instead of reconstructing the complete future scene, MoWAM models structured robot motion as a compact abstraction of the future, capturing how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing video generation to be removed entirely at inference while retaining an explicit representation of the future. The compact motion representation further enables efficient inference-time scaling by sampling multiple candidates of motion and action pairs and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines. In addition, performance improves as more candidates are explored, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.
Jiayu Wang, Bin Zhu, Yue Yu +1
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China · Singapore Management University, Singapore · Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China
World-Action Models (WAMs) augment robot policies with future visual prediction, but it remains unclear what the visual modality should learn for control. While photorealistic future prediction provides dense supervision, it also incurs substantial computation and can allocate capacity to texture, illumination, and background variations that are only weakly related to action selection. Recent efficient WAM variants suggest that the main benefit of the video branch may not lie in the rendered future itself, but in the control-relevant visual representations induced during training. In this work, we revisit future video prediction from a dynamic-centric perspective and ask whether an existing RGB-based WAM can be redirected from appearance-dominated reconstruction toward interaction-induced visual dynamics without introducing additional modality-specific predictions or online inputs at deployment. We propose DC-WAM, a dynamic-centric WAM framework that redistributes supervision and computation in the RGB video branch. At the supervision level, DC-WAM combines temporal-difference flow matching with trajectory-guided weighting, emphasizing dense temporal changes and localized regions where the gripper, manipulated objects, and contact areas move. At the reasoning level, DynaRoute predicts token-wise dynamic relevance and converts it into an attention bias, guiding the model toward control-relevant future tokens. Experiments in simulation and on real-world manipulation tasks show that DC-WAM consistently improves policy performance, especially under out-of-distribution perturbations in lighting, object appearance, and background texture.
Haoyuan Ji, Lingxiang Fan, Shang Su +4
Tsinghua University · Dense-AI · University of Michigan
World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-relevant state transitions are entangled with nuisance variations in texture, illumination, background, and viewpoint. We argue that WAMs should explicitly predict action-relevant future state rather than relying on RGB prediction alone. We introduce DreamWAM, which reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics. During training, DreamWAM combines joint latent denoising of RGB and motion with lightweight gated residual branches for geometry and semantics. Shared attention between VideoDiT and ActionDiT allows the action branch to learn from these future-state predictions, while all beyond-RGB supervision branches are disabled at inference and deployment remains RGB-only. Across both no-rollout and joint video-action inference, DreamWAM consistently improves the matched RGB-only baselines on LIBERO, from 97.30% to 98.40% and from 98.00% to 98.90%, respectively. The gains become larger under unseen LIBERO-Plus perturbations, from 51.36% to 63.44% and from 69.16% to 75.47%. The same robustness extends to real-world manipulation, where DreamWAM attains an average success rate of 74.4% across unseen changes in lighting, background, and object layout, compared with 55.6% for Fast-WAM-Joint. These results show that robust world-action learning depends not only on predicting the future, but on representing it in a form that matters for action. The code and models are publicly released at https://github.com/hustvl/DreamWAM.
Shanglin Yuan, Weiheng Zhao, Xin Shi +6
Huazhong University of Science and Technology · D-Robotics · Wuhan University +1