World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution, whereas wrist cameras move with the end effector, mixing local interaction changes with viewpoint shifts and self-occlusion. These contrasting predictive demands suggest that the two views may benefit from different future objectives. We introduce FutureDuet, which retains both views for control, while allowing each visual stream to receive a different future objective. For the main view, future RGB models task evolution, while interaction masks and robot skeletons focus supervision on task objects and robot motion. For the wrist stream, future latent prediction models short-horizon interaction changes without requiring pixel-level reconstruction. ActionDiT jointly reads the resulting Task State and Interaction State, combining scene-level progress with close-range interaction evidence. All auxiliary prediction modules are training-only, adding no inference overhead. FutureDuet achieves 94.2% clean and 94.1% randomized success on RoboTwin50 and 99.2% average success on LIBERO. The improvements are most pronounced on six RoboTwin50 tasks that require precise interaction, averaging gains of 9.2% and 12.8% over Fast-WAM in clean and randomized settings. Controlled studies further show complementary gains from separating the wrist pathway and designing future supervision separately for the two views.
Figures & tables
Figure 1: Decoupling future supervision across views. (a) A common multi-view WAM design concatenates observations and applies one future objective across views. (b) FutureDuet separately models task evolution from the main view and interaction evolution from wrist views; the resulting states jointly support action generation.
Figure 2: Overview of FutureDuet. Different training-only future objectives shape the Task State and Interaction State. ActionDiT reads both through the Duet Interface; auxiliary predictors are removed at deployment.
Table 3
Figure 3: RoboTwin50 performance varies substantially across tasks. (a) Number of tasks reaching at least 90% and 95% success under clean and randomized evaluation. (b) Gains over Fast-WAM on the six tasks that remain below 90% for FutureDuet in both settings. Gains are measured in percentage points (pp). OM, HM, TS, MSP, PCB, and SBT denote Open Microwave, Hanging Mug, Turn Switch, Move Stapler Pad, Place Can Basket, and Stack Bowls Three, respectively.
Figure 4: Failure cases on three challenging interaction tasks. The policy reaches the relevant interaction region but fails to complete the required opening, engagement, or contact.
Predictive configuration
Fine-grained execution
Additional tasks
Overall
Main-view targets
Wrist pathway
OM
HM
TS
AB
BR
GR
Avg.
RGB
Unified (future RGB)
43
60
57
88
90
88
71.0
RGB+Mask
Unified (future RGB)
43
60
58
92
93
89
72.5
RGB+Skel.
Unified (future RGB)
44
61
59
89
91
92
72.7
RGB+Mask+Skel.
Unified (future RGB)
45
62
59
91
94
90
73.5
RGB
Separate (future latent)
59
66
68
88
88
87
76.0
Table 3: Factorial study of main-view supervision and wrist pathways. The study crosses four main-view target configurations with unified future-RGB and separate future-latent wrist pathways. Fine-grained execution comprises Open Microwave (OM), Hanging Mug (HM), and Turn Switch (TS); additional tasks comprise Adjust Bottle (AB), Blocks Ranking RGB (BR), and Grab Roller (GR). Green and yellow shading indicate the best and second-best result in each column.
Figure 6: Structured targets focus action readout. Compared with RGB-only future supervision, joint mask-and-skeleton supervision focuses readout on the task-relevant interaction.
World Action Models (WAMs) connect visual prediction with robot control, but supplying predictive context often requires expensive future-video generation. Direct policies avoid this cost but lack an explicit interface for accessing future-indexed predictive information. We introduce ForeWAM, a World Action Model that separates forecasting from rendering to expose and shape latent predictive context for efficient control. Its core mechanism, Future-KV, performs a single Video DiT prefill over the current visual latent and noise-initialized future slots, then reuses the resulting key-value states throughout action denoising. To make this context relevant to control, we introduce dynamics registers supervised by latent actions from a frozen teacher during training, encouraging representations of interaction-induced transitions. This reusable context supports a lightweight, single-layer action decoder. We evaluate ForeWAM on LIBERO, LIBERO-Plus, RoboCasa, and real-world manipulation tasks. Without additional policy-level embodied pretraining, ForeWAM improves RoboCasa success by 9.7 percentage points over Fast-WAM at the same budget of 50 demonstrations per task, reaching 59.2%. With a single-layer decoder, it achieves 77.6% success on LIBERO-Plus and reduces policy-query latency to 88.7 ms on an NVIDIA A800, delivering a 6.27-fold speedup over Fast-WAM. These results show that latent predictive computation provides useful foresight for robust, efficient control without explicit future-video generation.
Jiakai Huang, Zhongbo Wu, Siyu Xu +5
Shanghai Jiao Tong University · ACE Robotics · Nanyang Technological University
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future observations. However, conditioning future prediction only on the task prompt and observation context risks capturing generic task progression rather than the action-specific consequences of the executed action. We introduce SelfWAM, a unified self-grounded WAM built on a modality-specialized Mixture-of-Transformers (MoT) architecture that jointly predicts actions, action-conditioned future RGB frames, and robot self-masks, thereby grounding future prediction in the robot's visible body and its action-induced motion. During joint training, SelfWAM allows future visual queries to attend to a clean copy of the demonstrated action, turning the video branch into an action-specific consequence model while leaving the fast action-only inference path unchanged. To focus video learning on action-relevant visual changes, we use prompt-specific objectives for future robot self-mask prediction, which removes appearance details and provides a target whose temporal evolution is tightly coupled with the conditioning action. Together, clean-action conditioning and future self-mask supervision make future predictions more directly reflect how the executed action changes the robot's visible motion and the surrounding scene. Experiments on RoboTwin 2.0 and real-world manipulation tasks show that SelfWAM produces more action-sensitive futures and preserves fast policy inference, while improving policy performance.
World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-relevant state transitions are entangled with nuisance variations in texture, illumination, background, and viewpoint. We argue that WAMs should explicitly predict action-relevant future state rather than relying on RGB prediction alone. We introduce DreamWAM, which reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics. During training, DreamWAM combines joint latent denoising of RGB and motion with lightweight gated residual branches for geometry and semantics. Shared attention between VideoDiT and ActionDiT allows the action branch to learn from these future-state predictions, while all beyond-RGB supervision branches are disabled at inference and deployment remains RGB-only. Across both no-rollout and joint video-action inference, DreamWAM consistently improves the matched RGB-only baselines on LIBERO, from 97.30% to 98.40% and from 98.00% to 98.90%, respectively. The gains become larger under unseen LIBERO-Plus perturbations, from 51.36% to 63.44% and from 69.16% to 75.47%. The same robustness extends to real-world manipulation, where DreamWAM attains an average success rate of 74.4% across unseen changes in lighting, background, and object layout, compared with 55.6% for Fast-WAM-Joint. These results show that robust world-action learning depends not only on predicting the future, but on representing it in a form that matters for action. The code and models are publicly released at https://github.com/hustvl/DreamWAM.
Shanglin Yuan, Weiheng Zhao, Xin Shi +6
Huazhong University of Science and Technology · D-Robotics · Wuhan University +1