World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise. We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video--action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents. Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model. Project page: https://ctrl-wam.github.io/
Figures & tables
Figure 1: Overview of CtrlWAM . Standard noising pairs perturbed actions with video tied to the recorded future. Aligned noising instead renders the consequences of the perturbed trajectories before adding visual noise, while retaining the recorded action–video pair as the training target. Combined with warped video–action noise schedules, the model supports joint prediction, action-conditioned video generation, video-conditioned action prediction, and control of surrounding agents, with applications to driving and robotic manipulation.
Figure 2: Safe video, unsafe action. This schematic illustrates intent–foresight misalignment: after noising and denoising, the video depicts a safe pass, while a small deviation in the predicted trajectory leads to a collision.
Figure 3: Visual and action predictions evolve at different rates during denoising. Left: driving and manipulation examples show that visual structure becomes established earlier, while the estimated ego and gripper trajectories continue to change. Columns indicate denoising steps and noise levels. Right: schematic schedules illustrate the mismatch under shared noising. Our proposed warp instead assigns higher video noise at a given action noise level.
Figure 4: The architecture of CtrlWAM . A shared diffusion transformer processes video, ego-action, and optional agent-action tokens, conditioned on observed history. Future tokens are denoised as prediction targets or supplied clean as conditions, enabling joint prediction, action-conditioned video generation, and video-conditioned action prediction. The additional agent streams support prediction and control of surrounding agents.
Figure 5: Counterfactual ego control with CtrlWAM . Rows show the generation under the factual ego, and stop, accelerate commands. Columns show corresponding times from 1.8 to 6.4 s; the small plots at right show the ego trajectories associated with each row.
Figure 6: Ego-command counterfactuals on 100 clips. Each level is averaged over four commands (brake to stop, accelerate, nudge left, nudge right). Warped schedules ( ς = 2, 3, 5) follow commands more closely than the shared schedule ( ς = 1) at every level, with the gap widening as perturbations grow.
Table 2: Counterfactual agent-command following. Scene history, the ego command, and sampling noise are held fixed while a selected agent’s future trajectory changes. Gemini 3.8 Flash evaluates agreement between the generated video and the supplied agent command; scores are reported on a percentage scale.
Category
Cosmos 3 Zero-shot
Cosmos 3 SFT
GT-Src
CtrlWAM
Oracle
Visual Quality
0.447
0.560
0.563
0.576
0.598
Motion Quality
0.247
0.316
0.306
0.327
0.348
Content Consistency
0.413
0.650
0.620
0.665
0.662
Physics Adherence
0.258
0.456
0.480
0.540
0.782
3D Accuracy
0.812
0.884
0.899
0.921
0.956
Controllability
0.635
0.774
0.788
0.819
0.869
Table 3: WorldArena Track-1 summary ( Shang et al., 2026 ) on RoboTwin manipulation ( 1,000 episodes). Rows report six category averages and EWMScore, which is 100× the mean of the 15 individual metrics. Oracle evaluates ground-truth videos and is excluded from rankings.
Figure 8: Qualitative RoboTwin manipulation comparison. Rows show the recorded video and generations from Cosmos-SFT, the GT-Src control, and CtrlWAM at matching frame indices. The trajectory-accuracy annotations apply to this episode, for which CtrlWAM more closely reproduces the recorded bowl motion and final arrangement.
Table 10
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Additional counterfactual ego control on a highway (top) and at an urban intersection (bottom). Each panel compares the recording with generations under factual, stop, accelerate, nudge-left, and nudge-right commands at matching times over 0.0–6.4 s. Plots at right show the corresponding ego trajectories.
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.
Jiajun Li, Tiecheng Guo, Yifan Ye +9
1The University of Hong Kong · 2Peking University · 3Muka Robotics +2
World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirely for acceleration. We find that latent WAMs, despite matching explicit ones on in-distribution tasks, fail to retain the generalization benefits that originally motivated WAMs. To demonstrate this, we evaluate generalization along three axes: environmental perturbation, data efficiency, and task generalization. Controlled comparisons with a matched backbone, training data, and budget reveal consistent degradation across all three axes when the action expert no longer conditions on future representations. Further analysis shows that the gap arises almost entirely from the first denoising step: the benefit comes from preparing the future, not generating it. We therefore propose Simple-WAM, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the training-time noise schedule to this inference behavior. Across simulation and real-world tasks, Simple-WAM achieves the best of both worlds, leading explicit WAMs in generalization performance with efficiency comparable to Latent WAMs. Project Page: https://zrporz.github.io/Simple-WAM-Web/
Renping Zhou, Zanlin Ni, Zihao Fan +8
Leap Lab, Tsinghua University · University of Science and Technology of China · Beijing Institute of Technology
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit pretraining. In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow-matching action expert on the KV caches produced by image-editing denoising, using them as a compact world-action context. ImageWAM outperforms standard VLA baselines and matching competitive WAMs without additional policy pretraining across different simulator and real-world experiments. It also reduces FLOPs to 1/6 and latency to 1/4 of video-based WAMs. Attention analysis further shows that editing caches focus on task-relevant change regions, supporting image editing as an effective alternative to video-based world-action modeling.
Yuyang Zhang, Wenyao Zhang, Zekun Qi +7
Shanghai Jiao Tong University · Tsinghua University · Tencent Robotics X +1