cs.CVOct 1, 2026

CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight

Authors: Chensheng Peng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, +4 more

Organizations: NVIDIA · UC Berkeley · UCLA

Abstract

World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise. We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video--action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents. Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model. Project page: https://ctrl-wam.github.io/

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

    Jun 8, 2026Jiajun Li, Tiecheng Guo, Yifan Ye +9Efficient World-Action ModelFaster-Wam

  2. What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

    Sep 28, 2026Renping Zhou, Zanlin Ni, Zihao Fan +8Latent World ModelsWorld Models

  3. ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

    Jun 17, 2026Yuyang Zhang, Wenyao Zhang, Zekun Qi +7World ModelsVideo Generation