cs.CVAug 25, 2026

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

Authors: Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai, Yi Xu, Can Cui, Jiayan Wang, Zichong Yang, +3 more

Organizations: Virginia Tech · Drexel University · Northeastern University · Purdue University

Abstract

Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control on a single shared video DiT backbone: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non-interfering attention mask that isolates video representations and staleness-robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed-success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24-task RoboCasa kitchen benchmark (vs. 67.1% for synchronous Cosmos Policy) and 99.0% on LIBERO while cutting per-chunk control latency 24×24\times relative to synchronous world-action models (48 ms on one A100). In single-arm and bimanual real-robot manipulation, it achieves higher average success than π0.5π_{0.5} without any robot-data pretraining. Code is available at https://github.com/linhanwang/GlanceWAM.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

    Jun 17, 2026Yuyang Zhang, Wenyao Zhang, Zekun Qi +7World ModelsVideo Generation

  2. Latent Action as Intention Enables Efficient Future Imagination for World Action Models

    Aug 25, 2026Xiang Li, Yupeng Zheng, Songen Gu +11Efficient World-Action ModelLatent Actions

  3. Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

    Jun 8, 2026Jiajun Li, Tiecheng Guo, Yifan Ye +9Efficient World-Action ModelFaster-Wam