GlanceWAM: Sparse Test-Time Imagination for World-Action Models
Organizations: Virginia Tech · Drexel University · Northeastern University · Purdue University
Abstract
Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control on a single shared video DiT backbone: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non-interfering attention mask that isolates video representations and staleness-robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed-success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24-task RoboCasa kitchen benchmark (vs. 67.1% for synchronous Cosmos Policy) and 99.0% on LIBERO while cutting per-chunk control latency relative to synchronous world-action models (48 ms on one A100). In single-arm and bimanual real-robot manipulation, it achieves higher average success than without any robot-data pretraining. Code is available at https://github.com/linhanwang/GlanceWAM.
Figures & tables
| (a) System | SR (%) |
|---|---|
| Cosmos Policy (external anchor) | 67.1 |
| plain co-training (no lookahead) | 64.4 |
| GlanceWAM, single-layer lookahead | 71.5 |
| same checkpoint, lookahead zeroed | 61.6 |
| same checkpoint, offset pinned † | 61.8 |
| retrained without mask (bidirectional) | 47.0 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| RoboCasa | LIBERO | Auxiliary supervision | World model | |
| kitchen | (avg) | beyond RGB demos | at test time | |
| DeVA ( Zhang et al., 2026 ) | 72.0 | 99.0 | affordance + depth decoders (contact labels, depth model) | joint denoising, every chunk |
| w/o guidance (their ablation) | 66.0 | — | none | (same) |
| Flex- ( Yan et al., 2026 ) , action-only | — | 98.7 | 3D pointmaps + DINOv3 semantic futures | dropped (fast path) |
| Flex- , full joint | — | 99.2 | (same) | joint denoising, every chunk (slow path) |
| GlanceWAM (ours) | 72.2 | 99.0 | none | async lookahead frame, 3 s |