Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control on a single shared video DiT backbone: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non-interfering attention mask that isolates video representations and staleness-robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed-success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24-task RoboCasa kitchen benchmark (vs. 67.1% for synchronous Cosmos Policy) and 99.0% on LIBERO while cutting per-chunk control latency 24× relative to synchronous world-action models (48 ms on one A100). In single-arm and bimanual real-robot manipulation, it achieves higher average success than π0.5 without any robot-data pretraining. Code is available at https://github.com/linhanwang/GlanceWAM.
Figures & tables
Figure 1: Synchronous imagination versus sparse lookahead foresight. (a) Existing WAMs either skip test-time foresight or imagine video synchronously with every action chunk, paying K heavy video denoising steps per chunk. (b) GlanceWAM glances ahead asynchronously to a single latent lookahead frame Hf≈3 s away and reuses it across ≈4 action chunks, each decoded in 48 ms. (c) Success rate on RoboCasa kitchen (24 tasks, demos only).
Figure 2: GlanceWAM architecture. (a) In co-training, the video DiT predicts the future target xt+Hf , while the action head learns inverse dynamics toward a lookahead frame xt+u at a random offset u∼U(0,Hf] , conditioned on Δ=u , so it sees every lookahead age it will meet at deployment. Two separate VAE passes and the prefix-LM mask M keep the lookahead out of video prediction. (b) At inference, the lookahead latent z^la is generated asynchronously once per Hf and never decoded; each 0.8 s action chunk reads the held z^la at the decaying offset Δ in 48 ms, within the offset range seen in training.
Figure 3: Non-interfering attention mask. 3-class prefix-LM mask M .
Figure 4: Imagined lookaheads. Observation ot , the GlanceWAM lookahead z^la (decoded for display only), and the real frame xt+3s , in RoboCasa kitchen (top) and real IMETA Y1 rollouts (bottom). The lookahead gets the target layout right while coarsening texture. Wrist view inset; more in Figure 10 .
Table 5
Figure 5: Real-robot evaluation on IMETA Y1. Success of GlanceWAM, π0.5 , and Fast-WAM over 50 trials per task; thumbnails show one GlanceWAM rollout per task.
Table 4: Comparison with concurrent world-action models. Benchmark numbers as reported by each paper under its own evaluation protocol. The DeVA ablation row removes its affordance and depth decoders, leaving RGB demonstrations only, as in GlanceWAM. Flex- π does not report RoboCasa kitchen; its two LIBERO rows are the reported FLEX- π∗ deployment modes.
Figure 9: Where the action head reads the lookahead (extends Figure 7 ). Two probed contexts, third-person view of the decoded lookahead z^la . Cross-attention: value-weighted contribution of each lookahead token to the action queries’ cross-attention output, i.e. the norm of its attention-weighted value after the output projection. The value-weighted share in § 4.4 is ∥lookahead part∥/(∥obs part∥+∥lookahead part∥) of the attention output. Action sensitivity: action shift when a 2×2 -token lookahead patch is overwritten with the observation tokens at the same location. Each map is normalized independently.
Figure 10: Lookahead frames on eight kitchen tasks (extends Figure 4 ). For eight RoboCasa kitchen tasks (articulation, knobs, buttons, pick-and-place): the observation at t , the GlanceWAM lookahead, and the environment frame at t+3 s. Lookaheads get the target layout right (closed doors, objects on target surfaces) while coarsening texture; they are decoded to RGB only for display. One third-person view per panel, wrist view inset.
Figure 11: Generated lookaheads vs. sampler budget K (RoboCasa kitchen demonstration contexts, matched seeds per row; last column: the demonstration frame at t+Hf ). Fidelity improves from K=1 to K=10 (the one-step lookahead smears the arm and the manipulated object, most visibly in the wrist insets), yet success barely changes over K=1 – 30 (Table 3 b). Layout as in Figure 10 .
Figure 12: Real-robot rollout keyframes. Initial, grasp, and final frames from successful GlanceWAM rollouts. Single-arm rows ( stack cups , pick and place , remove cup cap ) are side-view recordings cropped to the workspace; bimanual rows ( stack cups , bowl ) are frames from the policy’s top camera.
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit pretraining. In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow-matching action expert on the KV caches produced by image-editing denoising, using them as a compact world-action context. ImageWAM outperforms standard VLA baselines and matching competitive WAMs without additional policy pretraining across different simulator and real-world experiments. It also reduces FLOPs to 1/6 and latency to 1/4 of video-based WAMs. Attention analysis further shows that editing caches focus on task-relevant change regions, supporting image editing as an effective alternative to video-based world-action modeling.
Yuyang Zhang, Wenyao Zhang, Zekun Qi +7
Shanghai Jiao Tong University · Tsinghua University · Tencent Robotics X +1
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives, especially with scarce robot demonstrations and in out-of-distribution scenarios. To bridge this gap, we introduce LAWA, a WAM architecture that uses compact latent actions as an operational representation of future intentions, enabling efficient test-time future imagination without generating future observations. Specifically, a discrete tokenizer enhanced by action-free pre-training produces manipulation-centric codebook targets. LAWA jointly denoises a continuous latent state anchored to these targets with executable action chunks while omitting the future-video branch at inference. On RoboCasa, LAWA achieves state-of-the-art average success rates of 65.6% and 80.8% in the few-shot and full data settings, improving over the matched Fast-WAM baseline by 9.6 and 4.5 points, respectively. It also preserves the performance level of the matched Joint-WAM variant while requiring 42.9% lower inference latency. LAWA also demonstrates competitive zero-shot robustness on LIBERO-Plus and superior performance on real-world tasks. These results show that future imagination need not be discarded: retaining it with compact latent actions yields an effective trade-off among performance, generalization, and latency. Code and models will be released.
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.
Jiajun Li, Tiecheng Guo, Yifan Ye +9
1The University of Hong Kong · 2Peking University · 3Muka Robotics +2