cs.CVOct 5, 2026

RealtimeWAM: One-Step Asynchronous World Action Models

Authors: Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong, Shiqiao Gu, Shunzi Yang, Ruihao Gong, Shen Ren, +2 more

Organizations: Nanyang Technological University · Beihang University · Sensetime · Continental Automotive Singapore

Abstract

World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, <1%<1\% drop) across these benchmarks while delivering significant end-to-end speedup (\eg, ∼25×\sim25\times on H100). Our code and checkpoints are available via this link.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RealtimeWAM: How Fast Can I Run My World Action Model?

    Oct 7, 2026Huanan Liu, Ye Li, Kangye Ji +8Efficient Inference for World Action ModelsWorld Action Models

  2. Flash-WAM: Modality-Aware Distillation for World Action Models

    Jun 3, 2026Arman Akbari, Ci Zhang, Arash Akbari +6Video Diffusion ModelsAction-Conditioned World Models

  3. ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

    Jun 17, 2026Yuyang Zhang, Wenyao Zhang, Zekun Qi +7Action-Conditioned World ModelsWorld Action Models