cs.ROSep 28, 2026

Efficient World Action Model Inference with Adaptive Intermediate States

Authors: Zhinnan Liu, Haozhi Han, Ruge Zhang, Teng Ma, Tao Ma, Zheng Liu, Yifeng Chen, Yunquan Zhang, +3 more

Organizations: Xiamen University · Institute for AI Industry Research, Tsinghua University · School of Computer Science, Peking University · Institute of Computing Technology, Chinese Academy of Sciences · Alibaba Group

Abstract

World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present WAMACHINE\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that WAMACHINE\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}} achieves 1.47-3.05×\times speedups in observation-to-action latency and 2.23-3.27×\times speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Latent Action as Intention Enables Efficient Future Imagination for World Action Models

    Aug 25, 2026Xiang Li, Yupeng Zheng, Songen Gu +11Efficient World-Action ModelLatent Actions

  2. Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

    Jun 8, 2026Jiajun Li, Tiecheng Guo, Yifan Ye +9Efficient World-Action ModelFaster-Wam

  3. Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

    Aug 5, 2026Weiheng Zhao, Haoyi Jiang, Xin Shi +5Efficient World-Action ModelRecent World-Action Models