Efficient World Action Model Inference with Adaptive Intermediate States
Authors: Zhinnan Liu, Haozhi Han, Ruge Zhang, Teng Ma, Tao Ma, Zheng Liu, Yifeng Chen, Yunquan Zhang, +3 more
Organizations: Xiamen University · Institute for AI Industry Research, Tsinghua University · School of Computer Science, Peking University · Institute of Computing Technology, Chinese Academy of Sciences · Alibaba Group
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present WAMACHINE, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that WAMACHINE achieves 1.47-3.05× speedups in observation-to-action latency and 2.23-3.27× speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.
Figures & tables
Figure 1: The main idea of WAMachine . WAMachine adapts state across replans, denoising steps, and Transformer layers to reduce computation and observation-to-action latency. K denotes native denoising steps; k1 and k2 denote steps before and after real observation arrival, respectively; k1+k2<K describes the illustrated accepted warm-start path.
Figure 2: Evidence for state continuity. (a) Consecutive replans show high cosine similarity and small relative L2 distance between trajectory latents at matching early stages (top); remapped initialization enables three-step inference with low action RMSE and 96–100% success on LIBERO (bottom). (b) Most anticipatory prefixes are ready within the action execution window (below y=x ); color shows action RMSE after observation rebinding relative to full inference under the real condition. (c) Relative L2 changes between adjacent layer hidden states are small in the middle Action-DiT layers (top); residual rescaling and state refresh control relative output L2 error compared with repeatedly reusing the first-step residual R1 (bottom). Analysis protocols appear in Appendix E .
Figure 3: The framework of WAMachine . Left: Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan. Top right: Observation Rebinding advances an anticipatory prefix during action execution, then rebinds retained denoising state to the real observation for continuation if consistency checks pass; otherwise, inference restarts from the remapped initialization under the real condition. Bottom: Residual Rescaling uses a shallow probe to rescale retained residuals and skip the middle layers; if consistency checks fail, it performs full computation of these layers and refreshes the retained layer state.
GPU inference time per replan
O2A latency
Subset SR
Method
Time (ms) ↓
Speedup ↑
Time (ms) ↓
Speedup ↑
(%) ↑
Cosmos Policy / LIBERO
Native
284.29
1.00 ×
636.77
1.00 ×
98.0
RTI-DP
59.14
4.81 ×
399.15
1.60 ×
68.5
RTC
1024.53
0.28 ×
2203.99
0.29 ×
98.0
VLA-Cache
461.48
0.62 ×
813.54
0.78 ×
98.0
Table 1: Inference efficiency and subset success rate (SR) on the same fixed 200-episode subset for each WAM. GPU inference time per replan and observation-to-action (O2A) latency are averaged over replans and reported in milliseconds; speedups are relative to Native. Each acceleration method is applied independently to Native, and WAMachine includes CUDA Graphs. All methods use one policy worker on one A100 80GB GPU; WAMachine uses an additional rendering GPU for Motus. Large-sample task success appears in Table 2 . Lower ( ↓ ) or higher ( ↑ ) is better.
Table 2: Large-sample closed-loop task success (%) for Native and WAMachine . Each method uses 6,000 LIBERO episodes for each of Cosmos Policy and Fast-WAM-IDM, and 1,000 Clean and 1,000 Randomized RoboTwin 2.0 episodes for Motus. Avg./Ret. reports average task success and the percentage of Native average task success retained.
Variant
TR
OR
RR
Graph
GPU (ms) ↓
O2A (ms) ↓
Subset SR (%) ↑
Native
✗
✗
✗
✗
415.19
470.96
97.5
Native + Graph
✗
✗
✗
✓
295.29
350.25
98.0
WAMachine
✓
✓
✓
✓
114.75
144.75
98.0
w/o TR
✗
✓
✓
✓
188.02
177.20
97.0
w/o OR
✓
✗
✓
✓
108.45
165.62
95.5
w/o RR
✓
✓
✗
✓
188.99
205.45
99.0
Table 3: Ablation of the three mechanisms and CUDA Graphs on LIBERO with Fast-WAM-IDM, using the same fixed 200-episode subset as Table 1 . TR, OR, and RR denote Trajectory Remapping, Observation Rebinding, and Residual Rescaling, respectively; Graph denotes CUDA Graphs. GPU and O2A report mean GPU inference time per replan and observation-to-action latency in milliseconds; subset SR reports task success (%) from the same runs. Lower ( ↓ ) or higher ( ↑ ) is better.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Benchmark
Large sample
Fixed subset
Cosmos Policy
LIBERO
6,000
200
Fast-WAM-IDM
LIBERO
6,000
200
Motus
RoboTwin 2.0
1,000+1,000
200 Clean
Appendix
Table 4: Evaluation episodes per configuration. Fixed subsets supply both task success and timing; the Motus subset uses Clean scenes only.
Model
Trajectory
First replan
Accepted prefix
Refresh
Cosmos Policy
Joint video/action
5
1+2
3
Fast-WAM-IDM
Video
10
2+4
6
Action
10
5
5
Motus
Joint video/action
10
3+4
7
Appendix
Table 5: Denoising steps by execution path. In the accepted prefix column, a+b denotes a anticipatory steps followed by b steps under the real condition; a single number denotes only steps under the real condition. Refresh counts start from the remapped initialization.
Branch
L
Probe
Middle layers
Tail
Evaluated
Cosmos Policy
28
[2,4)
[4,24)
[24,28)
8
Fast-WAM-IDM video
30
[2,4)
[4,26)
[26,30)
8
Fast-WAM-IDM action
30
[2,4)
[4,26)
[26,30)
8
Motus joint
30
[2,7)
[7,26)
[26,30)
11
Appendix
Table 6: RR layer boundaries and layer evaluations remaining after accepted reuse. Motus counts each coupled layer group once.
Parameter
Fast-WAM-IDM
Cosmos Policy
Motus
Minimum cosine similarity
0.80
0.984
0.95
Maximum relative fitting error
0.60
0.176
0.40
Appendix
Table 7: RR consistency thresholds used in Equation 8 .
All 200 episodes
Pairwise common success
Method
Time (s) ↓
Speedup ↑
Time (s) ↓
Speedup ↑
Cosmos Policy / LIBERO
Native
10.34
1.00 ×
–
–
RTI-DP ( n=135 )
95.85
0.11×
63.79
0.14×
RTC ( n=193 )
27.13
0.38×
26.23
0.38×
VLA-Cache ( n=193 )
12.51
0.83×
12.08
0.83×
Appendix
Table 8: Episode time and speedup on all 200 conditions and on pairwise common-success subsets. n gives the common-success count for each Native–method pair. Times are mean seconds per episode; each speedup uses Native time on the same subset. Bold and underline mark the best and second-best values within each model, respectively. Measurement boundaries are defined in Appendix D.1 .
Residual Rescaling (RR)
Observation Rebinding (OR)
Model
Acceptance rate
Layer-skip fraction
Acceptance rate
Replan coverage
Cosmos Policy
81.55%
38.43%
83.75%
74.79%
Fast-WAM-IDM
79.20%
45.80%
83.47%
71.40%
Motus
49.34%
10.36%
54.34%
38.67%
Appendix
Table 9: RR acceptance rate, layer-skip fraction, OR acceptance rate, and replan coverage of the complete WAMachine on the 200-episode subsets. All values are percentages; metric definitions are given in Appendix D.2 .
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives, especially with scarce robot demonstrations and in out-of-distribution scenarios. To bridge this gap, we introduce LAWA, a WAM architecture that uses compact latent actions as an operational representation of future intentions, enabling efficient test-time future imagination without generating future observations. Specifically, a discrete tokenizer enhanced by action-free pre-training produces manipulation-centric codebook targets. LAWA jointly denoises a continuous latent state anchored to these targets with executable action chunks while omitting the future-video branch at inference. On RoboCasa, LAWA achieves state-of-the-art average success rates of 65.6% and 80.8% in the few-shot and full data settings, improving over the matched Fast-WAM baseline by 9.6 and 4.5 points, respectively. It also preserves the performance level of the matched Joint-WAM variant while requiring 42.9% lower inference latency. LAWA also demonstrates competitive zero-shot robustness on LIBERO-Plus and superior performance on real-world tasks. These results show that future imagination need not be discarded: retaining it with compact latent actions yields an effective trade-off among performance, generalization, and latency. Code and models will be released.
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.
Jiajun Li, Tiecheng Guo, Yifan Ye +9
1The University of Hong Kong · 2Peking University · 3Muka Robotics +2
World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approaches face a fundamental dilemma: Joint-WAMs preserve future-aware representations during inference but incur prohibitive computation costs, while efficient alternatives remove future modeling at inference time and may lose the robustness benefits of temporal reasoning. In this work, we revisit the role of future representations in WAMs and show that inference-time future conditioning is critical for generalization under distribution shifts. This observation motivates Faster-WAM, an efficient future-conditioning WAM that preserves future representations while avoiding expensive video-action interaction. Faster-WAM introduces a sparse future-conditioning framework that computes future representations once and selectively reuses them throughout action denoising. Specifically, we propose SparseMoT to replace ubiquitous layer-wise fusion with selective video-action interaction at a compact subset of network stages, and Interval KV-Fusion to aggregate multi-depth future representations without increasing attention complexity. Experiments demonstrate that Faster-WAM achieves a substantially better performance-efficiency trade-off than existing WAMs. On the out-of-distribution LIBERO-Plus benchmark, Faster-WAM improves success rate from 49.14% to 73.57% compared with Fast-WAM, while running 2.21× faster than Joint-WAM. It further achieves state-of-the-art performance on LIBERO and RoboTwin 2.0, while demonstrating strong robustness in real-world manipulation.
Weiheng Zhao, Haoyi Jiang, Xin Shi +5
Huazhong University of Science and Technology · D-Robotics · Horizon Robotics +1