Efficient World Action Model Inference with Adaptive Intermediate States
Authors: Zhinnan Liu, Haozhi Han, Ruge Zhang, Teng Ma, Tao Ma, Zheng Liu, Yifeng Chen, Yunquan Zhang, +3 more
Organizations: Xiamen University · Institute for AI Industry Research, Tsinghua University · School of Computer Science, Peking University · Institute of Computing Technology, Chinese Academy of Sciences · Alibaba Group
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present WAMACHINE, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that WAMACHINE achieves 1.47-3.05× speedups in observation-to-action latency and 2.23-3.27× speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.
Figures & tables
Figure 1: The main idea of WAMachine . WAMachine adapts state across replans, denoising steps, and Transformer layers to reduce computation and observation-to-action latency. K denotes native denoising steps; k1 and k2 denote steps before and after real observation arrival, respectively; k1+k2<K describes the illustrated accepted warm-start path.
Figure 2: Evidence for state continuity. (a) Consecutive replans show high cosine similarity and small relative L2 distance between trajectory latents at matching early stages (top); remapped initialization enables three-step inference with low action RMSE and 96–100% success on LIBERO (bottom). (b) Most anticipatory prefixes are ready within the action execution window (below y=x ); color shows action RMSE after observation rebinding relative to full inference under the real condition. (c) Relative L2 changes between adjacent layer hidden states are small in the middle Action-DiT layers (top); residual rescaling and state refresh control relative output L2 error compared with repeatedly reusing the first-step residual R1 (bottom). Analysis protocols appear in Appendix E .
Figure 3: The framework of WAMachine . Left: Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan. Top right: Observation Rebinding advances an anticipatory prefix during action execution, then rebinds retained denoising state to the real observation for continuation if consistency checks pass; otherwise, inference restarts from the remapped initialization under the real condition. Bottom: Residual Rescaling uses a shallow probe to rescale retained residuals and skip the middle layers; if consistency checks fail, it performs full computation of these layers and refreshes the retained layer state.
GPU inference time per replan
O2A latency
Subset SR
Method
Time (ms) ↓
Speedup ↑
Time (ms) ↓
Speedup ↑
(%) ↑
Cosmos Policy / LIBERO
Native
284.29
1.00 ×
636.77
1.00 ×
98.0
RTI-DP
59.14
4.81 ×
399.15
1.60 ×
68.5
RTC
1024.53
0.28 ×
2203.99
0.29 ×
98.0
VLA-Cache
461.48
0.62 ×
813.54
0.78 ×
98.0
Table 1: Inference efficiency and subset success rate (SR) on the same fixed 200-episode subset for each WAM. GPU inference time per replan and observation-to-action (O2A) latency are averaged over replans and reported in milliseconds; speedups are relative to Native. Each acceleration method is applied independently to Native, and WAMachine includes CUDA Graphs. All methods use one policy worker on one A100 80GB GPU; WAMachine uses an additional rendering GPU for Motus. Large-sample task success appears in Table 2 . Lower ( ↓ ) or higher ( ↑ ) is better.
Table 2: Large-sample closed-loop task success (%) for Native and WAMachine . Each method uses 6,000 LIBERO episodes for each of Cosmos Policy and Fast-WAM-IDM, and 1,000 Clean and 1,000 Randomized RoboTwin 2.0 episodes for Motus. Avg./Ret. reports average task success and the percentage of Native average task success retained.
Variant
TR
OR
RR
Graph
GPU (ms) ↓
O2A (ms) ↓
Subset SR (%) ↑
Native
✗
✗
✗
✗
415.19
470.96
97.5
Native + Graph
✗
✗
✗
✓
295.29
350.25
98.0
WAMachine
✓
✓
✓
✓
114.75
144.75
98.0
w/o TR
✗
✓
✓
✓
188.02
177.20
97.0
w/o OR
✓
✗
✓
✓
108.45
165.62
95.5
w/o RR
✓
✓
✗
✓
188.99
205.45
99.0
Table 3: Ablation of the three mechanisms and CUDA Graphs on LIBERO with Fast-WAM-IDM, using the same fixed 200-episode subset as Table 1 . TR, OR, and RR denote Trajectory Remapping, Observation Rebinding, and Residual Rescaling, respectively; Graph denotes CUDA Graphs. GPU and O2A report mean GPU inference time per replan and observation-to-action latency in milliseconds; subset SR reports task success (%) from the same runs. Lower ( ↓ ) or higher ( ↑ ) is better.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Benchmark
Large sample
Fixed subset
Cosmos Policy
LIBERO
6,000
200
Fast-WAM-IDM
LIBERO
6,000
200
Motus
RoboTwin 2.0
1,000+1,000
200 Clean
Appendix
Table 4: Evaluation episodes per configuration. Fixed subsets supply both task success and timing; the Motus subset uses Clean scenes only.
Model
Trajectory
First replan
Accepted prefix
Refresh
Cosmos Policy
Joint video/action
5
1+2
3
Fast-WAM-IDM
Video
10
2+4
6
Action
10
5
5
Motus
Joint video/action
10
3+4
7
Appendix
Table 5: Denoising steps by execution path. In the accepted prefix column, a+b denotes a anticipatory steps followed by b steps under the real condition; a single number denotes only steps under the real condition. Refresh counts start from the remapped initialization.
Branch
L
Probe
Middle layers
Tail
Evaluated
Cosmos Policy
28
[2,4)
[4,24)
[24,28)
8
Fast-WAM-IDM video
30
[2,4)
[4,26)
[26,30)
8
Fast-WAM-IDM action
30
[2,4)
[4,26)
[26,30)
8
Motus joint
30
[2,7)
[7,26)
[26,30)
11
Appendix
Table 6: RR layer boundaries and layer evaluations remaining after accepted reuse. Motus counts each coupled layer group once.
Parameter
Fast-WAM-IDM
Cosmos Policy
Motus
Minimum cosine similarity
0.80
0.984
0.95
Maximum relative fitting error
0.60
0.176
0.40
Appendix
Table 7: RR consistency thresholds used in Equation 8 .
All 200 episodes
Pairwise common success
Method
Time (s) ↓
Speedup ↑
Time (s) ↓
Speedup ↑
Cosmos Policy / LIBERO
Native
10.34
1.00 ×
–
–
RTI-DP ( n=135 )
95.85
0.11×
63.79
0.14×
RTC ( n=193 )
27.13
0.38×
26.23
0.38×
VLA-Cache ( n=193 )
12.51
0.83×
12.08
0.83×
Appendix
Table 8: Episode time and speedup on all 200 conditions and on pairwise common-success subsets. n gives the common-success count for each Native–method pair. Times are mean seconds per episode; each speedup uses Native time on the same subset. Bold and underline mark the best and second-best values within each model, respectively. Measurement boundaries are defined in Appendix D.1 .
Residual Rescaling (RR)
Observation Rebinding (OR)
Model
Acceptance rate
Layer-skip fraction
Acceptance rate
Replan coverage
Cosmos Policy
81.55%
38.43%
83.75%
74.79%
Fast-WAM-IDM
79.20%
45.80%
83.47%
71.40%
Motus
49.34%
10.36%
54.34%
38.67%
Appendix
Table 9: RR acceptance rate, layer-skip fraction, OR acceptance rate, and replan coverage of the complete WAMachine on the 200-episode subsets. All values are percentages; metric definitions are given in Appendix D.2 .