World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural design. More general caching strategies exploit feature redundancy, but redundancy alone does not capture the changing computational demands of closed-loop control. To address these challenges, we present RealtimeWAM, a general, training-free framework that coordinates parallel execution with adaptive computation for low-latency inference across diverse WAM architectures. We exploit layerwise dependencies to overlap observation processing with prediction. However, concurrent branches still compete for GPU resources, limiting the benefit of parallel execution. We therefore adapt computation throughout the pipeline through selective reuse, caching observation features in visually stable regions and reusing Transformer residuals while reserving additional refinement for small predicted adjustments. We evaluate RealtimeWAM on FastWAM and OpenWAM across RoboTwin, LIBERO, and LIBERO-Plus. On an RTX 4090, measured mean inference latencies are 24.09 and 63.09 ms, corresponding to average speedups of 8.90× and 10.67×. Average success rates are 82.75% and 87.41%, respectively, within 0.02 and 0.53 percentage points of native inference. Across five real-world tasks, RealtimeWAM improves average success rates over native inference by 17.2 and 37.2 percentage points on FastWAM and OpenWAM, respectively.
Figures & tables
Figure 1: Responsive robot control with RealtimeWAM. Left: stacking three shuttlecocks on the conveyor requires rapid responses to scene changes. Right: mean measured inference latency on this task from Table 4 , on RTX 4090. The dashed line marks the empirically observed conveyor-task threshold of approximately 90 ms, above which task success was substantially lower (Appendix A ).
Figure 2: Parallel execution across WAM architectures. Left: action prediction alone. Right: joint video and action prediction. Each pair contrasts native and parallel execution. Layerwise release of observation K/V allows prediction to proceed alongside Prefill; shaded regions group work available for concurrent execution.
Figure 3: Changing motion and precision requirements. (a) Smoothed predicted speed (mm/control step) and relative deviation from a reference using more computation. Darker cells indicate larger deviations; green marks identify selected observations. (b) Illustrative cases: the same offset δ fits the large movement’s wide tolerance but exceeds the small adjustment’s narrow tolerance. Measurement details are in Appendix E .
Figure 4: Coordinating token reuse and motion-adaptive refinement. (a) Observation processing proceeds alongside Video and Action prediction. (b) Selected Prefill tokens receive fresh FFN outputs; the remaining positions reuse cached outputs. Future-video and action prediction execute their full branches at each selected Transformer evaluation. (c) Transformer refreshes at Euler updates 0 and 3 provide the motion estimate that selects additional refreshes at 6 and 9. Both schedules perform ten updates, with current input features and output heads evaluated at every step.
Figure 5: Prefill token allocation. (a) RGB change identifies required tokens. (b) Feature change ranks the remaining tokens to fill spare capacity. (c) Required and added tokens are refreshed; the rest stay cached. Each cell is one observation token.
RoboTwin
LIBERO
Average
Backbone
Method
Clean
Random
LIBERO
Plus
SR (%)
Lat. (ms)
Speedup
FastWAM
Native
91.88
91.78
97.60
49.80
82.77
214.33
1.00 ×
VLA-Cache
80.00
83.20
97.00
48.60
77.20
115.55
1.85 ×
TeaCache
81.42
90.86
97.05
49.20
79.63
49.28
4.35 ×
ToCa
91.35
91.24
94.55
49.15
81.57
102.66
2.09 ×
BAC
91.67
89.53
97.50
51.55
82.56
67.14
3.19 ×
Table 1: Training-free acceleration across manipulation benchmarks. Columns report success rates (%) and mean latency (ms); Average weights settings equally. Caching baselines use compilation and CUDA Graph replay. Speedup is original Native mean latency divided by method mean latency. Best and second-best results exclude Native; lower latency is better.
Stage
Optimization
FastWAM
OpenWAM
First
Full 10F
First
Full 10F
0
Native
67.24
214.33
81.37
672.83
1
+ CUDA Graph
36.83
104.09
69.34
524.69
2
+ Compile
36.21
99.68
60.62
450.90
3
+ Parallel
17.77
33.52
37.12
223.27
4
+ Token reuse
17.69
29.47
31.62
217.24
Table 2: Cumulative execution optimizations on RTX 4090 (ms). Values are mean latencies measured during complete closed-loop rollouts. First: latency from prepared CPU inputs through Prefill and the first predictive evaluation in a 2F run; Full 10F: request latency from a separate 10F run. Stages are cumulative. The plot shows first-forward speedup over Native. Darker/lighter fills mark the lowest/second-lowest latency among optimized stages.
Configuration
Token policy
Motion
Results
Reuse
Sel.
Fill
NFE
Gate
SR (%)
First
Request
Ours (full)
✓
✓
✓
Ada.
✓
93.4
30.96
63.76
w/o token reuse
×
n/a
n/a
Ada.
✓
93.8
37.41
72.10
Random tokens
✓
×
✓
Ada.
✓
91.3
30.75
62.32
w/o capacity filling
✓
✓
×
Ada.
✓
92.1
30.76
61.78
Random gate
✓
✓
✓
Rnd.
×
91.1
31.02
64.77
Table 3: Component ablations and latency diagnostics. Left: OpenWAM/RoboTwin success rate, measured first-forward latency (First), and rollout-mean request latency (Request); latencies are in ms. Fixed 2F/4F retain the normal token policy. ✓ / × : enabled/disabled; n/a: inactive; Sel.: RGB/feature selection; Gate: motion criterion; Ada./Rnd.: adaptive/random 2F/4F. Right: (a) historical OpenWAM first-stage profile (line), with measured first-forward latency for w/o filling (65 tokens) and Ours (80 tokens) marked by diamonds; (b) rollout-mean request latencies for Fixed 2F, Fixed 4F, and Ours, on RTX 4090.
Figure 6: Dynamic allocation over one rollout. Shared observations accompany evaluation budgets (a) and token allocation (b) across all 50 requests of an OpenWAM block-sorting rollout. Hatching marks initialization.
Figure 7: Real-world task sequences. Rows A–E show stacking three shuttlecocks, placing a carrot on a plate, hanging a mug, transferring a tennis ball, and intercepting a rolling can, respectively. Seven frames progress from left to right within each recording.
Task
FastWAM
OpenWAM
Native
Ours
Native
Ours
SR ↑
Lat. ↓
SR ↑
Lat. ↓
SR ↑
Lat. ↓
SR ↑
Lat. ↓
Three shuttlecocks
24.0
210.36
40.0
24.33
4.0
640.28
68.0
68.61
Carrot to plate
46.0
209.74
74.0
21.86
66.0
638.56
96.0
62.12
Hang mug
72.0
211.08
70.0
25.07
82.0
642.15
82.0
65.36
Tennis ball to bowl
26.0
210.52
34.0
23.24
16.0
639.47
76.0
64.85
Table 4: Real-world manipulation results. Success rate (SR, %) and mean inference latency (Lat., ms). Average weights tasks equally.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Real-world experimental setups. Front views (top) and oblique views (bottom) of the three setups. (a) Conveyor. (b) Inclined tabletop. (c) Level tabletop.
Task
Setup
Task goal
Trials
A: Three shuttlecocks
Conveyor
Stack the three shuttlecocks together.
50
B: Carrot to plate
Conveyor
Move the carrot from the belt onto the plate.
50
C: Hang mug
Level tabletop
Hang the mug by its handle on the rack.
50
D: Tennis ball to bowl
Conveyor
Move the tennis ball from the belt into the bowl.
50
E: Rolling can to bowl
Inclined table
Intercept the rolling can and place it in the bowl.
50
Appendix
Table 5: Real-world task goals and trial counts. Trials are counted per task, backbone, and inference configuration.
(a) FastWAM
Benchmark
Subset
Native
VLA-Cache
TeaCache
ToCa
BAC
Ours
RoboTwin
Clean
91.88
80.00
81.42
91.35
91.67
92.04
Random
91.78
83.20
90.86
91.24
89.53
90.50
Avg. SR
91.83
81.60
86.14
91.30
90.60
91.27
LIBERO
Spatial
98.20
98.00
97.80
94.40
98.20
96.20
Object
100.00
100.00
99.80
98.80
100.00
99.60
Appendix
Table 6: Category-level success rates (%). Each backbone uses the same methods and success-rate evaluations as Table 1 . Plus denotes LIBERO-Plus.
Model
N
Native
Compile + Graph
Parallel
OpenWAM
10
672.83
450.90
223.27
OpenWAM
4
240.53
205.49
101.95
OpenWAM
2
127.31
109.09
59.11
FastWAM
10
214.33
99.68
33.52
FastWAM
4
104.61
50.27
27.99
FastWAM
2
71.55
38.69
21.56
Appendix
Table 7: Dense execution controls at fixed evaluation budgets. Mean request latency measured during complete closed-loop rollouts in ms, following the main-text timing protocol. All three execution variants compute full FFNs.
Model
N
Cases
Mean case MAE
Max
OpenWAM
10
3
0.001288
0.007812
OpenWAM
4
3
0.001245
0.011719
OpenWAM
2
3
0.002375
0.031250
FastWAM
10
8
0.001449
0.015625
FastWAM
4
8
0.001384
0.015625
FastWAM
2
8
0.002515
0.031250
Appendix
Table 8: Parallel versus native output checks at matched budgets. Mean case MAE averages recorded action MAEs; Max is the largest coordinate difference. Each model uses its own effective action space.
Method
FastWAM
OpenWAM
Original
+ C/G
Reduction
Original
+ C/G
Reduction
Native
214.33
99.68
53.5%
672.83
450.90
33.0%
VLA-Cache
236.13
115.55
51.1%
743.32
400.26
46.2%
TeaCache
112.85
49.28
56.3%
362.31
245.99
32.1%
ToCa
225.50
102.66
54.5%
719.60
440.42
38.8%
BAC
151.17
67.14
55.6%
460.04
305.08
33.7%
Appendix
Table 9: Effect of compilation and CUDA Graph replay on inference latency. Original denotes each method’s original implementation; + C/G enables compilation and CUDA Graph replay with its execution schedule preserved. All methods report measured rollout-mean latencies (ms), with task settings, timing boundaries, and four-setting averaging following Table 1 . Native uses the 10F results in Table 7 . Reduction is the percentage decrease from Original to + C/G.
Preparation schedule
First
Later
Request P50
P90
P95
Sequential
23.06
20.06
51.32
51.65
51.79
VAE/condition overlap
23.06
20.08
50.55
50.91
50.97
Overlap + early RGB
22.97
20.05
50.26
50.62
50.73
Appendix
Table 10: Preparation scheduling at a fixed 80-token budget. Historical OpenWAM profiling, N=2 . Each row summarizes 100 requests in ms. First and later are GPU-event medians; request percentiles use wall time.
Figure 9: Stability of a selection across layers. (a) Overlap with each layer’s lowest-change set. (b) Relative FFN change at selected positions. Curves average 44 observation pairs. The future-video branch is a separate diagnostic selection; its mean overlap is 61.0%.
Figure 10: RGB and layer-0 FFN changes: typical examples. Observation pairs are eight environment steps apart. Camera regions are spatially aligned; each model and metric uses a fixed color range.
Figure 11: Examples with weaker spatial correspondence. Color ranges match Figure 10 ; low pixel change does not always imply low feature change.
Figure 12: Token selection and reuse diagnostics. (a,b) Simulation and real-video observations, reuse masks, and RGB/feature changes. (c) Historical OpenWAM profiling: request latency and normalized action MAE versus refreshed-token count at N=2 . (d) Random versus RGB-guided selection at matched counts; bars show mean action MAE, dots show individual masks, and whiskers show one standard deviation.
History / selector
Refresh
Mean MAE
Worst MAE
Max. error
Dense previous history / L0
64/120
0.002722
0.022480
0.140625
Recursive / L0
64/120
0.002571
0.011384
0.269531
Recursive / L0 + age cap
64/120
0.004006
0.035773
0.473633
Recursive / L0 + RGB + age cap
64/120
0.003292
0.016858
0.273438
Recursive / L0 + age cap
96/120
0.001596
0.012808
0.117188
Appendix
Table 11: Controlled OpenWAM token reuse in Prefill. All 44 transitions, N=2 ; one L0-feature mask per observation. Mean and worst-observation MAE and maximum component error use normalized active actions. Future-video and action branches remain dense.
Policy
Components
Mean NFE
XYZ deviation ↓
Refine
Motion gate
Abs. (mm)
Relative
Fixed 2F
×
×
2.00
7.65
0.0869
Fixed 4F
✓
×
4.00
3.68
0.0280
Ours
✓
✓
2.65
6.48
0.0468
Appendix
Table 12: Motion refinement ablation. Means over 93 paired observations; green ✓ /red × : enabled/disabled. Refine adds two evaluations; the motion gate selects requests. XYZ deviations are relative to 10F; absolute values are in mm.
Figure 13: Prediction deviations over all 93 paired observations. Each request contributes a 2F and a 4F point at the same coarse predicted displacement, joined by a faint vertical line. Both compare with the same-input 10F prediction over the first 24 XYZ steps. Left: deviation normalized by the 10F amplitude, with a 0.02 m floor. Right: absolute mean XYZ deviation on a logarithmic scale, in mm. The dashed vertical line marks the 8 cm motion-gate threshold.
Evaluation positions
Mean
Median
P95
Max
{0,3}
0.007304
0.006876
0.010497
0.027701
{0,3,6,9}
0.001224
0.001186
0.001713
0.005981
Appendix
Table 13: Fixed-budget action deviations on FastWAM. Executed-prefix normalized MAE against the same-observation 10F prediction over 1,798 observations.
World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, <1% drop) across these benchmarks while delivering significant end-to-end speedup (\eg, ∼25× on H100). Our code and checkpoints are available via this link.
Chengtao Lv, Jinyang Du, Shuyi Feng +7
Nanyang Technological University · Beihang University · Sensetime +1
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present WAMACHINE, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that WAMACHINE achieves 1.47-3.05× speedups in observation-to-action latency and 2.23-3.27× speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.
Zhinnan Liu, Haozhi Han, Ruge Zhang +8
Xiamen University · Institute for AI Industry Research, Tsinghua University · School of Computer Science, Peking University +2
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.
Jiajun Li, Tiecheng Guo, Yifan Ye +9
1The University of Hong Kong · 2Peking University · 3Muka Robotics +2