World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods address this by token pruning that prioritizes visual fidelity to reduce denoising costs in video diffusion models. However, these methods do not use action relevance to determine which future-frame tokens to retain during joint denoising in WAMs. In this paper, we propose Sparse-WAM, a training-free framework for action-guided sparse imagination that selectively processes future-frame tokens to accelerate WAM inference. We observe substantial overlap in the spatial distribution of attention from action tokens to future-frame tokens (action-to-future attention) between consecutive denoising steps, despite continued updates to the future representations. Motivated by this, we develop Action-Guided Token Selection to retain frame-specific action-relevant regions together with cross-frame context. However, a naive implementation can incur attention-scoring and token-packing overhead that offsets the computational savings from pruning. We therefore introduce Pilot, an efficient engine that reduces sparse inference overhead through lightweight scoring and cross-step reuse of token selections. On LIBERO with FastWAM-Joint and RoboLab-120 with Cosmos 3 Edge, Sparse-WAM achieves inference speedups of approximately 2.0× and 1.8×, respectively, over dense eager inference on an NVIDIA RTX 4090, while largely preserving task performance.
Figures & tables
Figure 1: Representative architectures of VLA and WAM. In WAM, noisy imagination tokens constitute the majority of the visual-action sequence and evolve throughout denoising.
Figure 2: Insight 1: (a) Action queries exhibit temporal alignment with future frames, with attention extending to neighboring frames near temporal group boundaries. (b) Action-relevant hotspots shift spatially across future frames. Insight 2: Excluding hotspots yields higher attention overlap between adjacent frames. Insight 3: Attention maps exhibit substantial spatial overlap between consecutive denoising steps in all three camera views.
Figure 3: Overview of Sparse-WAM. Future-frame token positions are selected using action-to-future attention and reused during sparse denoising.
Method
Success Rate (%)
Avg. SR (%)
Avg. Speedup
FLOPs
Spatial
Object
Goal
Long
FastWAM-Joint
99.00%
100.00%
98.20%
97.80%
98.75%
1.00×
100.00%
+ ToCa
98.80%
99.80%
98.60%
98.60%
98.95%
1.09×
69.19%
+ C 3 ache
98.60%
99.40%
97.60%
97.40%
98.25%
1.30×
75.00%
+ SpecPrune-VLA
97.20%
99.80%
96.80%
92.40%
96.55%
1.43×
73.53%
+ WorldCache
99.60%
99.80%
98.00%
98.60%
99.00%
1.62×
60.00%
Table 1: Performance and inference efficiency on LIBERO.
Method
Success Rate (%)
Avg. SR (%)
Avg. Speedup
FLOPs
Simple
Moderate
Complex
Cosmos3-Edge
25.60%
23.30%
11.80%
22.90%
1.00×
100.00%
+ ToCa
9.84%
14.10%
1.18%
10.00%
1.26×
65.68%
+ C 3 ache
9.06%
12.82%
0.59%
9.08%
1.51×
75.14%
+ SpecPrune-VLA
10.31%
14.10%
0.00%
10.08%
1.52×
50.35%
+ WorldCache
21.56%
21.03%
11.18%
19.92%
1.52×
75.14%
Table 2: Performance and inference efficiency on RoboLab-120.
Figure 4: Tasks on LIBERO, RoboLab-120 and Real World.
Method
Success Rate (%)
Latency (ms)
Speedup
Pack objects
Stack cups
Battery assembly
Average
FastWAM-Joint
75.00
75.00
83.33
77.78
501
1.00×
+ Sparse-WAM
83.33
75.00
66.67
75.00
242
2.08×
Table 3: Real-world manipulation performance on AgileX Cobot Magic.
Figure 5: Ablation studies of token selection components and future-token pruning ratios.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Cosmos 3 Nano
Cosmos 3 Edge
FastWAM-Joint
Transformer layers
36
28
30
Camera views
3
3
2
Tokens per future frame
360
340
98
Future latent frames
8
8
2
Denoising steps
4
4
10
Guidance scale
3.0
3.0
1.0
Appendix
Table 4: Backbone and default Sparse-WAM configurations.
Configuration / Method
Latency (ms)
Speedup
(a) Execution-backend ablation
Dense eager
859.31
1.00×
Sparse-WAM eager
555.55
1.55×
+ CUDA Graph
467.16
1.84×
+ torch.compile
464.95
1.85×
(b) Matched-backend comparison
Appendix
Table 5: Cosmos 3 Edge latency per action chunk. Panel (a) uses dense eager as the reference; panel (b) enables CUDA Graphs and torch.compile for every method, including dense inference. SpecPrune-VLA retains at least 184 future tokens per frame.
Method
Success Rate (%)
Latency (ms)
Camera
Robot
Lang.
Light
Backg.
Noise
Layout
Avg.
FastWAM-Joint
51.25
51.25
87.50
95.00
66.25
61.25
83.75
70.89
431.7
+ Sparse-WAM
31.25
46.25
85.00
92.50
57.50
47.50
81.25
63.04
207.2
FastWAM (action-only)
26.25
41.25
63.75
77.50
57.50
36.25
61.25
51.96
90.1
Appendix
Table 6: LIBERO-Plus success rates (%) and generation latency per action chunk.
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.
Jiajun Li, Tiecheng Guo, Yifan Ye +9
1The University of Hong Kong · 2Peking University · 3Muka Robotics +2
World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirely for acceleration. We find that latent WAMs, despite matching explicit ones on in-distribution tasks, fail to retain the generalization benefits that originally motivated WAMs. To demonstrate this, we evaluate generalization along three axes: environmental perturbation, data efficiency, and task generalization. Controlled comparisons with a matched backbone, training data, and budget reveal consistent degradation across all three axes when the action expert no longer conditions on future representations. Further analysis shows that the gap arises almost entirely from the first denoising step: the benefit comes from preparing the future, not generating it. We therefore propose Simple-WAM, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the training-time noise schedule to this inference behavior. Across simulation and real-world tasks, Simple-WAM achieves the best of both worlds, leading explicit WAMs in generalization performance with efficiency comparable to Latent WAMs. Project Page: https://zrporz.github.io/Simple-WAM-Web/
Renping Zhou, Zanlin Ni, Zihao Fan +8
Leap Lab, Tsinghua University · University of Science and Technology of China · Beijing Institute of Technology
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit pretraining. In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow-matching action expert on the KV caches produced by image-editing denoising, using them as a compact world-action context. ImageWAM outperforms standard VLA baselines and matching competitive WAMs without additional policy pretraining across different simulator and real-world experiments. It also reduces FLOPs to 1/6 and latency to 1/4 of video-based WAMs. Attention analysis further shows that editing caches focus on task-relevant change regions, supporting image editing as an effective alternative to video-based world-action modeling.
Yuyang Zhang, Wenyao Zhang, Zekun Qi +7
Shanghai Jiao Tong University · Tsinghua University · Tencent Robotics X +1