World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods address this by token pruning that prioritizes visual fidelity to reduce denoising costs in video diffusion models. However, these methods do not use action relevance to determine which future-frame tokens to retain during joint denoising in WAMs. In this paper, we propose Sparse-WAM, a training-free framework for action-guided sparse imagination that selectively processes future-frame tokens to accelerate WAM inference. We observe substantial overlap in the spatial distribution of attention from action tokens to future-frame tokens (action-to-future attention) between consecutive denoising steps, despite continued updates to the future representations. Motivated by this, we develop Action-Guided Token Selection to retain frame-specific action-relevant regions together with cross-frame context. However, a naive implementation can incur attention-scoring and token-packing overhead that offsets the computational savings from pruning. We therefore introduce Pilot, an efficient engine that reduces sparse inference overhead through lightweight scoring and cross-step reuse of token selections. On LIBERO with FastWAM-Joint and RoboLab-120 with Cosmos 3 Edge, Sparse-WAM achieves inference speedups of approximately 2.0× and 1.8×, respectively, over dense eager inference on an NVIDIA RTX 4090, while largely preserving task performance.
Figures & tables
Figure 1: Representative architectures of VLA and WAM. In WAM, noisy imagination tokens constitute the majority of the visual-action sequence and evolve throughout denoising.
Figure 2: Insight 1: (a) Action queries exhibit temporal alignment with future frames, with attention extending to neighboring frames near temporal group boundaries. (b) Action-relevant hotspots shift spatially across future frames. Insight 2: Excluding hotspots yields higher attention overlap between adjacent frames. Insight 3: Attention maps exhibit substantial spatial overlap between consecutive denoising steps in all three camera views.
Figure 3: Overview of Sparse-WAM. Future-frame token positions are selected using action-to-future attention and reused during sparse denoising.
Method
Success Rate (%)
Avg. SR (%)
Avg. Speedup
FLOPs
Spatial
Object
Goal
Long
FastWAM-Joint
99.00%
100.00%
98.20%
97.80%
98.75%
1.00×
100.00%
+ ToCa
98.80%
99.80%
98.60%
98.60%
98.95%
1.09×
69.19%
+ C 3 ache
98.60%
99.40%
97.60%
97.40%
98.25%
1.30×
75.00%
+ SpecPrune-VLA
97.20%
99.80%
96.80%
92.40%
96.55%
1.43×
73.53%
+ WorldCache
99.60%
99.80%
98.00%
98.60%
99.00%
1.62×
60.00%
Table 1: Performance and inference efficiency on LIBERO.
Method
Success Rate (%)
Avg. SR (%)
Avg. Speedup
FLOPs
Simple
Moderate
Complex
Cosmos3-Edge
25.60%
23.30%
11.80%
22.90%
1.00×
100.00%
+ ToCa
9.84%
14.10%
1.18%
10.00%
1.26×
65.68%
+ C 3 ache
9.06%
12.82%
0.59%
9.08%
1.51×
75.14%
+ SpecPrune-VLA
10.31%
14.10%
0.00%
10.08%
1.52×
50.35%
+ WorldCache
21.56%
21.03%
11.18%
19.92%
1.52×
75.14%
Table 2: Performance and inference efficiency on RoboLab-120.
Figure 4: Tasks on LIBERO, RoboLab-120 and Real World.
Method
Success Rate (%)
Latency (ms)
Speedup
Pack objects
Stack cups
Battery assembly
Average
FastWAM-Joint
75.00
75.00
83.33
77.78
501
1.00×
+ Sparse-WAM
83.33
75.00
66.67
75.00
242
2.08×
Table 3: Real-world manipulation performance on AgileX Cobot Magic.
Figure 5: Ablation studies of token selection components and future-token pruning ratios.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Cosmos 3 Nano
Cosmos 3 Edge
FastWAM-Joint
Transformer layers
36
28
30
Camera views
3
3
2
Tokens per future frame
360
340
98
Future latent frames
8
8
2
Denoising steps
4
4
10
Guidance scale
3.0
3.0
1.0
Appendix
Table 4: Backbone and default Sparse-WAM configurations.
Configuration / Method
Latency (ms)
Speedup
(a) Execution-backend ablation
Dense eager
859.31
1.00×
Sparse-WAM eager
555.55
1.55×
+ CUDA Graph
467.16
1.84×
+ torch.compile
464.95
1.85×
(b) Matched-backend comparison
Appendix
Table 5: Cosmos 3 Edge latency per action chunk. Panel (a) uses dense eager as the reference; panel (b) enables CUDA Graphs and torch.compile for every method, including dense inference. SpecPrune-VLA retains at least 184 future tokens per frame.
Method
Success Rate (%)
Latency (ms)
Camera
Robot
Lang.
Light
Backg.
Noise
Layout
Avg.
FastWAM-Joint
51.25
51.25
87.50
95.00
66.25
61.25
83.75
70.89
431.7
+ Sparse-WAM
31.25
46.25
85.00
92.50
57.50
47.50
81.25
63.04
207.2
FastWAM (action-only)
26.25
41.25
63.75
77.50
57.50
36.25
61.25
51.96
90.1
Appendix
Table 6: LIBERO-Plus success rates (%) and generation latency per action chunk.