World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert's cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.
Figures & tables
Figure 1: The WAM-Cache inference pipeline. (a) Sparse refresh selection at chunk t : three complementary signals construct the refresh set Rt . Latent surprise measures feature drift from a token’s reference embedding; action-conditioned selection identifies tokens heavily attended to by the action expert in the preceding chunk; and an age counter forces recomputation after G consecutive reuses. (b) Sparse prefill and mixed-cache execution : the video DiT executes a sparse prefill exclusively over Rt , updating layerwise keys and values in-place. The action expert then denoises the next action chunk over the resulting mixed cache, querying fresh representations for Rt alongside retained keys and values for unselected tokens.
Figure 2: Sequential ablation of caching signals on clean RoboTwin 2.0. The first two configurations run without the age bound.
Clean
Randomized
Method
FLOPs
Δ FLOPs
SR
Δ SR
FLOPs
Δ FLOPs
SR
Δ SR
(T) ↓
(%)
(%) ↑
(pp)
(T) ↓
(%)
(%) ↑
(pp)
Fast-WAM (dense prefill)
1.053
0.0
90.16
0.0
1.053
0.0
89.66
0.0
+ FastV ( Chen et al., 2024 )
0.618
−41.3
62.72
−27.4
0.642
−39.0
60.18
−29.5
+ ToMe ( Bolya et al., 2022 )
0.622
−41.0
77.92
−12.2
0.639
−39.3
74.14
−15.5
+ Eventful ( Dutson et al., 2023 )
0.623
−40.8
52.94
−37.2
0.641
−39.2
51.36
−38.3
Table 1: Results on RoboTwin 2.0 (50 tasks, 100 episodes per task). FLOPs is the per-chunk video DiT prefill cost, and Δ FLOPs and Δ SR are relative to the dense prefill. Each baseline is calibrated to the prefill FLOPs of the first WAM-Cache row of each condition. The two WAM-Cache rows use (ba,bs)=(0.4,0.1) and (0.6,0.1) under clean observations and (0.4,0.2) and (0.4,0.3) under randomization.
SR (%) ↑
FLOPs
Δ FLOPs
Method
Spatial
Object
Goal
Long
Avg.
(T) ↓
(%)
Fast-WAM (dense prefill)
96.80
99.40
96.00
93.80
96.50
0.859
0.0
+ FastV ( Chen et al., 2024 )
73.20
93.60
75.40
73.00
78.80
0.588
−31.5
+ ToMe ( Bolya et al., 2022 )
94.20
97.60
92.20
83.80
91.95
0.587
−31.7
+ Eventful ( Dutson et al., 2023 )
95.40
97.60
94.20
87.40
93.65
0.599
−30.2
+ VLA-Cache ( Xu et al., 2025 )
96.00
98.40
94.80
90.20
94.85
0.588
−31.5
Table 2: Results on LIBERO (50 episodes per task, 500 per suite). Each baseline is calibrated to the prefill FLOPs of the first WAM-Cache row. The two WAM-Cache rows use (ba,bs)=(0.5,0.1) and (0.6,0.1) .
Figure 5
Figure 4: Hyperparameter sweeps on clean RoboTwin 2.0. (a) Age bound G at (ba,bs)=(0.4,0.1) , where w/o runs without the bound. (b) Attention budget ba at bs=0.1 and G=2 . Blue circles (left axis) are the success rate and gray squares (right axis) the prefill FLOPs saved. The dotted line is the dense prefill (90.16%), and the shaded band marks the configuration we use.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
RoboTwin 2.0
LIBERO
N
Clean
Rand.
Avg.
Spatial
Object
Goal
Long
Avg.
1
69.72
67.72
68.72
97.4
99.0
95.0
91.2
95.65
2
90.16
89.66
89.91
96.8
99.4
96.0
93.8
96.50
5
91.94
90.80
91.37
96.8
99.6
96.6
94.2
96.80
10
91.46
91.18
91.32
97.2
99.8
96.6
94.6
97.05
Appendix
Table 5: Dense Fast-WAM success rate (%) versus the number of action-denoising steps N on RoboTwin 2.0 and LIBERO.
Scene
Method
(ba,bs)
SR (%) ↑
Δ SR (pp)
Δ FLOPs (%)
Clean
Fast-WAM (dense prefill)
–
91.46
0.0
0.0
+ WAM-Cache
(0.4,0.1)
90.82
−0.6
−41.6
+ WAM-Cache, larger budget
(0.6,0.1)
91.46
0.0
−28.2
Randomized
Fast-WAM (dense prefill)
–
91.18
0.0
0.0
+ WAM-Cache
(0.4,0.2)
90.04
−1.1
−39.5
+ WAM-Cache, larger budget
(0.4,0.3)
90.94
−0.2
−37.3
Appendix
Table 6: WAM-Cache on RoboTwin 2.0 at N=10 action-denoising steps (50 tasks × 100 episodes per cell), with the budget settings of Table 1 . Δ SR and Δ FLOPs are relative to the dense prefill at N=10 .
FLOPs (T)
Prefill share
N
Prefill
Action expert
VAE
DiTs only
incl. VAE
2
1.053
0.209
0.571
83.5%
57.5%
10
1.053
1.044
0.571
50.2%
39.5%
Appendix
Table 7: Per-chunk FLOPs (T) of Fast-WAM at Sv=120 and the share taken by the video DiT prefill, with and without the VAE encoding.
Figure 5: Real-world tasks on the AIRBOT Play arm, viewed from the external camera. Top: picking up a carrot and placing it in a bowl. Bottom: stacking a red cube on a blue cube. Frames run left to right, from the initial scene through grasping and transport to placement.
World Action Models (WAMs) generalize better than standard Vision-Language-Action (VLA) policies to novel motions and environments, because a video-modeling objective lets them learn from abundant unlabeled video rather than scarce labeled robot demonstrations. This generalization is computationally expensive. To complete a task, a WAM runs over multiple inference chunks, and each chunk requires a costly denoising process. Existing acceleration methods reduce this cost by caching and reusing computation within a single chunk's denoising trajectory. Our empirical analysis reveals a substantial source of redundancy they overlook: redundancy across chunks. When a robot executes a smooth behavior, the residuals computed at a given denoising step are strongly correlated from one chunk to the next. We introduce C3ache, a training-free method that caches and reuses these residuals across inference chunks at the same denoising step. Experiments on benchmarks with a Fast-WAM backbone show that C3ache achieves up to a 2.5× speedup in total wall-clock inference time, with negligible degradation in task success rate.
Weisen Zhao, Lam Nguyen, Zhicong Lu +1
George Mason University · University of Central Florida
World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods address this by token pruning that prioritizes visual fidelity to reduce denoising costs in video diffusion models. However, these methods do not use action relevance to determine which future-frame tokens to retain during joint denoising in WAMs. In this paper, we propose Sparse-WAM, a training-free framework for action-guided sparse imagination that selectively processes future-frame tokens to accelerate WAM inference. We observe substantial overlap in the spatial distribution of attention from action tokens to future-frame tokens (action-to-future attention) between consecutive denoising steps, despite continued updates to the future representations. Motivated by this, we develop Action-Guided Token Selection to retain frame-specific action-relevant regions together with cross-frame context. However, a naive implementation can incur attention-scoring and token-packing overhead that offsets the computational savings from pruning. We therefore introduce Pilot, an efficient engine that reduces sparse inference overhead through lightweight scoring and cross-step reuse of token selections. On LIBERO with FastWAM-Joint and RoboLab-120 with Cosmos 3 Edge, Sparse-WAM achieves inference speedups of approximately 2.0× and 1.8×, respectively, over dense eager inference on an NVIDIA RTX 4090, while largely preserving task performance.
World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, <1% drop) across these benchmarks while delivering significant end-to-end speedup (\eg, ∼25× on H100). Our code and checkpoints are available via this link.
Chengtao Lv, Jinyang Du, Shuyi Feng +7
Nanyang Technological University · Beihang University · Sensetime +1