World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert's cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.
Figures & tables
Figure 1: The WAM-Cache inference pipeline. (a) Sparse refresh selection at chunk t : three complementary signals construct the refresh set Rt . Latent surprise measures feature drift from a token’s reference embedding; action-conditioned selection identifies tokens heavily attended to by the action expert in the preceding chunk; and an age counter forces recomputation after G consecutive reuses. (b) Sparse prefill and mixed-cache execution : the video DiT executes a sparse prefill exclusively over Rt , updating layerwise keys and values in-place. The action expert then denoises the next action chunk over the resulting mixed cache, querying fresh representations for Rt alongside retained keys and values for unselected tokens.
Figure 2: Sequential ablation of caching signals on clean RoboTwin 2.0. The first two configurations run without the age bound.
Clean
Randomized
Method
FLOPs
Δ FLOPs
SR
Δ SR
FLOPs
Δ FLOPs
SR
Δ SR
(T) ↓
(%)
(%) ↑
(pp)
(T) ↓
(%)
(%) ↑
(pp)
Fast-WAM (dense prefill)
1.053
0.0
90.16
0.0
1.053
0.0
89.66
0.0
+ FastV ( Chen et al., 2024 )
0.618
−41.3
62.72
−27.4
0.642
−39.0
60.18
−29.5
+ ToMe ( Bolya et al., 2022 )
0.622
−41.0
77.92
−12.2
0.639
−39.3
74.14
−15.5
+ Eventful ( Dutson et al., 2023 )
0.623
−40.8
52.94
−37.2
0.641
−39.2
51.36
−38.3
Table 1: Results on RoboTwin 2.0 (50 tasks, 100 episodes per task). FLOPs is the per-chunk video DiT prefill cost, and Δ FLOPs and Δ SR are relative to the dense prefill. Each baseline is calibrated to the prefill FLOPs of the first WAM-Cache row of each condition. The two WAM-Cache rows use (ba,bs)=(0.4,0.1) and (0.6,0.1) under clean observations and (0.4,0.2) and (0.4,0.3) under randomization.
SR (%) ↑
FLOPs
Δ FLOPs
Method
Spatial
Object
Goal
Long
Avg.
(T) ↓
(%)
Fast-WAM (dense prefill)
96.80
99.40
96.00
93.80
96.50
0.859
0.0
+ FastV ( Chen et al., 2024 )
73.20
93.60
75.40
73.00
78.80
0.588
−31.5
+ ToMe ( Bolya et al., 2022 )
94.20
97.60
92.20
83.80
91.95
0.587
−31.7
+ Eventful ( Dutson et al., 2023 )
95.40
97.60
94.20
87.40
93.65
0.599
−30.2
+ VLA-Cache ( Xu et al., 2025 )
96.00
98.40
94.80
90.20
94.85
0.588
−31.5
Table 2: Results on LIBERO (50 episodes per task, 500 per suite). Each baseline is calibrated to the prefill FLOPs of the first WAM-Cache row. The two WAM-Cache rows use (ba,bs)=(0.5,0.1) and (0.6,0.1) .
Figure 5
Figure 4: Hyperparameter sweeps on clean RoboTwin 2.0. (a) Age bound G at (ba,bs)=(0.4,0.1) , where w/o runs without the bound. (b) Attention budget ba at bs=0.1 and G=2 . Blue circles (left axis) are the success rate and gray squares (right axis) the prefill FLOPs saved. The dotted line is the dense prefill (90.16%), and the shaded band marks the configuration we use.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
RoboTwin 2.0
LIBERO
N
Clean
Rand.
Avg.
Spatial
Object
Goal
Long
Avg.
1
69.72
67.72
68.72
97.4
99.0
95.0
91.2
95.65
2
90.16
89.66
89.91
96.8
99.4
96.0
93.8
96.50
5
91.94
90.80
91.37
96.8
99.6
96.6
94.2
96.80
10
91.46
91.18
91.32
97.2
99.8
96.6
94.6
97.05
Appendix
Table 5: Dense Fast-WAM success rate (%) versus the number of action-denoising steps N on RoboTwin 2.0 and LIBERO.
Scene
Method
(ba,bs)
SR (%) ↑
Δ SR (pp)
Δ FLOPs (%)
Clean
Fast-WAM (dense prefill)
–
91.46
0.0
0.0
+ WAM-Cache
(0.4,0.1)
90.82
−0.6
−41.6
+ WAM-Cache, larger budget
(0.6,0.1)
91.46
0.0
−28.2
Randomized
Fast-WAM (dense prefill)
–
91.18
0.0
0.0
+ WAM-Cache
(0.4,0.2)
90.04
−1.1
−39.5
+ WAM-Cache, larger budget
(0.4,0.3)
90.94
−0.2
−37.3
Appendix
Table 6: WAM-Cache on RoboTwin 2.0 at N=10 action-denoising steps (50 tasks × 100 episodes per cell), with the budget settings of Table 1 . Δ SR and Δ FLOPs are relative to the dense prefill at N=10 .
FLOPs (T)
Prefill share
N
Prefill
Action expert
VAE
DiTs only
incl. VAE
2
1.053
0.209
0.571
83.5%
57.5%
10
1.053
1.044
0.571
50.2%
39.5%
Appendix
Table 7: Per-chunk FLOPs (T) of Fast-WAM at Sv=120 and the share taken by the video DiT prefill, with and without the VAE encoding.
Figure 5: Real-world tasks on the AIRBOT Play arm, viewed from the external camera. Top: picking up a carrot and placing it in a bowl. Bottom: stacking a red cube on a blue cube. Frames run left to right, from the initial scene through grasping and transport to placement.