Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework that adapts computation to the current control sequence. Specifically, the action-aware scheduling and refresh policy detects action changes across and within chunks, and labels each chunk as initial, transition, turning, or steady state. At one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual from the most recent fully computed step in the same chunk. To exploit the persistence of low-frequency structure during steady interaction, we further introduce a frequency-mixed history prior guidance that incorporates complementary information from the preceding clean latent without an additional DiT forward pass. Evaluated on Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21x to 1.41x DiT-backbone speedups without model retraining while improving WBench Overall scores over original inference across all three models.
Figures & tables
Figure 1: Our CtrlCache adapts DiT computation to the incoming controls. The left panel shows how the control-aware policy selects full computation, residual reuse, or refresh across representative control states. The right panel summarizes WBench Overall and DiT-backbone speedup on Matrix-Game 2.0 ( He et al., 2025 ) , LingBot-World v1 ( Robbyant Team et al., 2026 ) , and LingBot-World v2 ( Gao et al., 2026 ) . Here, Full denotes a regular full DiT forward, Reuse denotes residual reuse at the selected denoising step, and Refresh denotes a full forward at an action transition.
Figure 2: Cosine similarity between the clean latents of adjacent chunks under different control signals for three representative sequences. Each panel compares the full latent, low-frequency and high-frequency components. The shading background marks the control regime of each chunk, i.e., non-turning (steady), sustained turning, and action change (transition). Similarity drops at the action changes, and the low-frequency component stays more persistent than the high-frequency one.
Figure 3: Overview of CtrlCache. (a) Chunk-wise autoregressive denoising retains full computation at the first and final steps. (b) The action-aware policy enables residual reuse at the selected interior step for turning and steady chunks, while preserving full computation for initial and transition chunks. (c) Frequency-mixed history prior guidance transfers structure from the preceding chunk’s last clean latent frame to guide the first-step clean latent estimate during steady interaction.
Model
Methods
WBench
vs. Original
Efficiency
Quality ↑
Cons. ↑
Inter. ↑
Setting ↑
Physics ↑
Overall ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Latency (s) ↓
Speedup ↑
Matrix-Game 2.0
Original
0.7360
0.6838
0.8261
0.5129
0.5127
0.6543
–
–
–
14.99
1.00 ×
TeaCache
0.7252
0.6820
0.8145
0.4982
0.5275
0.6495
13.89
0.3344
0.4810
10.33
1.45 ×
EasyCache
0.7245
0.6803
0.8093
0.5464
0.5518
0.6625
13.46
0.3152
0.5021
10.17
1.47 ×
Our CtrlCache
0.7370
0.6867
0.8265
0.5527
0.5900
0.6786
14.51
0.3504
0.4757
10.67
1.41 ×
LingBot-World v1
Original
0.8073
0.8715
0.8088
0.7868
0.6338
0.7816
–
–
–
104.54
1.00 ×
Table 1: Quality and efficiency comparison of CtrlCache and baselines on the WBench navigation track ( Ying et al., 2026 ) . “ vs. Original” compares each accelerated method with the original full-computation model under identical prompts, controls, and initial frames. “Efficiency” reports DiT-backbone denoising latency (mean seconds per case), excluding KV-cache commit/update time, and speedup over Original. Bold denotes the best result within each model.
Figure 4: Qualitative comparisons of CtrlCache and baselines. For each case, all methods use the same initial condition and control sequence. The blue dashed rectangles mark reference regions for comparing layout, object placement, and motion-related details across methods. The red rectangles highlight distortions in local appearance and scene structure, while the green rectangles highlight better-preserved details in our results. Control strings follow the original WBench convention, where W / S / A / D denote translational commands and left / right / up / down denote view changes.
Model
Variant
WBench
Efficiency
Quality ↑
Cons. ↑
Inter. ↑
Setting ↑
Physics ↑
Overall ↑
Latency (s) ↓
Speedup ↑
Matrix-Game 2.0
Original
0.7360
0.6838
0.8261
0.5129
0.5127
0.6543
14.99
1.00 ×
+ Residual reuse
0.7353
0.6870
0.8418
0.5285
0.4843
0.6554
10.31
1.45 ×
+ Action-aware policy
0.7310
0.6802
0.8372
0.5417
0.5169
0.6614
10.39
1.44 ×
+ History prior (St.+Tu.+Tr.)
0.7299
0.6772
0.8014
0.5436
0.5272
0.6559
10.65
1.41 ×
+ History prior (St.+Tu.)
0.7301
0.6816
0.8200
0.5030
0.5856
0.6641
10.67
1.41 ×
Table 2: Ablation studies on Matrix-Game 2.0 and LingBot-World v1 using WBench navigation track. The first two variants evaluate naive residual reuse (Eq. ( 1 )) alone and with our proposed action-aware scheduling and refresh policy (abbreviated as Action-aware policy ; Sec. 3.2 ), respectively. The last three fix both components and apply the frequency-mixed history prior guidance (abbreviated as History prior ; Sec. 3.3 ) on steady (St.), steady and turning (St.+Tu.), or steady , turning and transition (St.+Tu.+Tr.) chunks. Efficiency reports DiT-backbone denoising latency and speedup relative to Original. Bold denotes the best result within each model.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Step s
Timestep
Overall ↑
Matrix-Game 2.0
2
908
0.6554
3
713
0.4930
LingBot-World v1
2
978
0.8014
3
947
0.8021
4
825
0.5624
Appendix
Table D.1: Hyperparameter studies on Matrix-Game 2.0 and LingBot-World v1. Each entry is WBench Overall; bold denotes the best result within each model and panel. (a) Reuse denoising step s . (b) Temporal weights wk ( α=0.5 , λ=0.5 ). (c) Frequency composition ( λ=0.5 ). (d) Blend strength λ ( α=0.5 ).
Figure 8Figure 9
Figure E.1: Additional qualitative comparisons on Matrix-Game 2.0, LingBot-World v1, and LingBot-World v2. Each comparison uses the same initial condition and WBench control sequence across methods; the corresponding prompts and controls are shown with the frames. Rows correspond to Original, TeaCache, EasyCache, and Our CtrlCache.
Diffusion models achieve state-of-the-art video generation quality, but their inference remains expensive due to the large number of sequential denoising steps. This has motivated a growing line of research on accelerating diffusion inference. Among training-free acceleration methods, caching reduces computation by reusing previously computed model outputs across timesteps. Existing caching methods rely on heuristic criteria to choose cache/reuse timesteps and require extensive tuning. We address this limitation with a principled sensitivity-aware caching framework. Specifically, we formalize the caching error through an analysis of the model output sensitivity to perturbations in the denoising inputs, i.e., the noisy latent and the timestep, and show that this sensitivity is a key predictor of caching error. Based on this analysis, we propose Sensitivity-Aware Caching (SenCache), a dynamic caching policy that adaptively selects caching timesteps on a per-sample basis. Our framework provides a theoretical basis for adaptive caching, explains why prior empirical heuristics can be partially effective, and extends them to a dynamic, sample-specific approach. Experiments on Wan 2.1, CogVideoX, and LTX-Video show that SenCache achieves better visual quality than existing caching methods under similar computational budgets.
Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpoint yields consistent content, remains an open problem. Full KV-cache attention preserves this consistency but breaks real-time constraints: memory footprint and attention cost grow linearly with rollout length. Sliding window inference restores throughput but discards long-term consistency. We propose WorldKV, a training-free framework with two components: World Retrieval and World Compression. World Retrieval stores evicted KV-cache chunks in GPU/CPU memory and selectively retrieves scene-relevant chunks via camera/ action correspondence, inserting them back into the native attention window without re-encoding. World Compression prunes redundant tokens within each chunk via key-key similarity to an anchor frame, halving per-chunk storage to fit 2x more history under a fixed budget. On Matrix-Game-2.0 and LingBot- World-Fast, WorldKV matches or exceeds full-KV memory fidelity at roughly 2x the throughput, and is competitive with memory-trained baselines without any fine-tuning. Project Page: https://cvlab-kaist.github.io/WorldKV/
Real-time world simulation is becoming a key infrastructure for scalable evaluation and online reinforcement learning of autonomous driving systems. Recent driving world models built on autoregressive video diffusion achieve high-fidelity, controllable multi-camera generation, but their inference cost remains a bottleneck for interactive deployment. However, existing diffusion caching methods are designed for offline video generation with multiple denoising steps, and do not transfer to this scenario. Few-step distilled models have no inter-step redundancy left for these methods to reuse, and sequence-level parallelization techniques require future conditioning that closed-loop interactive generation does not provide. We present X-Cache, a training-free acceleration method that caches along a different axis: across consecutive generation chunks rather than across denoising steps. X-Cache maintains per-block residual caches that persist across chunks, and applies a dual-metric gating mechanism over a structure- and action-aware block-input fingerprint to independently decide whether each block should recompute or reuse its cached residual. To prevent approximation errors from permanently contaminating the autoregressive KV cache, X-Cache identifies KV update chunks (the forward passes that write clean keys and values into the persistent cache) and unconditionally forces full computation on these chunks, cutting off error propagation. We implement X-Cache on X-world, a production multi-camera action-conditioned driving world model built on multi-block causal DiT with few-step denoising and rolling KV cache. X-Cache achieves 71% block skip rate with 2.6x wall-clock speedup while maintaining minimum degradation.