In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion
Organizations: Meta · S-Lab, Nanyang Technological University
Abstract
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value (KV) cache by additional forwards to build the cache without advancing an output latent. However, every denoising forward itself already computes the in-flight KV of the current chunk. We introduce FlashForward, which directly reuses this cache to avoid the heavy cache-update-only model forwards. After the current chunk completes one denoising stage, its stage-specific cache is already available for the next chunk. Assigning one GPU to each stage therefore lets different chunks occupy different stages concurrently. This early availability has a quality cost: the resulting stage-matched history is noisy, causing appearance and motion drift among chunks. To complement it, FlashForward produces sparse auxiliary clean anchor latents before the corresponding region is generated so the generation trajectories can be stabilized by this two-sided conditioning. The two memories operate at different temporal scales: sparse clean anchor KV supplies coarse, long-range two-sided structural guidance, while dense stage-matched history preserves fine, recent evolution. With up to four GPUs, FlashForward runs -- faster than HiAR and -- faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbone scales at 480p and 720p. On VBench, for the 1.3B model at 480p, it achieves higher scores and remains stable at longer durations, demonstrating that FlashForward generates high-quality and temporally consistent videos across durations of 20s, 35s and 65s at a much faster generation speed.
Figures & tables
| Method | Memory noise level | When it becomes usable | Cache-update-only forwards | Chunk execution |
|---|---|---|---|---|
| Self-Forcing | After the chunk is re-encoded | per chunk | Serial | |
| N-C-Causal-rCM | After the last denoising stage | Serial | ||
| HiAR | After each stage-specific re-encoding | per chunk | Overlapped | |
| FlashForward | Immediately after each denoising stage | Overlapped |
| VBench-1.0 | LTX- | Wan2.1- | NOVA | Pyramid | SkyReels- | MAGI-1- | CausVid | Self- | Causal | HiAR | FlashForward |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Metric | Video | T2V-1.3B | Flow | V2-1.3B | 4.5B | Forcing | Forcing | ||||
| Total | 0.766 | 0.802 | 0.773 | 0.775 | 0.788 | 0.757 | 0.764 | 0.805 | 0.810 | 0.821 | 0.838 |
| Quality | 0.789 | 0.813 | 0.777 | 0.804 | 0.808 | 0.785 | 0.771 | 0.829 | 0.837 | 0.846 | 0.859 |
| Semantic | 0.685 | 0.766 | 0.757 | 0.670 | 0.707 | 0.647 | 0.740 | 0.708 | 0.701 | 0.723 | 0.753 |
| SFT phase (50-step sampling) | Distillation phase (4-step sampling) | |||||
|---|---|---|---|---|---|---|
| Ablation | Total | Quality | Semantic | Total | Quality | Semantic |
| Dual-role SFT | 0.8172 | 0.8371 | 0.7376 | 0.8184 | 0.8332 | 0.7593 |
| time-rebased supervision | 0.8150 | 0.8345 | 0.7371 | 0.8198 | 0.8338 | 0.7638 |
| role-specific embedding and LoRA | 0.8205 | 0.8377 | 0.7516 | 0.8380 | 0.8593 | 0.7528 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Wan2.1-T2V-1.3B (H100) | ||||||||||||
| 480p, GPU | 480p, multi-GPU | 720p, multi-GPU | ||||||||||
| Schedule | 5 | 20 | 35 | 65 | 5 | 20 | 35 | 65 | 5 | 20 | 35 | 65 |
| Bidirectional [-0.15em] | 3.31 [-0.15em] 0.00 | 32.02 [-0.15em] 0.01 | 91.22 [-0.15em] 0.04 | 295.83 [-0.15em] 0.12 | 1.13 [-0.15em] 0.01 | 8.95 [-0.15em] 0.01 | 24.01 [-0.15em] 0.02 | 75.28 [-0.15em] 0.05 | 3.74 [-0.15em] 0.00 | 39.59 [-0.15em] 0.06 | 115.79 [-0.15em] 0.04 | 379.17 [-0.15em] 0.64 |
| Self-Forcing [-0.15em] | 3.69 [-0.15em] 0.01 | 17.44 [-0.15em] 0.11 | 31.48 [-0.15em] 0.14 | 57.89 [-0.15em] 0.20 | 5.10 [-0.15em] 0.05 | 20.25 [-0.15em] 0.10 | 35.40 [-0.15em] 0.31 | 65.14 [-0.15em] 0.55 | 6.04 [-0.15em] 0.05 | 27.17 [-0.15em] 0.11 | 47.71 [-0.15em] 0.27 | 97.68 [-0.15em] 0.48 |
| HiAR [-0.15em] | 5.57 [-0.15em] 0.04 | 27.86 [-0.15em] 0.03 | 50.41 [-0.15em] 0.20 | 94.74 [-0.15em] 0.15 | 2.06 [-0.15em] 0.05 | 7.67 [-0.15em] 0.00 | 13.31 [-0.15em] 0.01 | 24.67 [-0.15em] 0.03 | 6.49 [-0.15em] 0.01 | 26.82 [-0.15em] 0.02 | 46.94 [-0.15em] 0.03 | 86.85 [-0.15em] 0.06 |
| FlashForward [-0.15em] | 4.56 [-0.15em] 0.07 | 17.77 [-0.15em] 0.14 | 32.80 [-0.15em] 0.21 | 59.64 [-0.15em] 0.46 | 2.21 [-0.15em] 0.08 | 6.32 [-0.15em] 0.02 | 10.79 [-0.15em] 0.22 | 21.29 [-0.15em] 0.08 | 5.71 [-0.15em] 0.05 | 19.33 [-0.15em] 0.27 | 35.14 [-0.15em] 0.13 | 66.53 [-0.15em] 0.12 |
| Wan2.1-T2V-1.3B (H100) | ||||||||||||
| 480p, GPU | 480p, GPU | 720p, GPU | ||||||||||
| Schedule | 5 | 20 | 35 | 65 | 5 | 20 | 35 | 65 | 5 | 20 | 35 | 65 |
| Bidirectional | 22.1 | 31.7 | 41.2 | 60.3 | 15.7 | 18.4 | 21.1 | 26.5 | 17.0 | 23.2 | 29.4 | 41.8 |
| Self-Forcing | 24.9 | 24.9 | 25.0 | 25.1 | 20.6 | 20.6 | 20.7 | 20.7 | 28.1 | 28.2 | 28.3 | 28.5 |
| HiAR | 24.9 | 24.9 | 25.0 | 25.1 | 24.9 | 24.9 | 25.0 | 25.1 | 32.9 | 33.0 | 33.1 | 33.3 |
| FlashForward | 24.9 | 25.0 | 25.2 | 25.4 | 24.9 | 24.9 | 25.0 | 25.1 | 32.9 | 33.0 | 33.1 | 33.3 |
| (a) Compared modes: KV source and final anchor source | ||||||
| Denoising steps | Final anchor | Anchor KV | Anchor seam | Chunk seam | Interior | Periodic |
| 50 | Planner | Planner | 3.93 | 2.71 | 1.32 | 7.66 |
| 50 | Planner | Renderer | 3.34 | 2.37 | 1.29 | 5.99 |
| 50 | Renderer (ours) | Planner | 1.83 | 2.27 | 1.32 | 3.54 |
| 4 | Planner | Renderer | 6.20 | 1.68 | 1.40 | 9.55 |
| 4 | Renderer (ours) | Planner | 2.43 | 2.50 | 1.26 | 5.61 |