Autoregressive video diffusion enables causal video streaming without a bidirectional pass over the full clip, but existing few-step systems usually retain only the opening and most recent frames in a fixed-size KV cache. Once an event leaves this window, later frames can no longer attend to it, a failure we term mid-horizon forgetting. We present Memory Forcing, a few-step streaming method that preserves this missing history without increasing the cache size. Its Archive & Working Banks partition the cache into sink, archive, and working regions, retaining diverse intermediate events alongside recent motion under fixed memory. Because absolute temporal indices drift outside the training range, Bank-aware RoPE reassigns indices at attention time so each bank remains distinguishable. At 1.3B, Memory Forcing leads on longer clips, shows the smallest drop from 5s to 60s among methods reporting all four lengths, and preserves subjects and scenes through leave-and-return. The same design scales to Wan2.2 5B, producing more physically plausible, realistic, and dynamic videos and, to our knowledge, the first public 5B model on this forcing line.
Figures & tables
Figure 1: Qualitative samples from Memory Forcing. Each row is one generated clip, frames in time order. Subjects, motion, and scene layout stay consistent across the clip.
Figure 2: Archive & Working Banks, same cache each row. (a) Only the current window is filled. (b) The first chunk goes into sink and stays. (c) Later chunks enter working while slots remain. (d) When working is full, Similarity writes into archive. (e) FIFO then drops the oldest archive chunk.
Figure 3: Bank-aware RoPE assigns τ on one window. Absolute τabs lets later indices grow past the trained range. Contiguous τctg numbers from 0 . Collapsed τcol puts sink and archive at 0 . Stepped τstp puts sink at 0 , archive at 1 , working from 2 . Interpolated τitp places archive in (0,1) .
Figure 4: Further qualitative samples from Memory Forcing. Each row is one generated clip, frames in time order. Fine action and wide scenes keep identity and layout across the clip.
VBench (Total)
VBench-Complex (Quality)
Gen-Complex (Quality)
MovieGen (Quality)
Method
5s
15s
30s
60s
Δ↓5s→60s
5s
15s
30s
60s
Δ↓5s→60s
5s
15s
30s
60s
Δ↓5s→60s
5s
15s
30s
60s
Δ↓5s→60s
CausVid
82.99
—
—
—
—
80.83
—
—
—
—
80.35
—
—
—
—
79.62
—
—
—
—
NOVA
78.54
74.78
—
—
—
73.32
69.18
—
—
—
74.86
69.41
—
—
—
74.08
68.86
—
—
—
SkyReels-V2
82.65
80.11
—
—
—
81.08
75.07
—
—
—
78.83
71.93
—
—
—
79.88
75.35
—
—
—
Self-Forcing
84.09
—
—
—
—
81.37
—
—
—
—
79.99
—
—
—
—
81.09
—
—
—
—
LongLive
83.23
82.95
82.10
82.21
-1.02
79.73
80.01
79.07
78.81
-0.92
81.51
80.59
79.98
80.36
-1.15
80.15
79.79
79.20
79.13
-1.02
Table 1: Comparison on official VBench Total and Quality on VBench-Complex, Gen-Complex, and MovieGen at 5s, 15s, 30s, and 60s. CFODE and SFODE are the CausalForcing and Self-Forcing pretraining. Red is best and blue is second-best in each column.
Table 2: Ablations on Wan2.1-T2V-1.3B ( Wang et al., 2025a ) scored on official VBench ( Huang et al., 2024 ) at 30s and 60s. Every panel reports Total/Quality/Semantic at 16 FPS and 832×480 .
Figure 5: Qualitative comparison at 1.3B against four recent open-source models. Memory Forcing keeps the subject’s identity and the scene layout, while the baselines drift or collapse.
Figure 6: Memory Forcing at Wan2.1-T2V-1.3B and Wan2.2 5B on the same prompts. The 5B run is more physically plausible, more realistic, and more dynamic.
Streaming autoregressive diffusion makes minute-scale video synthesis practical, but its bounded context and fixed denoising schedule allocate resources uniformly across a highly non-stationary sequence. A rolling key-value cache forgets distant visual evidence even when that evidence remains important, while every generated chunk receives the same number of denoising passes irrespective of its actual difficulty. We introduce Surprise Forcing, a training-free framework that treats both limitations as online resource-allocation problems. A Surprise-Gated Memory Bank summarizes evicted frames with value-token descriptors, evaluates them using complementary global-deviation and nearest-neighbor novelty signals, and regulates admission through a feedback-controlled budget in normalized score space. Priority-based replacement and relevance-aware routing then keep the external memory compact and useful. In parallel, Surprise-Aware Denoising estimates chunk difficulty from the maximum adjacent-frame cosine distance after the first denoising pass and uses a local percentile scheduler to skip intermediate steps for comparatively easy chunks. Experiments on VBench, VBench-Long, and VBench-2.0 show that the proposed allocation strategy improves long-horizon consistency and visual quality while retaining real-time streaming throughput.
Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: https://anonymous.4open.science/w/self-aligned-forcing/.
Weiqiang Wang, Zhuokun Chen, Yusheng Dai +5
Monash University · Vivix AI · Lancaster University
Autoregressive video diffusion models support real-time synthesis but suffer from error accumulation and context loss over long horizons. We discover that attention heads in AR video diffusion transformers serve functionally distinct roles as local heads for detail refinement, anchor heads for structural stabilization, and memory heads for long-range context aggregation, yet existing methods treat them uniformly, leading to suboptimal KV cache allocation. We propose Head Forcing, a training-free framework that assigns each head type a tailored KV cache strategy: local and anchor heads retain only essential tokens, while memory heads employ a hierarchical memory system with dynamic episodic updates for long-range consistency. A head-wise RoPE re-encoding scheme further ensures positional encodings remain within the pretrained range. Without additional training, Head Forcing extends generation from 5 seconds to minute-level duration, supports multi-prompt interactive synthesis, and consistently outperforms existing baselines. Project Page: https://jiahaotian-sjtu.github.io/headforcing.github.io/.
Jiahao Tian, Yiwei Wang, Gang Yu +1
AGI Lab, Westlake University · University of California at Merced · StepFun