Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects or modifies cached key-value (KV) entries to mitigate drifting. However, we observe that KV conditioning alone is insufficient as it assumes cached KV entries remain in-distribution. This assumption fails beyond the training horizon: nothing constrains the construction of KV entries during rollout, giving rise to the KV-provenance problem where cached entries themselves become out-of-distribution (OOD). To address this, we propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations. Its key mechanism, self-caching, prevents OOD KV entries at their source. Each chunk is cached without attending to prior KV entry, keeping the rolling window exactly in-distribution. Consequently, ID-Forcing seamlessly extends short-horizon models to minute-scale video generation. Extensive evaluations show that our method remains competitive on standard video generation benchmark while substantially outperforming prior work in mitigating drifting, as validated by both our drift metrics and a user study.
Figures & tables
Figure 1: ID-Forcing extends autoregressive video generation far beyond the training horizon. In-distribution KV operations suppress color and motion drift, enabling stable minute-scale generation.
Figure 2: Same conditioning, different provenance. We generate the next chunk x7 from the same conditioning window κ1:6 , varying only how the last chunk of the horizon, x6 , is cached into κ6 . Case 1 (self-caching): x6 attends to no prior entry, and x7 shows no drifting. Case 2 (self-forcing caching): x6 attends to a window never seen in training, and x7 degrades. Case 3 (sink caching): keeping the first chunk does not help, as the window is still unseen in training. Thus, how a chunk is cached alone can determine whether the next chunk drifts.
Figure 3: KV operations comparison. Both illustrations show the caching window of each KV entry and the conditioning window of size L used to generate chunk xi . Left : Self-Forcing caches every entry autoregressively, so entries beyond the horizon are cached under windows unseen in training. Right : ID-Forcing caches the earlier entries with self-caching and the rest autoregressively on top of them, so every entry is cached under a window seen in training.
Method
Aesthetic Quality ↑
Background Consistency ↑
Imaging Quality ↑
Motion Smoothness ↑
Subject Consistency ↑
Dynamic Degree ↑
Color Drift ↑
Motion Drift ↑
120 seconds
Results on Self-Forcing
Self-Forcing
49.95
96.17
61.60
98.26
96.20
30.04
25.54
46.88
Deep Forcing
57.26
96.28
65.94
98.21
97.22
45.49
59.65
71.88
∞ -RoPE
55.54
95.58
67.42
97.64
96.18
62.30
54.15
85.16
MemRoPE
55.45
95.78
67.70
97.69
96.32
62.90
59.25
82.81
Table 1: Quantitative evaluation on long video generation. We report VBench-Long ( Huang et al., 2024 ) metrics and drift metrics (color drift and motion drift) on 120- and 240-second videos, with methods grouped by base model.
Figure 4: Qualitative comparison on 2-minute videos. Full videos are available at this link .
Method
Color Cons.
Bg. Cons.
Subj. Cons.
Dyn. Deg.
Temp. Flick.
Over. Pref.
Self-Forcing
100.0
99.0
96.0
93.0
98.0
100.0
Deep Forcing
90.0
84.0
85.0
81.0
97.0
95.0
∞ -RoPE
91.0
90.0
87.0
75.0
91.0
93.0
MemRoPE
88.0
82.0
78.0
59.0
89.0
86.0
Table 2: User study. Preference rate (%) of ID-Forcing over each baseline in pairwise comparisons (Baseline vs. Ours). Each value denotes the percentage of participants who preferred ours ( ↑ ).
Figure 5: Qualitative ablation on 2-minute video. Components are added from top to bottom. Without self-caching, colors and textures still drift even with the sink and re-rotation, whereas the full ID-Forcing preserves the scene and color throughout the 2-minute video.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Aesthetic Quality ↑
Background Consistency ↑
Imaging Quality ↑
Motion Smoothness ↑
Subject Consistency ↑
Dynamic Degree ↑
Color Drift ↑
Motion Drift ↑
30 seconds
Results on Self-Forcing
Self-Forcing
56.85
96.06
67.48
98.37
96.33
44.58
37.45
60.94
Deep Forcing
57.79
96.05
67.00
98.13
96.79
59.22
63.61
82.03
∞ -RoPE
57.38
95.87
67.84
97.80
96.65
64.79
56.77
85.94
MemRoPE
57.32
96.08
67.73
98.05
96.78
66.35
65.61
89.84
Appendix
Table 3: Quantitative evaluation on VBench-Long ( Huang et al., 2024 ) (30s and 60s).
Method
Aesthetic Quality ↑
Background Consistency ↑
Imaging Quality ↑
Motion Smoothness ↑
Subject Consistency ↑
Dynamic Degree ↑
Color Drift ↑
Motion Drift ↑
Self-Forcing
49.95
96.17
61.60
98.26
96.20
30.04
25.54
46.88
+ Sink
55.23
95.59
64.68
96.88
95.70
46.90
48.65
67.19
+ Re-rotation
54.14
95.68
66.89
97.36
96.08
65.44
60.66
87.50
+ Self-Caching (ID-Forcing)
58.23
96.13
68.20
97.77
96.76
65.65
72.37
89.06
Appendix
Table 4: Quantitative results of ablation study on VBench-Long ( Huang et al., 2024 ) (120s).
Figure 6: Additional Qualitative results on 2-minute videos.
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limited KV-cache budget prevents the model from retaining the full history, while repeatedly conditioning on self-generated frames induces a context distribution shift that accumulates over time, leading to visual artifacts, quality degradation, and temporal drift. In this paper, we propose TetherCache, a training-free and plug-and-play cache management strategy for drift-resistant long video generation. TetherCache organizes the cache into sink, memory, and recent regions, and introduces two complementary mechanisms. First, GRAB (Gated Recall with Attention-Diversity Balancing) selects long-range memory frames using a gated score that combines attention-based relevance with temporal diversity, preserving informative yet diverse historical context under a fixed cache budget. Second, TAME (Trusted Alignment via Memory Editing) lightly edits newly recalled memory tokens by aligning their statistics to a trusted context distribution, reducing the pollution caused by drifted historical features. Built on Self-Forcing, TetherCache consistently improves long-video generation quality on VBench-Long across 30s, 60s, and 240s settings. In particular, for 240s generation, it substantially improves overall and semantic scores while reducing quality drift from 7.84 to 1.33, demonstrating its effectiveness for stable long-horizon autoregressive video diffusion.
Autoregressive video diffusion models support real-time synthesis but suffer from error accumulation and context loss over long horizons. We discover that attention heads in AR video diffusion transformers serve functionally distinct roles as local heads for detail refinement, anchor heads for structural stabilization, and memory heads for long-range context aggregation, yet existing methods treat them uniformly, leading to suboptimal KV cache allocation. We propose Head Forcing, a training-free framework that assigns each head type a tailored KV cache strategy: local and anchor heads retain only essential tokens, while memory heads employ a hierarchical memory system with dynamic episodic updates for long-range consistency. A head-wise RoPE re-encoding scheme further ensures positional encodings remain within the pretrained range. Without additional training, Head Forcing extends generation from 5 seconds to minute-level duration, supports multi-prompt interactive synthesis, and consistently outperforms existing baselines. Project Page: https://jiahaotian-sjtu.github.io/headforcing.github.io/.
Jiahao Tian, Yiwei Wang, Gang Yu +1
AGI Lab, Westlake University · University of California at Merced · StepFun
Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long horizons, manifesting as color drift, motion stagnation, and eventual visual collapse. In this paper, we characterize this phenomenon from a frequency-domain perspective: error accumulation appears as a pronounced energy drift in the low-frequency bands. We further investigate the effectiveness of attention sink in the frequency domain, and find that it improves the video quality by alleviating the spectral energy drift to some extent, but cannot fully resolve it. Motivated by the above analysis, we propose FreqForcing, a training-free framework that addresses error accumulation in long-video generation via Spectral Self-Anchoring (SSA). The proposed SSA leverages the low-frequency components of anchor attention to maintain long-horizon visual stability, while preserving dynamic motion through the high-frequency components of local attention. Our FreqForcing extends Self-Forcing pretrained on 5s clips to two-minute generation, achieving 24x extrapolation. Extensive experiments show that FreqForcing outperforms existing training-free methods quantitatively and qualitatively while remaining competitive with representative training-based approaches.