Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: https://anonymous.4open.science/w/self-aligned-forcing/.
Figures & tables
Figure 1: History stage and gradient flow. The model fθ generates blocks from left to right, and the K/V of earlier blocks form the history of the last block, whose loss L is shown. Superscripts give noise levels ( ts : current block, tc : history, 0 : clean). Self-Forcing (a) and HiAR (b) detach the history, so the loss cannot reach it; Self Gradient Forcing (c) and SAF (d) keep this gradient with the history at tc=0 and tc=ts , respectively, but the former needs an extra no-grad rollout.
Figure 2: Two observations motivating differentiable noisy history. (a) A fixed LongLive model with varying history noise level tc ; imaging quality and dynamic degree ( Huang et al., 2024 ) are normalized to their values at tc=0 and tc=750 , respectively. (b) With all settings identical except whether gradients flow through the history K/V, differentiable history raises the gradient cosine similarity to a matched-weight bidirectional reference from 0.643 to 0.972 . Error bars show the range over multiple samples.
Figure 3: Overview of SAF. (a) Rollout is parallel across blocks and sequential only across stages (Section 3.3 ). (b) Stage alignment enables joint denoising with differentiable history in one block-causal forward pass (Section 3.2 ). (c) Reusing the K/V computed during denoising as history allows multi-GPU stage pipelining and removes per-block recaching (Section 3.5 ).
Table 4
Figure 6: Chunkwise Interactive qualitative comparison. Ours tracks the changing ten-second prompts with faithful semantics and stable visual quality, while competing methods exhibit quality degradation, object-count errors, or missed interactions.
Figure 7: Chunkwise MovieGen-100s qualitative comparison. Ours sustains the long-horizon table-wiping action with stable scene structure, whereas competing methods accumulate artifacts, duplicate subjects, or produce inconsistent reflections.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Two-pass training of SDF. (a) Pass 1 rolls out all blocks serially without gradients, recaches the sink from XΩ(d) , and records the trajectory; the history of each block is read at the stage chosen by the random, fixed, or aligned policy. (b) Pass 2 replays the recorded states in one differentiable forward pass over a context stream C and a query stream D , so the DMD loss on the queries reaches the context K/V they attend to. Under the aligned policy ( ci=d ), context and query coincide and the two streams collapse into one. For clarity, the attention masks omit the clean sink; as in SAF, block i also attends to the clean K/V of the preceding sink blocks, KVΩ,<i (Section 3.4 ). At the top of (a), Xi is block i , whose K/V enters the history at the bold stage; the context and query streams in (b) hold its recorded states Xi(ci) and Xi(d) , and cg is shared by context group g (Eq. 7 ).
Benchmark
Prompts
Samples
Length
Dimensions
Aggregate metrics
VBench
944
5
5 s
16
Total, Quality, Semantic
Interactive
100×6
1
60 s
7
Quality, ViCLIP
MovieGen-100s
128
1
100 s
7
Quality
Appendix
Table 5: Evaluation benchmarks. All models are trained on 5-second clips, so Interactive and MovieGen-100s test generation at 12× and 20× the training length. Each Interactive video is driven by six prompts in turn, each lasting 10 seconds, and the benchmark provides 100 such prompt sequences ( 100×6 ). Samples is the number of videos per prompt (per prompt sequence for Interactive), and Dimensions is the number of VBench dimensions scored.
Figure 9: Additional chunkwise Interactive comparison. Zoom in for better visualization.
Figure 10: Additional chunkwise MovieGen-100s example 1. Zoom in for better visualization.
Figure 11: Additional chunkwise MovieGen-100s example 2. Zoom in for better visualization.
Figure 12: Additional framewise Interactive comparison. Zoom in for better visualization.
Figure 13: Additional framewise MovieGen-100s example 1. Zoom in for better visualization.
Figure 14: Additional framewise MovieGen-100s example 2. Zoom in for better visualization.
Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. While recent works address this via post-training, they typically rely on a bidirectional teacher model or discriminator. To achieve an end-to-end solution, we introduce Resampling Forcing, a teacher-free framework that enables training autoregressive video models from scratch and at scale. Central to our approach is a self-resampling scheme that simulates inference-time model errors on history frames during training. Conditioned on these degraded histories, a sparse causal mask enforces temporal causality while enabling parallel training with frame-level diffusion loss. To facilitate efficient long-horizon generation, we further introduce history routing, a parameter-free mechanism that dynamically retrieves the top-k most relevant history frames for each query. Experiments demonstrate that our approach achieves performance comparable to distillation-based baselines while exhibiting superior temporal consistency on longer videos owing to native-length training.
Yuwei Guo, Ceyuan Yang, Hao He +5
The Chinese University of Hong Kong, Hong Kong, China · ByteDance Seed, China · ByteDance, China
Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long horizons, manifesting as color drift, motion stagnation, and eventual visual collapse. In this paper, we characterize this phenomenon from a frequency-domain perspective: error accumulation appears as a pronounced energy drift in the low-frequency bands. We further investigate the effectiveness of attention sink in the frequency domain, and find that it improves the video quality by alleviating the spectral energy drift to some extent, but cannot fully resolve it. Motivated by the above analysis, we propose FreqForcing, a training-free framework that addresses error accumulation in long-video generation via Spectral Self-Anchoring (SSA). The proposed SSA leverages the low-frequency components of anchor attention to maintain long-horizon visual stability, while preserving dynamic motion through the high-frequency components of local attention. Our FreqForcing extends Self-Forcing pretrained on 5s clips to two-minute generation, achieving 24x extrapolation. Extensive experiments show that FreqForcing outperforms existing training-free methods quantitatively and qualitatively while remaining competitive with representative training-based approaches.
Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous video stream. However, streaming video diffusion models introduce a fundamental train-inference mismatch: inference follows a specialized denoising order, whereas advanced training strategies typically require diverse noise-level configurations. To address this trade-off between train-inference consistency and training coverage, we reformulate the video diffusion sampling as a frame-indexed stochastic process over noise levels. Within this stochastic process space, we construct a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling. We further introduce a joint calibration algorithm and a temporal correlative sampling algorithm to ensure trajectory smoothness and cross-frame correlation. Building on these designs, we propose Stream Forcing, a unified training framework for streaming video generation that balances training sufficiency and inference efficiency. Extensive experiments demonstrate that Stream Forcing significantly improves generation quality with a 36.6% FVD improvement on the UCF-101 benchmark. Furthermore, our method facilitates robust zero-shot extrapolation to long-horizon video generation with a 27.9% FVD improvement on the UCF-101 benchmark.
Yueting Zhu, Yuehao Song, Kaicheng Zhang +5
Huazhong University of Science & Technology · Anyverse Dynamics · Horizon Robotics