cs.CVSep 29, 2026

Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History

Authors: Weiqiang Wang, Zhuokun Chen, Yusheng Dai, Boying Li, Yi Zhang, Hossein Rahmani, Qiuhong Ke, Jianfei Cai

Organizations: Monash University · Vivix AI · Lancaster University

Abstract

Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: https://anonymous.4open.science/w/self-aligned-forcing/.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

    Dec 17, 2025Yuwei Guo, Ceyuan Yang, Hao He +5Autoregressive Video Diffusion ModelsResampling

  2. FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

    Jul 29, 2026Jiatong Li, Leo Liang, Linghe Kong +1Autoregressive Video Diffusion ModelsAutoregressive Video Generation

  3. Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation

    Aug 11, 2026Yueting Zhu, Yuehao Song, Kaicheng Zhang +5Video Diffusion ModelsVideo Generation