Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint optimization of visual quality and temporal consistency. We introduce Self Gradient Forcing Plus (SGF+), which assigns separate parameters to context writing and denoising while preserving their interaction through causal attention. Both roles are jointly optimized using the original generation objective without auxiliary losses, with context writing supervised through its contribution to future predictions. This simple change improves visual quality and long-horizon consistency over the evaluated baselines in both framewise and chunkwise generation, without additional video training data or a longer training horizon. Trained on only 5s rollouts, SGF+ supports continuous generation for up to 24 hours without long-video fine-tuning. These results highlight role-specific parameterization as an effective design principle for high-quality autoregressive video generation and native long-horizon extrapolation.
Figures & tables
Figure 2 : Method lineage for autoregressive video generation. SF addresses train–test history mismatch via self-generated rollouts. SGF restores context-writing gradients, while SGF+ decouples context writing from denoising to resolve their gradient conflict.
Figure 3 : Visualization of gradient conflict. We analyze SGF gradients over 128 prompts and 4 denoising timesteps, separately for Attention and FFN. (a) Angular-distance t-SNE visualizes the distributions of context-writing and denoising gradients. t0=0 denotes context writing at timestep 0, while t1,t2,t3,t4 denote denoising at timesteps 250, 500, 750, and 1000, respectively. (b) Gradients are aggregated to compute the mean directions. (c) At each denoising timestep, we compute the paired cosine similarity between the denoising and context-writing gradients for the same prompt. Layer-wise gradient distributions are provided in Appendix E .
Figure 4 : Training pipeline of SGF and SGF+. Pass 1 performs a no-gradient autoregressive rollout and records detached clean context latents X together with noisy target latents Z⋆ at a sampled denoising exit. Pass 2 reconstructs the forward computation from the recorded latents, allowing future-generation losses to supervise context writing through differentiable context KV states. SGF accumulates context-writing and denoising gradients on shared parameters θ , whereas SGF+ assigns separate parameters θc and θd to the two roles.
Figure 5 : Qualitative comparisons. The top two cases show chunkwise generation, and the bottom case shows framewise generation. Additional qualitative comparisons covering a wider variety of scenes are provided in Appendix F .
Method
Subject
Background
Flickering
Motion
Dynamics
Aesthetics
Imaging
60s Generation Horizon · Chunkwise
Self Forcing
94.97
94.85
98.88
96.69
93.46
58.09
68.12
Self Gradient Forcing
98.17
97.12
98.94
98.53
64.16
65.38
71.06
Self Gradient Forcing Plus
98.48
97.41
99.35
98.69
63.93
66.63
71.51
60s Generation Horizon · Framewise
Self Forcing
97.62
97.42
99.08
98.55
63.89
64.80
70.50
Table 1 : Framewise and chunkwise long-horizon metrics at 60s and 240s. The 60s setting uses VBench-Long prompts, and the 240s setting uses MovieGen-128 prompts. All methods share the same initialization, prompt set, and random seed within each setting.
Method
Subject
Background
Flickering
Motion
Dynamics
Aesthetics
Imaging
Attention only
95.37
95.43
99.30
97.40
93.41
58.49
68.18
FFN only
97.43
96.48
98.33
98.27
65.37
63.49
70.41
Full model (Ours)
98.48
97.41
99.35
98.69
63.93
66.63
71.51
Table 2 : Ablation of modules for parameter separation. Attention only and FFN only apply parameter separation only to Attention and FFN modules, respectively. The full model applies role-specific separation to all model parameters.
Method
Train Peak Memory
Train Stable Memory
Train Time / step
Infer Memory
Infer Time / 81 frames
Self Forcing
86.36 GB
86.19 GB
10.02 s
24.85 GB
4.963 s
Self Gradient Forcing
97.83 GB
70.06 GB
11.79 s
24.85 GB
4.962 s
Self Gradient Forcing Plus
98.15 GB
73.60 GB
12.76 s
27.96 GB
4.969 s
Table 3 : Training and inference efficiency. Compared with SGF, SGF+ modestly increases training time per step, with a small increase in inference memory and nearly unchanged inference latency.
Figure 6 : Adjacent-timestep gradient directions at TF initialization. We measure gradient directions on 128 real videos over a 50-step diffusion schedule. The curve shows the mean angle between adjacent timesteps. The sharp 0→1 transition reveals distinct gradient directions for clean-context writing and denoising. Additional visualizations of context-writing and denoising gradients at TF initialization are provided in Appendix C .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Subject
Background
Flickering
Motion
Dynamics
Aesthetics
Imaging
Chunkwise
Self Forcing
93.74
94.74
97.83
97.58
89.72
66.35
69.46
Self Gradient Forcing
97.01
96.42
98.80
98.50
65.28
68.04
70.56
Self Gradient Forcing Plus
97.24
96.81
99.30
98.64
62.78
68.45
70.64
Framewise
Self Forcing
96.06
96.15
98.81
98.54
63.89
67.11
69.83
Appendix
Table 4 : Framewise and chunkwise VBench metrics at 5s. SF, SGF, and SGF+ are evaluated within the training horizon using standard VBench. Bold marks the best value within each generation mode.
Figure 10 : Layer-wise gradient distributions in FFN. We concatenate the up/down weight gradients within each block and use the settings in Figure 9 . Block 29 has zero context-writing gradients for all 512 prompt–timestep pairs, so its angular distances are undefined and no t-SNE is shown. sil. denotes the role silhouette score before projection, with higher values indicating stronger role separation.
Figure 11 : Additional qualitative comparisons in chunkwise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a fluffy monster gazing at a candle, a neon-lit Tokyo street, and a person reading.
Figure 12 : Additional qualitative comparisons in chunkwise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: an astronaut, a boy riding a bicycle, and a woman in a flower garden.
Figure 13 : Additional qualitative comparisons in chunkwise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a red balloon drifting through an urban street, a woman aboard a train, and a sunflower by a windowsill.
Figure 14 : Additional qualitative comparisons in chunkwise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a grandmother celebrating her birthday, a woman in front of fireworks, and a man eating noodles with chopsticks.
Figure 15 : Additional qualitative comparisons in framewise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a porcelain toilet in an old-fashioned bathroom, a white cat driving a red car, and a volcanic eruption in a coffee cup.
Figure 16 : Additional qualitative comparisons in framewise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a kangaroo dancing disco, Art Deco lampposts, and an old wooden armchair in a living room.
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.
Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and structural drift. In this paper, we investigate this failure mode from a spectral kinematic perspective and identify discrete latent acceleration as an effective signal for revealing unstable high-frequency temporal perturbations. To this end, we propose SAGA, a training-free \textbf{\textit{s}}table \textbf{\textit{a}}cceleration \textbf{\textit{g}}uidance approach for \textbf{\textit{a}}utoregressive video generation. SAGA integrates an acceleration domain spectral guidance objective based on finite-window Slepian projections with a structured autoregressive noise initialization strategy that suppresses short-range temporal correlations while preserving long-range motion structure. Without retraining or modifying the backbone, SAGA can be directly applied to existing chunk-wise autoregressive diffusion models, which is the prevalent setting for high-quality generation. Extensive experiments show that SAGA consistently improves temporal quality across multiple autoregressive diffusion models. On Self-Forcing, SAGA improves Temporal Quality from 97.30 to 97.91 and Image Quality from 69.60 to 70.51. Moreover, spectral analysis and human preference studies demonstrate that SAGA reduces temporal instability while maintaining visual fidelity.
Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le +2
University of Science, VNU-HCM, Vietnam · Vietnam National University, Ho Chi Minh City, Vietnam · University of Dayton, U.S.A.
Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response R(Δ,tdenoise), a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from R, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.