Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint optimization of visual quality and temporal consistency. We introduce Self Gradient Forcing Plus (SGF+), which assigns separate parameters to context writing and denoising while preserving their interaction through causal attention. Both roles are jointly optimized using the original generation objective without auxiliary losses, with context writing supervised through its contribution to future predictions. This simple change improves visual quality and long-horizon consistency over the evaluated baselines in both framewise and chunkwise generation, without additional video training data or a longer training horizon. Trained on only 5s rollouts, SGF+ supports continuous generation for up to 24 hours without long-video fine-tuning. These results highlight role-specific parameterization as an effective design principle for high-quality autoregressive video generation and native long-horizon extrapolation.
Figures & tables
Figure 2 : Method lineage for autoregressive video generation. SF addresses train–test history mismatch via self-generated rollouts. SGF restores context-writing gradients, while SGF+ decouples context writing from denoising to resolve their gradient conflict.
Figure 3 : Visualization of gradient conflict. We analyze SGF gradients over 128 prompts and 4 denoising timesteps, separately for Attention and FFN. (a) Angular-distance t-SNE visualizes the distributions of context-writing and denoising gradients. t0=0 denotes context writing at timestep 0, while t1,t2,t3,t4 denote denoising at timesteps 250, 500, 750, and 1000, respectively. (b) Gradients are aggregated to compute the mean directions. (c) At each denoising timestep, we compute the paired cosine similarity between the denoising and context-writing gradients for the same prompt. Layer-wise gradient distributions are provided in Appendix E .
Figure 4 : Training pipeline of SGF and SGF+. Pass 1 performs a no-gradient autoregressive rollout and records detached clean context latents X together with noisy target latents Z⋆ at a sampled denoising exit. Pass 2 reconstructs the forward computation from the recorded latents, allowing future-generation losses to supervise context writing through differentiable context KV states. SGF accumulates context-writing and denoising gradients on shared parameters θ , whereas SGF+ assigns separate parameters θc and θd to the two roles.
Figure 5 : Qualitative comparisons. The top two cases show chunkwise generation, and the bottom case shows framewise generation. Additional qualitative comparisons covering a wider variety of scenes are provided in Appendix F .
Method
Subject
Background
Flickering
Motion
Dynamics
Aesthetics
Imaging
60s Generation Horizon · Chunkwise
Self Forcing
94.97
94.85
98.88
96.69
93.46
58.09
68.12
Self Gradient Forcing
98.17
97.12
98.94
98.53
64.16
65.38
71.06
Self Gradient Forcing Plus
98.48
97.41
99.35
98.69
63.93
66.63
71.51
60s Generation Horizon · Framewise
Self Forcing
97.62
97.42
99.08
98.55
63.89
64.80
70.50
Table 1 : Framewise and chunkwise long-horizon metrics at 60s and 240s. The 60s setting uses VBench-Long prompts, and the 240s setting uses MovieGen-128 prompts. All methods share the same initialization, prompt set, and random seed within each setting.
Method
Subject
Background
Flickering
Motion
Dynamics
Aesthetics
Imaging
Attention only
95.37
95.43
99.30
97.40
93.41
58.49
68.18
FFN only
97.43
96.48
98.33
98.27
65.37
63.49
70.41
Full model (Ours)
98.48
97.41
99.35
98.69
63.93
66.63
71.51
Table 2 : Ablation of modules for parameter separation. Attention only and FFN only apply parameter separation only to Attention and FFN modules, respectively. The full model applies role-specific separation to all model parameters.
Method
Train Peak Memory
Train Stable Memory
Train Time / step
Infer Memory
Infer Time / 81 frames
Self Forcing
86.36 GB
86.19 GB
10.02 s
24.85 GB
4.963 s
Self Gradient Forcing
97.83 GB
70.06 GB
11.79 s
24.85 GB
4.962 s
Self Gradient Forcing Plus
98.15 GB
73.60 GB
12.76 s
27.96 GB
4.969 s
Table 3 : Training and inference efficiency. Compared with SGF, SGF+ modestly increases training time per step, with a small increase in inference memory and nearly unchanged inference latency.
Figure 6 : Adjacent-timestep gradient directions at TF initialization. We measure gradient directions on 128 real videos over a 50-step diffusion schedule. The curve shows the mean angle between adjacent timesteps. The sharp 0→1 transition reveals distinct gradient directions for clean-context writing and denoising. Additional visualizations of context-writing and denoising gradients at TF initialization are provided in Appendix C .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Subject
Background
Flickering
Motion
Dynamics
Aesthetics
Imaging
Chunkwise
Self Forcing
93.74
94.74
97.83
97.58
89.72
66.35
69.46
Self Gradient Forcing
97.01
96.42
98.80
98.50
65.28
68.04
70.56
Self Gradient Forcing Plus
97.24
96.81
99.30
98.64
62.78
68.45
70.64
Framewise
Self Forcing
96.06
96.15
98.81
98.54
63.89
67.11
69.83
Appendix
Table 4 : Framewise and chunkwise VBench metrics at 5s. SF, SGF, and SGF+ are evaluated within the training horizon using standard VBench. Bold marks the best value within each generation mode.
Figure 10 : Layer-wise gradient distributions in FFN. We concatenate the up/down weight gradients within each block and use the settings in Figure 9 . Block 29 has zero context-writing gradients for all 512 prompt–timestep pairs, so its angular distances are undefined and no t-SNE is shown. sil. denotes the role silhouette score before projection, with higher values indicating stronger role separation.
Figure 11 : Additional qualitative comparisons in chunkwise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a fluffy monster gazing at a candle, a neon-lit Tokyo street, and a person reading.
Figure 12 : Additional qualitative comparisons in chunkwise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: an astronaut, a boy riding a bicycle, and a woman in a flower garden.
Figure 13 : Additional qualitative comparisons in chunkwise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a red balloon drifting through an urban street, a woman aboard a train, and a sunflower by a windowsill.
Figure 14 : Additional qualitative comparisons in chunkwise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a grandmother celebrating her birthday, a woman in front of fireworks, and a man eating noodles with chopsticks.
Figure 15 : Additional qualitative comparisons in framewise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a porcelain toilet in an old-fashioned bathroom, a white cat driving a red car, and a volcanic eruption in a coffee cup.
Figure 16 : Additional qualitative comparisons in framewise generation. SF, SGF, and SGF+ are compared in three scenes. From top to bottom: a kangaroo dancing disco, Art Deco lampposts, and an old wooden armchair in a living room.