To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching. To keep memory manageable, it detaches historical caches, preserving forward dependencies between chunks but severing the backward gradient paths. We introduce Connected Self Forcing, a training framework that reconnects gradient paths across autoregressive chunks, allowing feedback from later predictions to guide how earlier context is generated. These connections go beyond historical KV-writing: gradients pass through generated latents into the computations that produced them, linking the generation of earlier context to its use in later predictions. To make this connected training memory-efficient, we develop shortcut gradient replay, which recovers cross-chunk gradients without retaining the full rollout computation graph. Integrated with distribution matching distillation, Connected Self Forcing trains historical chunks according to both their direct supervision and their contribution to subsequent generation. Experiments on autoregressive video generation show improvements in long-horizon visual quality and temporal consistency, without changing the inference procedure.
Figures & tables
Figure 1: Connected Self Forcing restores gradient flow from later chunks to earlier ones, helping maintain visual consistency in long videos.
Figure 2: A controlled 1D Gaussian time-series toy experiment. We compare Self Forcing (SF), Connected Self Forcing (CSF), and Full BPTT under matched autoregressive training conditions. Given the initial context and known external inputs at each step, the target trajectory is deterministic, allowing MSE evaluation against ground truth. (a) Validation MSE. (b) Evaluation rollout MSE.
Figure 3: Overview of Connected Self Forcing. Self Forcing uses detached autoregressive rollouts, where each generated chunk x^i receives only direct DMD supervision. Connected Self Forcing restores cross-chunk feedback through Shortcut Gradient Replay : gradients from a later prediction are replayed to historical KV states KVj , then propagated through the KV writer GθKV and generator Gθ , which share parameters θ . Replay stops at earlier generated history to avoid recursive backpropagation through the earlier generation branches. This trains each chunk both for its own generation quality and as context for subsequent generation, without adding parameters or changing inference.
Causal CD Init
TF Init
Causal ODE Init
Metric
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
Aesthetic
0.5473
0.5592
0.5659
0.5990
0.5917
0.6152
0.5264
0.5772
0.5804
Background
0.9567
0.9583
0.9597
0.9601
0.9543
0.9721
0.9533
0.9625
0.9647
Imaging
0.6939
0.7031
0.7085
0.7046
0.7152
0.7286
0.7123
0.7224
0.7191
Motion
0.9884
0.9849
0.9848
0.9846
0.9764
0.9906
0.9817
0.9874
0.9896
Subject
0.9656
0.9665
0.9707
0.9695
0.9671
0.9821
0.9563
0.9733
0.9727
Table 1: 60-second chunk-wise generation on VBench-Long. Avg.(6) averages Aesthetic, Background, Imaging, Motion Smoothness, Subject Consistency, and Flickering; Dynamic Degree is reported separately. Best and second-best values are highlighted except for Dynamic Degree.
Causal CD Init
TF Init
Causal ODE Init
Metric
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
Aesthetic
0.5138
0.5220
0.5449
0.5639
0.5626
0.5948
0.5059
0.5357
0.5260
Background
0.9524
0.9539
0.9563
0.9554
0.9512
0.9673
0.9536
0.9559
0.9607
Imaging
0.6700
0.6607
0.6938
0.6937
0.7030
0.7161
0.6907
0.6965
0.7004
Motion
0.9887
0.9845
0.9850
0.9835
0.9761
0.9893
0.9838
0.9847
0.9895
Subject
0.9582
0.9581
0.9666
0.9647
0.9619
0.9786
0.9598
0.9658
0.9688
Table 2: 240-second chunk-wise generation on the fixed 128-prompt MovieGen Video Bench subset. Metrics and formatting follow Table 1 .
Figure 4: Qualitative comparison of 60-second chunk-wise autoregressive generation: SF, SGF, and CSF on VBench-Long prompts under Causal CD and TF initialization. CSF more consistently preserves subject appearance and scene composition throughout the rollout across both settings.
Figure 5: Qualitative comparison of 240-second chunk-wise autoregressive generation. We compare SF, SGF, and CSF on MovieGen-128 prompts under Causal CD and TF initialization. Over four-minute rollouts, CSF better preserves subject identity and scene semantics, with less accumulated drift in both the kangaroo and pianist examples.
Method
Aes. ↑
Back. ↑
Imag. ↑
Mot. ↑
Subj. ↑
Flick. ↑
Dyn.
Avg.(6) ↑
Full BPTT
0.5597
0.9562
0.6808
0.9857
0.9661
0.9688
0.5742
0.8529
Serial recursion
0.5432
0.9628
0.6907
0.9887
0.9753
0.9726
0.4032
0.8556
CSF ( λ=0.5 )
0.5471
0.9498
0.6852
0.9769
0.9486
0.9497
0.8484
0.8429
CSF (Ours)
0.5659
0.9597
0.7085
0.9848
0.9707
0.9637
0.5976
0.8589
Table 3: Ablation study on 60-second chunk-wise generation with Causal CD initialization. Avg.(6) and highlighting conventions follow Table 1 .
Figure 6: Qualitative ablation on 60-second chunk-wise generation with Causal CD initialization.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Layers / width / heads / history window
2 / 32 / 4 / 4
Context / training rollout / evaluation rollout
4 / 20 / 124 positions
Generator updates / fake updates per generator update
500 / 5
Batch size / precision
64 / float32
Generator / fake learning rate
5×10−7 / 8×10−7
AdamW betas / weight decay / gradient clipping
(0,0.999) / 0.01 / 5
Appendix
Table 4: Shared toy configuration. Budgets denote update counts.
Setting
Method
Device
Allocated
Reserved
Time ( × SF)
Chunk-wise
SF
103.65
70.64
98.46
1.00 ×
Full BPTT
148.65
118.70
143.45
1.15 ×
CSF
88.46
59.02
83.26
1.27 ×
Frame-wise
SF
184.48
149.44
179.28
1.00 ×
Full BPTT
OOM
–
–
–
CSF
88.42
59.01
83.22
1.41 ×
Appendix
Table 5: Training efficiency and peak single-GPU memory (GiB). Device denotes sampled NVML usage; Allocated and Reserved denote PyTorch peaks. Time is normalized to SF separately within each setting. Full BPTT’s frame-wise OOM has no completed-run peaks or time.
Causal CD
TF
Causal ODE
Metric
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
Aesthetic
0.6430
0.6393
0.6515
0.6593
0.6527
0.6726
0.6566
0.6496
0.6490
Appear.
0.1890
0.1871
0.1902
0.1895
0.1916
0.1886
0.1939
0.1951
0.1963
Background
0.9554
0.9551
0.9591
0.9457
0.9349
0.9739
0.9679
0.9327
0.9612
Color
0.9327
0.9099
0.8963
0.8876
0.9021
0.8987
0.9034
0.8908
0.9090
Action
0.7180
0.7500
0.7300
0.7460
0.7540
0.7620
0.7520
0.7640
0.7660
Appendix
Table 6: Full 16-dimensional VBench results for 5-second chunk-wise generation under three causal initializations. Avg.(6) is the unweighted mean of Aesthetic Quality, Background Consistency, Imaging Quality, Motion Smoothness, Subject Consistency, and Flickering, excluding Dynamic Degree. Except for Dynamic Degree, best and second-best results within each initialization are shown in bold and underlined , respectively.
Figure 7: Additional qualitative comparison of 5-second chunk-wise generation. We compare SF, SGF, and CSF under Causal CD, TF, and Causal ODE initialization using uniformly spaced snapshots.
Figure 8: Additional qualitative comparison of 60-second chunk-wise generation. We compare SF, SGF, and CSF under Causal CD, TF, and Causal ODE initialization using uniformly spaced snapshots throughout the rollout.
Figure 9: Additional qualitative comparison of 240-second chunk-wise generation. We compare SF, SGF, and CSF under Causal CD, TF, and Causal ODE initialization using uniformly spaced snapshots over the four-minute rollout.
Horizon
Init.
Method
Aes. ↑
Back. ↑
Imag. ↑
Mot. ↑
Subj. ↑
Flick. ↑
Dyn.
Avg.(6) ↑
SF
0.5159
0.9453
0.6366
0.9614
0.9366
0.9376
0.9411
0.8222
Causal CD
SGF
0.5103
0.9517
0.6408
0.9834
0.9475
0.9659
0.7234
0.8333
60 s
CSF (Ours)
0.5170
0.9586
0.6805
0.9861
0.9671
0.9699
0.6589
0.8465
SF
0.6013
0.9601
0.7380
0.9777
0.9705
0.9629
0.7927
0.8684
TF
SGF
0.5866
0.9533
0.7234
0.9773
0.9632
0.9478
0.8782
0.8586
CSF (Ours)
0.5871
0.9634
0.7328
0.9869
0.9763
0.9723
0.6516
0.8698
Appendix
Table 7: Frame-wise long-horizon quantitative comparison at 60 and 240 seconds. We compare SF, SGF, and CSF under Causal CD and TF initialization using the same evaluation protocol as the chunk-wise experiments. Avg.(6) and highlighting conventions follow Table 6 .
Figure 10: Additional qualitative comparison of 60-second frame-wise generation. We compare SF, SGF, and CSF under Causal CD and TF initialization using uniformly spaced snapshots throughout the rollout.
Figure 11: Additional qualitative comparison of 240-second frame-wise generation. We compare SF, SGF, and CSF under Causal CD and TF initialization using uniformly spaced snapshots over the four-minute rollout.
Initialization
Duration
N
CSF − SF
CSF − SGF
(a) Chunk-wise
Causal CD
60 s
40
+0.519[−0.228,+1.250]
+0.261[−0.350,+0.826]
240 s
128
+1.000[+0.657,+1.328]
+1.144[+0.870,+1.426]
TF
60 s
40
+1.326[+0.798,+1.861]
+1.977[+1.424,+2.540]
240 s
128
+1.560[+1.287,+1.829]
+2.065[+1.739,+2.395]
Causal ODE
60 s
40
+1.813[+1.278,+2.390]
+0.068[−0.385,+0.528]
Appendix
Table 8: Paired bootstrap differences in Avg.(6) for the reported long-video settings. Each entry is the difference computed from unrounded per-video scores with its 95% percentile interval, multiplied by 100 (score points). Positive values favor CSF; intervals are not adjusted for multiple comparisons.
Duration
Method
Aes. ↑
Back. ↑
Imag. ↑
Mot. ↑
Subj. ↑
Flick. ↑
Dyn.
Avg.(6) ↑
5 s
CSF
0.6515
0.9591
0.7117
0.9853
0.9703
0.9900
0.6028
0.8780
+ GradProj
0.6365
0.9471
0.7075
0.9822
0.9606
0.9775
0.7083
0.8686
60 s
CSF
0.5659
0.9597
0.7085
0.9848
0.9707
0.9637
0.5976
0.8589
+ GradProj
0.5436
0.9543
0.7175
0.9835
0.9617
0.9614
0.7605
0.8537
240 s
CSF
0.5449
0.9563
0.6938
0.9850
0.9666
0.9656
0.6899
0.8520
+ GradProj
0.5254
0.9527
0.6795
0.9832
0.9598
0.9630
0.7997
0.8439
Appendix
Table 9: Effect of conflict-aware gradient projection on CSF under Causal CD at 5, 60, and 240 seconds. GradProj removes the component of gcross that is negatively aligned with gDMD .
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.
Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint optimization of visual quality and temporal consistency. We introduce Self Gradient Forcing Plus (SGF+), which assigns separate parameters to context writing and denoising while preserving their interaction through causal attention. Both roles are jointly optimized using the original generation objective without auxiliary losses, with context writing supervised through its contribution to future predictions. This simple change improves visual quality and long-horizon consistency over the evaluated baselines in both framewise and chunkwise generation, without additional video training data or a longer training horizon. Trained on only 5s rollouts, SGF+ supports continuous generation for up to 24 hours without long-video fine-tuning. These results highlight role-specific parameterization as an effective design principle for high-quality autoregressive video generation and native long-horizon extrapolation.
Zihan Su, Junhao Zhuang, Yaowei Li +10
Tsinghua University · Joy Future Academy, JD · The Chinese University of Hong Kong
Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response R(Δ,tdenoise), a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from R, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.