To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching. To keep memory manageable, it detaches historical caches, preserving forward dependencies between chunks but severing the backward gradient paths. We introduce Connected Self Forcing, a training framework that reconnects gradient paths across autoregressive chunks, allowing feedback from later predictions to guide how earlier context is generated. These connections go beyond historical KV-writing: gradients pass through generated latents into the computations that produced them, linking the generation of earlier context to its use in later predictions. To make this connected training memory-efficient, we develop shortcut gradient replay, which recovers cross-chunk gradients without retaining the full rollout computation graph. Integrated with distribution matching distillation, Connected Self Forcing trains historical chunks according to both their direct supervision and their contribution to subsequent generation. Experiments on autoregressive video generation show improvements in long-horizon visual quality and temporal consistency, without changing the inference procedure.
Figures & tables
Figure 1: Connected Self Forcing restores gradient flow from later chunks to earlier ones, helping maintain visual consistency in long videos.
Figure 2: A controlled 1D Gaussian time-series toy experiment. We compare Self Forcing (SF), Connected Self Forcing (CSF), and Full BPTT under matched autoregressive training conditions. Given the initial context and known external inputs at each step, the target trajectory is deterministic, allowing MSE evaluation against ground truth. (a) Validation MSE. (b) Evaluation rollout MSE.
Figure 3: Overview of Connected Self Forcing. Self Forcing uses detached autoregressive rollouts, where each generated chunk x^i receives only direct DMD supervision. Connected Self Forcing restores cross-chunk feedback through Shortcut Gradient Replay : gradients from a later prediction are replayed to historical KV states KVj , then propagated through the KV writer GθKV and generator Gθ , which share parameters θ . Replay stops at earlier generated history to avoid recursive backpropagation through the earlier generation branches. This trains each chunk both for its own generation quality and as context for subsequent generation, without adding parameters or changing inference.
Causal CD Init
TF Init
Causal ODE Init
Metric
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
Aesthetic
0.5473
0.5592
0.5659
0.5990
0.5917
0.6152
0.5264
0.5772
0.5804
Background
0.9567
0.9583
0.9597
0.9601
0.9543
0.9721
0.9533
0.9625
0.9647
Imaging
0.6939
0.7031
0.7085
0.7046
0.7152
0.7286
0.7123
0.7224
0.7191
Motion
0.9884
0.9849
0.9848
0.9846
0.9764
0.9906
0.9817
0.9874
0.9896
Subject
0.9656
0.9665
0.9707
0.9695
0.9671
0.9821
0.9563
0.9733
0.9727
Table 1: 60-second chunk-wise generation on VBench-Long. Avg.(6) averages Aesthetic, Background, Imaging, Motion Smoothness, Subject Consistency, and Flickering; Dynamic Degree is reported separately. Best and second-best values are highlighted except for Dynamic Degree.
Causal CD Init
TF Init
Causal ODE Init
Metric
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
Aesthetic
0.5138
0.5220
0.5449
0.5639
0.5626
0.5948
0.5059
0.5357
0.5260
Background
0.9524
0.9539
0.9563
0.9554
0.9512
0.9673
0.9536
0.9559
0.9607
Imaging
0.6700
0.6607
0.6938
0.6937
0.7030
0.7161
0.6907
0.6965
0.7004
Motion
0.9887
0.9845
0.9850
0.9835
0.9761
0.9893
0.9838
0.9847
0.9895
Subject
0.9582
0.9581
0.9666
0.9647
0.9619
0.9786
0.9598
0.9658
0.9688
Table 2: 240-second chunk-wise generation on the fixed 128-prompt MovieGen Video Bench subset. Metrics and formatting follow Table 1 .
Figure 4: Qualitative comparison of 60-second chunk-wise autoregressive generation: SF, SGF, and CSF on VBench-Long prompts under Causal CD and TF initialization. CSF more consistently preserves subject appearance and scene composition throughout the rollout across both settings.
Figure 5: Qualitative comparison of 240-second chunk-wise autoregressive generation. We compare SF, SGF, and CSF on MovieGen-128 prompts under Causal CD and TF initialization. Over four-minute rollouts, CSF better preserves subject identity and scene semantics, with less accumulated drift in both the kangaroo and pianist examples.
Method
Aes. ↑
Back. ↑
Imag. ↑
Mot. ↑
Subj. ↑
Flick. ↑
Dyn.
Avg.(6) ↑
Full BPTT
0.5597
0.9562
0.6808
0.9857
0.9661
0.9688
0.5742
0.8529
Serial recursion
0.5432
0.9628
0.6907
0.9887
0.9753
0.9726
0.4032
0.8556
CSF ( λ=0.5 )
0.5471
0.9498
0.6852
0.9769
0.9486
0.9497
0.8484
0.8429
CSF (Ours)
0.5659
0.9597
0.7085
0.9848
0.9707
0.9637
0.5976
0.8589
Table 3: Ablation study on 60-second chunk-wise generation with Causal CD initialization. Avg.(6) and highlighting conventions follow Table 1 .
Figure 6: Qualitative ablation on 60-second chunk-wise generation with Causal CD initialization.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Layers / width / heads / history window
2 / 32 / 4 / 4
Context / training rollout / evaluation rollout
4 / 20 / 124 positions
Generator updates / fake updates per generator update
500 / 5
Batch size / precision
64 / float32
Generator / fake learning rate
5×10−7 / 8×10−7
AdamW betas / weight decay / gradient clipping
(0,0.999) / 0.01 / 5
Appendix
Table 4: Shared toy configuration. Budgets denote update counts.
Setting
Method
Device
Allocated
Reserved
Time ( × SF)
Chunk-wise
SF
103.65
70.64
98.46
1.00 ×
Full BPTT
148.65
118.70
143.45
1.15 ×
CSF
88.46
59.02
83.26
1.27 ×
Frame-wise
SF
184.48
149.44
179.28
1.00 ×
Full BPTT
OOM
–
–
–
CSF
88.42
59.01
83.22
1.41 ×
Appendix
Table 5: Training efficiency and peak single-GPU memory (GiB). Device denotes sampled NVML usage; Allocated and Reserved denote PyTorch peaks. Time is normalized to SF separately within each setting. Full BPTT’s frame-wise OOM has no completed-run peaks or time.
Causal CD
TF
Causal ODE
Metric
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
SF
SGF
CSF (Ours)
Aesthetic
0.6430
0.6393
0.6515
0.6593
0.6527
0.6726
0.6566
0.6496
0.6490
Appear.
0.1890
0.1871
0.1902
0.1895
0.1916
0.1886
0.1939
0.1951
0.1963
Background
0.9554
0.9551
0.9591
0.9457
0.9349
0.9739
0.9679
0.9327
0.9612
Color
0.9327
0.9099
0.8963
0.8876
0.9021
0.8987
0.9034
0.8908
0.9090
Action
0.7180
0.7500
0.7300
0.7460
0.7540
0.7620
0.7520
0.7640
0.7660
Appendix
Table 6: Full 16-dimensional VBench results for 5-second chunk-wise generation under three causal initializations. Avg.(6) is the unweighted mean of Aesthetic Quality, Background Consistency, Imaging Quality, Motion Smoothness, Subject Consistency, and Flickering, excluding Dynamic Degree. Except for Dynamic Degree, best and second-best results within each initialization are shown in bold and underlined , respectively.
Figure 7: Additional qualitative comparison of 5-second chunk-wise generation. We compare SF, SGF, and CSF under Causal CD, TF, and Causal ODE initialization using uniformly spaced snapshots.
Figure 8: Additional qualitative comparison of 60-second chunk-wise generation. We compare SF, SGF, and CSF under Causal CD, TF, and Causal ODE initialization using uniformly spaced snapshots throughout the rollout.
Figure 9: Additional qualitative comparison of 240-second chunk-wise generation. We compare SF, SGF, and CSF under Causal CD, TF, and Causal ODE initialization using uniformly spaced snapshots over the four-minute rollout.
Horizon
Init.
Method
Aes. ↑
Back. ↑
Imag. ↑
Mot. ↑
Subj. ↑
Flick. ↑
Dyn.
Avg.(6) ↑
SF
0.5159
0.9453
0.6366
0.9614
0.9366
0.9376
0.9411
0.8222
Causal CD
SGF
0.5103
0.9517
0.6408
0.9834
0.9475
0.9659
0.7234
0.8333
60 s
CSF (Ours)
0.5170
0.9586
0.6805
0.9861
0.9671
0.9699
0.6589
0.8465
SF
0.6013
0.9601
0.7380
0.9777
0.9705
0.9629
0.7927
0.8684
TF
SGF
0.5866
0.9533
0.7234
0.9773
0.9632
0.9478
0.8782
0.8586
CSF (Ours)
0.5871
0.9634
0.7328
0.9869
0.9763
0.9723
0.6516
0.8698
Appendix
Table 7: Frame-wise long-horizon quantitative comparison at 60 and 240 seconds. We compare SF, SGF, and CSF under Causal CD and TF initialization using the same evaluation protocol as the chunk-wise experiments. Avg.(6) and highlighting conventions follow Table 6 .
Figure 10: Additional qualitative comparison of 60-second frame-wise generation. We compare SF, SGF, and CSF under Causal CD and TF initialization using uniformly spaced snapshots throughout the rollout.
Figure 11: Additional qualitative comparison of 240-second frame-wise generation. We compare SF, SGF, and CSF under Causal CD and TF initialization using uniformly spaced snapshots over the four-minute rollout.
Initialization
Duration
N
CSF − SF
CSF − SGF
(a) Chunk-wise
Causal CD
60 s
40
+0.519[−0.228,+1.250]
+0.261[−0.350,+0.826]
240 s
128
+1.000[+0.657,+1.328]
+1.144[+0.870,+1.426]
TF
60 s
40
+1.326[+0.798,+1.861]
+1.977[+1.424,+2.540]
240 s
128
+1.560[+1.287,+1.829]
+2.065[+1.739,+2.395]
Causal ODE
60 s
40
+1.813[+1.278,+2.390]
+0.068[−0.385,+0.528]
Appendix
Table 8: Paired bootstrap differences in Avg.(6) for the reported long-video settings. Each entry is the difference computed from unrounded per-video scores with its 95% percentile interval, multiplied by 100 (score points). Positive values favor CSF; intervals are not adjusted for multiple comparisons.
Duration
Method
Aes. ↑
Back. ↑
Imag. ↑
Mot. ↑
Subj. ↑
Flick. ↑
Dyn.
Avg.(6) ↑
5 s
CSF
0.6515
0.9591
0.7117
0.9853
0.9703
0.9900
0.6028
0.8780
+ GradProj
0.6365
0.9471
0.7075
0.9822
0.9606
0.9775
0.7083
0.8686
60 s
CSF
0.5659
0.9597
0.7085
0.9848
0.9707
0.9637
0.5976
0.8589
+ GradProj
0.5436
0.9543
0.7175
0.9835
0.9617
0.9614
0.7605
0.8537
240 s
CSF
0.5449
0.9563
0.6938
0.9850
0.9666
0.9656
0.6899
0.8520
+ GradProj
0.5254
0.9527
0.6795
0.9832
0.9598
0.9630
0.7997
0.8439
Appendix
Table 9: Effect of conflict-aware gradient projection on CSF under Causal CD at 5, 60, and 240 seconds. GradProj removes the component of gcross that is negatively aligned with gDMD .