LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation
Organizations: KAIST · Georgia Tech
Abstract
World models, game simulators, and long-take video creation require coherent scene evolution and sustained dynamics over extended durations. Autoregressive (AR) video diffusion provides a natural framework for long-horizon generation, yet extended rollouts often become near-static or lose visual quality. We hypothesize that these failures reflect the limited guidance provided by short-video supervision on how ongoing scene dynamics develops over longer durations. This motivates us to introduce LongTake, a two-stage training pipeline built around Long-Horizon Teacher Forcing (TF) on curated real long videos. Long-Horizon TF trains the AR model to predict later frames conditioned on long ground-truth video prefixes, extending direct supervision beyond the short training horizon. This supervision is designed to help the model sustain dynamics and preserve visual quality during long-horizon generation. Our central finding is that this training stage strengthens direct initialization for distribution matching distillation (DMD) under student self-rollout, without the intermediate few-step distillation stage used in standard pipelines. Under the same five-second DMD training setup, our initialization yields substantially higher dynamic degree than short horizon TF initialization on 30-second rollouts at comparable aesthetic quality, and surpasses the evaluated baselines in both measures. Hybrid DMD further reuses this teacher to extend supervision to later frames of the self-rollout while retaining bidirectional joint supervision over the initial window. On long-horizon self-rollouts, LongTake lies on the Pareto front of dynamic degree and aesthetic quality, and Hybrid DMD attains the highest dynamic degree among evaluated methods at both 30s and 60s.
Figures & tables
| 30s | 60s | |||||||||||
| Method | Subj. | Dyn. | Aesth. | Quality | IQ | VLM | Subj. | Dyn. | Aesth. | Quality | IQ | VLM |
| Short-video supervision | ||||||||||||
| Self Forcing ( Huang et al., 2025a ) | 2.69 | |||||||||||
| Causal Forcing ( Zhu et al., 2026 ) | 2.07 | |||||||||||
| LongLive ( Yang et al., 2026 ) | 3.83 | |||||||||||
| Reward Forcing ( Lu et al., 2026 ) | 3.75 | |||||||||||
| Subj. | Dyn. | Aesth. | Quality | |
|---|---|---|---|---|
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Long-Horizon TF | Joint DMD | Hybrid DMD |
|---|---|---|---|
| Iterations | 3,000 | 1,200 / 800 | 400 |
| Generator learning rate | |||
| Fake-score learning rate | — | ||
| Effective batch size | 64 | 8 | 8 |
| Gradient accumulation | 8 | 1 | 1 |
| Optimizer | AdamW | AdamW | AdamW |
| Source | Prefix length | Probability |
|---|---|---|
| Causal Forcing clean GT | 0 | 0.10 |
| OpenVidHD | 0–15 | 0.20 |
| OpenVidHD | 18–39 | 0.45 |
| OpenVidHD | 42–60 | 0.20 |
| OpenVidHD | 63–99 | 0.05 |
| 30s | ||||||||
| Method | Subj. | Backg. | Motion | Dyn. | Aesth. | Imaging | Flicker | Quality |
| Short-video supervision | ||||||||
| Self Forcing ( Huang et al., 2025a ) | ||||||||
| Causal Forcing ( Zhu et al., 2026 ) | ||||||||
| LongLive ( Yang et al., 2026 ) | ||||||||
| Reward Forcing ( Lu et al., 2026 ) | ||||||||
| Dimension | 30s | 60s |
|---|---|---|
| Subject consistency | 128 | 72 |
| Background consistency | 128 | 86 |
| Motion smoothness | 128 | 72 |
| Dynamic degree | 128 | 72 |
| Aesthetic quality | 128 | 93 |
| Imaging quality | 128 | 93 |
| Score | Exposure criterion |
|---|---|
| 0 | Near-total whiteout or blackout makes the scene unreadable. |
| 1 | Widespread exposure failure severely reduces visibility. |
| 2 | Persistent highlight or shadow clipping causes substantial detail loss. |
| 3 | Exposure problems have limited spatial extent or duration. |
| 4 | Occasional local exposure flaws cause little visibility loss. |
| 5 | Balanced exposure preserves visibility without disruptive clipping or darkening. |
| Prefix (video frames) | Teacher | VR | IF | DINO | RGB L1 |
|---|---|---|---|---|---|
| 81 | Short TF | 17.26 | 96 | 0.9904 | 0.0202 |
| Long TF | 18.28 | 100 | 0.9919 | 0.0176 | |
| 321 | Short TF | 16.60 | 96 | 0.9930 | 0.0183 |
| Long TF | 17.83 | 100 | 0.9941 | 0.0166 |