Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
Organizations: MMLab, The Chinese University of Hong Kong · Tencent AIPD
Abstract
Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in the surrounding context merely to preserve temporal consistency. To provide a clearer visual-quality signal, we introduce Rollout-Marginal Distillation (RMD). RMD retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context. To compensate for the lack of temporal context in independent chunk scoring, RMD subsequently applies video-level DMD to restore temporal coherence. Extensive experiments demonstrate that RMD maintains high visual quality far beyond its training horizon and outperforms video-level DMD baselines. Code and video results are available at https://cjeen.github.io/RMD
Figures & tables
| Chunk size | Method | Total | Quality | Semantic |
|---|---|---|---|---|
| Self Forcing | 70.94 | 77.89 | 43.17 | |
| 1 | Causal Forcing | 69.88 | 78.34 | 36.03 |
| RMD (ours) | 81.26 | 84.52 | 68.24 | |
| Self Forcing | 77.44 | 81.99 | 59.22 | |
| 3 | Causal Forcing | 74.88 | 80.53 | 52.27 |
| RMD (ours) | 81.48 | 85.03 | 67.24 |
| Variant | Total | Quality | Semantic |
|---|---|---|---|
| video DMD only | 72.44 | 78.68 | 47.50 |
| marginal only | 80.64 | 84.90 | 63.59 |
| random exits | 78.29 | 82.48 | 61.56 |
| RMD (full) | 81.26 | 84.52 | 68.24 |