Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in the surrounding context merely to preserve temporal consistency. To provide a clearer visual-quality signal, we introduce Rollout-Marginal Distillation (RMD). RMD retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context. To compensate for the lack of temporal context in independent chunk scoring, RMD subsequently applies video-level DMD to restore temporal coherence. Extensive experiments demonstrate that RMD maintains high visual quality far beyond its training horizon and outperforms video-level DMD baselines. Code and video results are available at https://cjeen.github.io/RMD
Figures & tables
Figure 1: Long-horizon autoregressive video generation. Trained on 5-second rollouts, RMD maintains high visual quality over 500-second rollouts ( 100× the training horizon). Notably, generation relies solely on a fully sliding context window, without fixed anchors, additional explicit memory mechanisms, or modifications to the generator architecture.
Figure 2: Comparison of real-score predictions under video-level DMD and RMD. (a) The autoregressive rollout accumulates visual artifacts. (b) Video-level DMD scores the sequence jointly, causing the real-score prediction to retain these historical artifacts. (c) RMD scores the current chunk independently, providing a clean target regardless of the degraded context.
Figure 3: Overview of Rollout-Marginal Training. The AR generator produces a causal rollout, sampling the initial chunk ( x1 ) at a random-step exit and continuation chunks ( x2,…,xT ) at the last-step exit. Adapted score networks evaluate them independently without temporal context. The resulting gradients ( g1,…,gT ) are backpropagated via differentiable causal replay.
Figure 4: Qualitative comparison of long-horizon generation. While baseline methods suffer from severe color saturation and structural collapse beyond their training horizon, RMD consistently preserves sharp subject details and stable scene structures.
Chunk size
Method
Total ↑
Quality ↑
Semantic ↑
Self Forcing
70.94
77.89
43.17
1
Causal Forcing
69.88
78.34
36.03
RMD (ours)
81.26
84.52
68.24
Self Forcing
77.44
81.99
59.22
3
Causal Forcing
74.88
80.53
52.27
RMD (ours)
81.48
85.03
67.24
Table 1: Quantitative comparison on long-horizon generation. Evaluated on VBench-Long, RMD consistently outperforms video-level distillation baselines across all metrics.
Figure 5: Performance over increasing rollout lengths. VBench-Long scores are evaluated at durations ranging from 10 to 60 seconds. As generation extends, RMD maintains stable visual quality and significantly mitigates the performance degradation suffered by baseline methods.
Figure 6: Qualitative ablation results. The spatio-temporal slice is extracted along the red line. Removing the rollout-marginal objective ( video DMD only ) causes severe color artifacts. Removing video-level refinement ( marginal only ) introduces temporal flickering (jagged slice artifacts). Removing asymmetric denoising exits ( random exits ) destroys the temporal prior, resulting in severe temporal jitter. Full RMD achieves both high visual quality and temporal coherence.
Variant
Total ↑
Quality ↑
Semantic ↑
video DMD only
72.44
78.68
47.50
marginal only
80.64
84.90
63.59
random exits
78.29
82.48
61.56
RMD (full)
81.26
84.52
68.24
Table 2: Quantitative ablation results. (a) Removing refinement with video DMD ( marginal only ) or asymmetric exits ( random exits ) degrades overall performance. (b) Temporal metrics confirm that this refinement is essential for temporal coherence.
Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to produce highly similar videos with weakened motion dynamics. We analyze this degradation in Distribution Matching Distillation (DMD)-distilled AR video generators and find that, in the autoregressive setting, it takes the form of a structured uncertainty collapse: the mode-seeking bias of DMD maps different noise samples to nearly identical first chunks, and the deterministic AR cache then propagates this collapsed state to all subsequent chunks, turning a local loss of stochasticity at the rollout root into a global suppression of temporal variation. Based on this analysis, we propose Uncertainty DMD, a simple uncertainty-injection framework that restores stochasticity at two key stages of AR generation: a timestep perturbation for the first chunk to increase first-chunk diversity, and a stochastic cache-writing mechanism for later chunks to preserve uncertainty in autoregressive conditioning. The method requires no architectural changes and introduces only lightweight perturbation operations. The same perturbation mechanisms are used during both training and inference. Experiments show that Uncertainty DMD consistently improves diversity and motion dynamics while maintaining comparable per-sample visual quality.
Zixuan Duan, Xunzhi Xiang, Yabo Chen +6
Nanjing University, Institute of Artificial Intelligence, China Telecom (TeleAI), China · Institute of Artificial Intelligence, China Telecom (TeleAI), China · Fudan University, Institute of Artificial Intelligence, China Telecom (TeleAI), China
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.
Chi Zhang, Yueyi Liu, Shi Haoyang +5
College of AI, Tsinghua University · IAIR, Xi’an Jiaotong University · Xianghui Academy, Fudan University +1
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.