cs.CVSep 28, 2026

From Scores to Samples: Elastic Forcing for Autoregressive Video Generation

Authors: Chi Zhang, Yueyi Liu, Shi Haoyang, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, Miao Liu

Organizations: College of AI, Tsinghua University · IAIR, Xi’an Jiaotong University · Xianghui Academy, Fudan University · Peking University

Abstract

Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation

    Apr 3, 2026Xingtong Ge, Yi Zhang, Yushi Huang +6Autoregressive Video GenerationVideo Generation

  2. One-Forcing: Towards Stable One-Step Autoregressive Video Generation

    May 22, 2026Jiaqi Feng, Justin Cui, Yuanhao Ban +1Autoregressive Video Generation