cs.CVSep 30, 2026

Uncertainty-Aware Consistency Distillation for Few-Step Video Generation

Authors: Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng

Organizations: School of Software Xi’an Jiaotong University Xi’an, China. · School of Computer and Information Science Hefei University of Technology Hefei, China. · Faculty of Science and Technology, and Institute of Collaborative Innovation University of Macau Macau, China.

Abstract

We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a common recipe, in which a multi-step teacher provides the consistency targets for a few-step student. However, these teacher-guided targets are not equally trustworthy, and the content is harder to learn where it varies rapidly over time, e.g., moving foliage shadows or flowing water. We observe that supervision reliability follows the local difficulty of the content rather than semantic complexity: regions that change little yield consistent endpoint predictions, whereas regions with large temporal variation produce larger discrepancies that coincide with the largest perceptual errors. Motivated by this observation, we propose Uncertainty-Aware Consistency Distillation (UACD), which reweights consistency supervision at each spatiotemporal region using a local, parameter-free uncertainty estimate. Specifically, we construct two independently perturbed teacher-guided consistency paths, whose student endpoint predictions provide a consensus target; the discrepancy between the student's direct prediction and this target is the uncertainty proxy. We then relax the consistency penalty on high-uncertainty regions through an exponential weight, while keeping the full penalty elsewhere, since the student cannot be expected to match targets that are hard to learn. To preserve perceptual quality under aggressive step reduction, we integrate feature-space adversarial training with semantic alignment. With parameter-efficient LoRA adaptation of the 50-step Wan model, our method achieves state-of-the-art 4-step generation on VBench 2.0 (0.556 mean score) and is preferred over competing methods in a user study.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

    May 13, 2026Yuchao Gu, Guian Fang, Yuxin Jiang +4Video Diffusion ModelsFew-Step Distillation

  2. Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation

    May 14, 2026Min Zhao, Hongzhou Zhu, Kaiwen Zheng +6Video GenerationAutoregressive Model

  3. Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation

    Sep 11, 2026Zixuan Duan, Xunzhi Xiang, Yabo Chen +6Autoregressive Video Diffusion ModelsAutoregressive Video Generation