Uncertainty-Aware Consistency Distillation for Few-Step Video Generation
Authors: Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng
Organizations: School of Software Xi’an Jiaotong University Xi’an, China. · School of Computer and Information Science Hefei University of Technology Hefei, China. · Faculty of Science and Technology, and Institute of Collaborative Innovation University of Macau Macau, China.
We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a common recipe, in which a multi-step teacher provides the consistency targets for a few-step student. However, these teacher-guided targets are not equally trustworthy, and the content is harder to learn where it varies rapidly over time, e.g., moving foliage shadows or flowing water. We observe that supervision reliability follows the local difficulty of the content rather than semantic complexity: regions that change little yield consistent endpoint predictions, whereas regions with large temporal variation produce larger discrepancies that coincide with the largest perceptual errors. Motivated by this observation, we propose Uncertainty-Aware Consistency Distillation (UACD), which reweights consistency supervision at each spatiotemporal region using a local, parameter-free uncertainty estimate. Specifically, we construct two independently perturbed teacher-guided consistency paths, whose student endpoint predictions provide a consensus target; the discrepancy between the student's direct prediction and this target is the uncertainty proxy. We then relax the consistency penalty on high-uncertainty regions through an exponential weight, while keeping the full penalty elsewhere, since the student cannot be expected to match targets that are hard to learn. To preserve perceptual quality under aggressive step reduction, we integrate feature-space adversarial training with semantic alignment. With parameter-efficient LoRA adaptation of the 50-step Wan model, our method achieves state-of-the-art 4-step generation on VBench 2.0 (0.556 mean score) and is preferred over competing methods in a user study.
Figures & tables
Figure 1: Uncertainty-aware supervision according to local consistency reliability. (a) In low-variation regions, two independently constructed teacher-guided consistency paths yield similar student endpoint predictions, producing a small student–consensus discrepancy and thus low uncertainty U (dark heatmap). (b) In high-variation regions, where the content is hard to learn ( e . g ., moving foliage shadows or flowing water), the two paths produce more divergent endpoint predictions, resulting in a larger discrepancy between the student’s direct prediction and the consensus target, and thus higher U (bright heatmap). (c) Risk-coverage analysis showing that uncertainty-guided filtering consistently reduces LPIPS (lower is better) as high-uncertainty videos are removed, while random discarding has little effect. Our loss relaxes the consistency penalty where uncertainty is high and keeps the full penalty elsewhere.
Figure 2: Overview of Uncertainty-Aware Consistency Distillation. Two independently perturbed paths are each advanced by one Euler step using the frozen teacher, producing two teacher-guided intermediate states. The student then predicts the endpoint x0(1) and x0(2) from each intermediate state, and their average forms a detached consensus consistency target x^0 . One of the two perturbed states is randomly selected for the direct student prediction input xt , and its discrepancy with the consensus target defines the normalized local uncertainty map U~ , which reweights the consistency distillation loss to attenuate supervision at regions with higher uncertainty. A feature-space discriminator provides adversarial and feature-matching objectives, together with a semantic alignment loss.
Method
NFE
Teacher
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
Wan2.1-1.3B
50
-
0.711
0.551
0.099
0.519
0.382
0.452
CogVideoX-5B
50
-
0.797
0.344
0.199
0.539
0.405
0.457
Distillation Methods
CausalForcing
4+4×(ℓ//4)
Wan2.1-14B
0.855
0.462
0.184
0.565
0.492
0.512
OneForcing
4+1×(ℓ//4)
Wan2.1-14B
0.754
0.631
0.213
0.571
0.460
0.526
DCM
4
Wan2.1-1.3B
0.828
0.599
0.209
0.603
0.394
0.527
Table 1: Quantitative comparison on VBench 2.0 across 5 video quality dimensions. All videos are generated at 832×480 resolution with ℓ=81 frames. NFE denotes the number of function evaluations during inference. Our method achieves the highest mean score with 4 NFEs using a 1.3B teacher. DCM also uses a 1.3B teacher, while CausalForcing and OneForcing use a 14B teacher.
Figure 3: Qualitative comparison with competitive methods. Our method generates videos with superior visual quality, accurate temporal transitions, and full prompt adherence compared to all baselines, highlighting that our uncertainty modeling enables the student to surpass its own teacher. More sample videos are available on our anonymous project website.
Figure 5
Variants
VBench 2.0
CD
DP
UT
DSA
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
(1)
✓
0.998
0.396
0.146
0.509
0.169
0.444
(2)
✓
✓
0.597
0.632
0.108
0.551
0.422
0.462
(3)
✓
ED
✓
0.840
0.584
0.197
0.487
0.456
0.513
(4)
✓
✓
LT
✓
0.812
0.605
0.153
0.516
0.458
0.509
(5)
✓
✓
ED
0.854
0.682
0.142
0.551
0.422
0.530
Table 2: Ablation study on primary components. CD : Consistency Distillation baseline with adversarial training. DP : Dual-path Prediction, where two independently perturbed teacher-guided paths are constructed and the student predicts the endpoint from each path. UT : Uncertainty Type. ED : Uncertainty Estimation via Discrepancy, where the discrepancy between the student’s direct endpoint prediction and the consensus consistency target is used as a parameter-free uncertainty estimate. LT : Uncertainty Estimation via a Learnable Token. DSA : Discriminator Semantic Alignment loss applied during adversarial training. Variant 1 achieves near-perfect Human Fidelity but a substantially lower mean, since Human Fidelity rewards frame-level realism but not semantic or physical correctness. Our full model (Variant 6) achieves the highest mean with balanced performance, confirming that dual-path uncertainty and semantic alignment are complementary.
Figure 6: Qualitative comparison of ablation variants. Our full model generates videos that faithfully align with the text prompt and closely resemble real-world video content in motion dynamics and visual fidelity.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
K
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
1
0.840
0.584
0.197
0.487
0.456
0.513
2
0.861
0.642
0.203
0.556
0.516
0.556
3
0.860
0.645
0.192
0.545
0.527
0.554
4
0.826
0.674
0.179
0.582
0.519
0.556
Appendix
Table 3: Ablation on the number of perturbed paths.
λ
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
0
0.597
0.632
0.108
0.551
0.422
0.462
0.1
0.747
0.637
0.160
0.574
0.502
0.524
0.5
0.794
0.630
0.154
0.571
0.484
0.526
1
0.861
0.642
0.203
0.556
0.516
0.556
5
0.862
0.692
0.195
0.574
0.518
0.568
10
0.824
0.711
0.180
0.568
0.530
0.562
Appendix
Table 4: Ablation on uncertainty weighting strength.
λalign
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
0.1
0.807
0.641
0.164
0.541
0.477
0.526
0.5
0.842
0.622
0.166
0.552
0.495
0.535
1
0.861
0.642
0.203
0.556
0.516
0.556
5
0.838
0.640
0.189
0.528
0.493
0.538
10
0.812
0.611
0.158
0.553
0.488
0.524
Appendix
Table 5: Ablation on semantic alignment strength.
Depth
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
1
0.860
0.627
0.166
0.555
0.490
0.540
5
0.853
0.647
0.185
0.548
0.514
0.549
10
0.844
0.613
0.190
0.591
0.506
0.549
all
0.861
0.642
0.203
0.556
0.516
0.556
Appendix
Table 6: Ablation on semantic alignment depth.
Figure 7: Additional qualitative comparison with competitive methods.
Figure 8: Additional qualitative comparison with competitive methods.
Figure 9: Additional qualitative comparison with competitive methods.
Few-step video generation has been significantly advanced by consistency distillation. However, the performance of consistency-distilled models often degrades as more sampling steps are allocated at test time, limiting their effectiveness for any-step video diffusion. This limitation arises because consistency distillation replaces the original probability-flow ODE trajectory with a consistency-sampling trajectory, weakening the desirable test-time scaling behavior of ODE sampling. To address this limitation, we introduce AnyFlow, the first any-step video diffusion distillation framework based on flow maps. Instead of distilling a model for only a few fixed sampling steps, AnyFlow optimizes the full ODE sampling trajectory. To this end, we shift the distillation target from endpoint consistency mapping (zt→z0) to flow-map transition learning (zt→zr) over arbitrary time intervals. We further propose Flow Map Backward Simulation, which decomposes a full Euler rollout into shortcut flow-map transitions, enabling efficient on-policy distillation that reduces test-time errors (i.e., discretization error in few-step sampling and exposure bias in causal generation). Extensive experiments across both bidirectional and causal architectures, at scales ranging from 1.3B to 14B parameters, demonstrate that AnyFlow achieves performance matches or surpasses consistency-based counterparts in the few-step regime, while scaling with sampling step budgets.
Yuchao Gu, Guian Fang, Yuxin Jiang +4
1NVIDIA · 2Show Lab, National University of Singapore · 3MIT
Real-time interactive video generation requires low-latency, streaming, and controllable rollout. Existing autoregressive (AR) diffusion distillation methods have achieved strong results in the chunk-wise 4-step regime by distilling bidirectional base models into few-step AR students, but they remain limited by coarse response granularity and non-negligible sampling latency. In this paper, we study a more aggressive setting: frame-wise autoregression with only 1--2 sampling steps. In this regime, we identify the initialization of a few-step AR student as the key bottleneck: existing strategies are either target-misaligned, incapable of few-step generation, or too costly to scale. We propose \textbf{Causal Forcing++}, a principled and scalable pipeline that uses \emph{causal consistency distillation} (causal CD) for few-step AR initialization. The core idea is that causal CD learns the same AR-conditional flow map as causal ODE distillation, but obtains supervision from a single online teacher ODE step between adjacent timesteps, avoiding the need to precompute and store full PF-ODE trajectories. This makes the initialization both more efficient and easier to optimize. The resulting pipeline, \ours, surpasses the SOTA 4-step chunk-wise Causal Forcing under the \textit{\textbf{frame-wise 2-step setting}} by 0.1 in VBench Total, 0.3 in VBench Quality, and 0.335 in VisionReward, while reducing first-frame latency by 50% and Stage 2 training cost by \sim$$4\times. We further extend the pipeline to action-conditioned world model generation in the spirit of Genie3. Project Page: https://github.com/thu-ml/Causal-Forcing and https://github.com/shengshu-ai/minWM .
Min Zhao, Hongzhou Zhu, Kaiwen Zheng +6
Tsinghua University · ShengShu · Renmin University of China
Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to produce highly similar videos with weakened motion dynamics. We analyze this degradation in Distribution Matching Distillation (DMD)-distilled AR video generators and find that, in the autoregressive setting, it takes the form of a structured uncertainty collapse: the mode-seeking bias of DMD maps different noise samples to nearly identical first chunks, and the deterministic AR cache then propagates this collapsed state to all subsequent chunks, turning a local loss of stochasticity at the rollout root into a global suppression of temporal variation. Based on this analysis, we propose Uncertainty DMD, a simple uncertainty-injection framework that restores stochasticity at two key stages of AR generation: a timestep perturbation for the first chunk to increase first-chunk diversity, and a stochastic cache-writing mechanism for later chunks to preserve uncertainty in autoregressive conditioning. The method requires no architectural changes and introduces only lightweight perturbation operations. The same perturbation mechanisms are used during both training and inference. Experiments show that Uncertainty DMD consistently improves diversity and motion dynamics while maintaining comparable per-sample visual quality.
Zixuan Duan, Xunzhi Xiang, Yabo Chen +6
Nanjing University, Institute of Artificial Intelligence, China Telecom (TeleAI), China · Institute of Artificial Intelligence, China Telecom (TeleAI), China · Fudan University, Institute of Artificial Intelligence, China Telecom (TeleAI), China