Uncertainty-Aware Consistency Distillation for Few-Step Video Generation
Authors: Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng
Organizations: School of Software Xi’an Jiaotong University Xi’an, China. · School of Computer and Information Science Hefei University of Technology Hefei, China. · Faculty of Science and Technology, and Institute of Collaborative Innovation University of Macau Macau, China.
We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a common recipe, in which a multi-step teacher provides the consistency targets for a few-step student. However, these teacher-guided targets are not equally trustworthy, and the content is harder to learn where it varies rapidly over time, e.g., moving foliage shadows or flowing water. We observe that supervision reliability follows the local difficulty of the content rather than semantic complexity: regions that change little yield consistent endpoint predictions, whereas regions with large temporal variation produce larger discrepancies that coincide with the largest perceptual errors. Motivated by this observation, we propose Uncertainty-Aware Consistency Distillation (UACD), which reweights consistency supervision at each spatiotemporal region using a local, parameter-free uncertainty estimate. Specifically, we construct two independently perturbed teacher-guided consistency paths, whose student endpoint predictions provide a consensus target; the discrepancy between the student's direct prediction and this target is the uncertainty proxy. We then relax the consistency penalty on high-uncertainty regions through an exponential weight, while keeping the full penalty elsewhere, since the student cannot be expected to match targets that are hard to learn. To preserve perceptual quality under aggressive step reduction, we integrate feature-space adversarial training with semantic alignment. With parameter-efficient LoRA adaptation of the 50-step Wan model, our method achieves state-of-the-art 4-step generation on VBench 2.0 (0.556 mean score) and is preferred over competing methods in a user study.
Figures & tables
Figure 1: Uncertainty-aware supervision according to local consistency reliability. (a) In low-variation regions, two independently constructed teacher-guided consistency paths yield similar student endpoint predictions, producing a small student–consensus discrepancy and thus low uncertainty U (dark heatmap). (b) In high-variation regions, where the content is hard to learn ( e . g ., moving foliage shadows or flowing water), the two paths produce more divergent endpoint predictions, resulting in a larger discrepancy between the student’s direct prediction and the consensus target, and thus higher U (bright heatmap). (c) Risk-coverage analysis showing that uncertainty-guided filtering consistently reduces LPIPS (lower is better) as high-uncertainty videos are removed, while random discarding has little effect. Our loss relaxes the consistency penalty where uncertainty is high and keeps the full penalty elsewhere.
Figure 2: Overview of Uncertainty-Aware Consistency Distillation. Two independently perturbed paths are each advanced by one Euler step using the frozen teacher, producing two teacher-guided intermediate states. The student then predicts the endpoint x0(1) and x0(2) from each intermediate state, and their average forms a detached consensus consistency target x^0 . One of the two perturbed states is randomly selected for the direct student prediction input xt , and its discrepancy with the consensus target defines the normalized local uncertainty map U~ , which reweights the consistency distillation loss to attenuate supervision at regions with higher uncertainty. A feature-space discriminator provides adversarial and feature-matching objectives, together with a semantic alignment loss.
Method
NFE
Teacher
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
Wan2.1-1.3B
50
-
0.711
0.551
0.099
0.519
0.382
0.452
CogVideoX-5B
50
-
0.797
0.344
0.199
0.539
0.405
0.457
Distillation Methods
CausalForcing
4+4×(ℓ//4)
Wan2.1-14B
0.855
0.462
0.184
0.565
0.492
0.512
OneForcing
4+1×(ℓ//4)
Wan2.1-14B
0.754
0.631
0.213
0.571
0.460
0.526
DCM
4
Wan2.1-1.3B
0.828
0.599
0.209
0.603
0.394
0.527
Table 1: Quantitative comparison on VBench 2.0 across 5 video quality dimensions. All videos are generated at 832×480 resolution with ℓ=81 frames. NFE denotes the number of function evaluations during inference. Our method achieves the highest mean score with 4 NFEs using a 1.3B teacher. DCM also uses a 1.3B teacher, while CausalForcing and OneForcing use a 14B teacher.
Figure 3: Qualitative comparison with competitive methods. Our method generates videos with superior visual quality, accurate temporal transitions, and full prompt adherence compared to all baselines, highlighting that our uncertainty modeling enables the student to surpass its own teacher. More sample videos are available on our anonymous project website.
Figure 5
Variants
VBench 2.0
CD
DP
UT
DSA
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
(1)
✓
0.998
0.396
0.146
0.509
0.169
0.444
(2)
✓
✓
0.597
0.632
0.108
0.551
0.422
0.462
(3)
✓
ED
✓
0.840
0.584
0.197
0.487
0.456
0.513
(4)
✓
✓
LT
✓
0.812
0.605
0.153
0.516
0.458
0.509
(5)
✓
✓
ED
0.854
0.682
0.142
0.551
0.422
0.530
Table 2: Ablation study on primary components. CD : Consistency Distillation baseline with adversarial training. DP : Dual-path Prediction, where two independently perturbed teacher-guided paths are constructed and the student predicts the endpoint from each path. UT : Uncertainty Type. ED : Uncertainty Estimation via Discrepancy, where the discrepancy between the student’s direct endpoint prediction and the consensus consistency target is used as a parameter-free uncertainty estimate. LT : Uncertainty Estimation via a Learnable Token. DSA : Discriminator Semantic Alignment loss applied during adversarial training. Variant 1 achieves near-perfect Human Fidelity but a substantially lower mean, since Human Fidelity rewards frame-level realism but not semantic or physical correctness. Our full model (Variant 6) achieves the highest mean with balanced performance, confirming that dual-path uncertainty and semantic alignment are complementary.
Figure 6: Qualitative comparison of ablation variants. Our full model generates videos that faithfully align with the text prompt and closely resemble real-world video content in motion dynamics and visual fidelity.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
K
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
1
0.840
0.584
0.197
0.487
0.456
0.513
2
0.861
0.642
0.203
0.556
0.516
0.556
3
0.860
0.645
0.192
0.545
0.527
0.554
4
0.826
0.674
0.179
0.582
0.519
0.556
Appendix
Table 3: Ablation on the number of perturbed paths.
λ
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
0
0.597
0.632
0.108
0.551
0.422
0.462
0.1
0.747
0.637
0.160
0.574
0.502
0.524
0.5
0.794
0.630
0.154
0.571
0.484
0.526
1
0.861
0.642
0.203
0.556
0.516
0.556
5
0.862
0.692
0.195
0.574
0.518
0.568
10
0.824
0.711
0.180
0.568
0.530
0.562
Appendix
Table 4: Ablation on uncertainty weighting strength.
λalign
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
0.1
0.807
0.641
0.164
0.541
0.477
0.526
0.5
0.842
0.622
0.166
0.552
0.495
0.535
1
0.861
0.642
0.203
0.556
0.516
0.556
5
0.838
0.640
0.189
0.528
0.493
0.538
10
0.812
0.611
0.158
0.553
0.488
0.524
Appendix
Table 5: Ablation on semantic alignment strength.
Depth
Human Fidelity
Creativity
Controllability
Commonsense
Physics
Mean
1
0.860
0.627
0.166
0.555
0.490
0.540
5
0.853
0.647
0.185
0.548
0.514
0.549
10
0.844
0.613
0.190
0.591
0.506
0.549
all
0.861
0.642
0.203
0.556
0.516
0.556
Appendix
Table 6: Ablation on semantic alignment depth.
Figure 7: Additional qualitative comparison with competitive methods.
Figure 8: Additional qualitative comparison with competitive methods.
Figure 9: Additional qualitative comparison with competitive methods.
Nanjing University, Institute of Artificial Intelligence, China Telecom (TeleAI), China · Institute of Artificial Intelligence, China Telecom (TeleAI), China · Fudan University, Institute of Artificial Intelligence, China Telecom (TeleAI), China