Diffusion distillation is widely adopted to accelerate sampling, and the resulting few-step models are broadly believed to match or even surpass their multi-step teachers in generation. However, standard evaluations such as GenEval2 typically draw only one sample per prompt, so improved scores may fail to reveal losses in distribution coverage. We therefore revisit whether distilled models truly match their teachers beyond single-draw performance using \textbf{pass@k}, which measures the probability that at least one of k independent samples satisfies a quality criterion. At k=1, pass@k reduces to standard single-draw evaluation. As k grows, the curve reveals whether additional draws find genuinely different successes or merely revisit the same modes, directly exposing how broadly a model covers the space of valid outputs. We first show that classifier-free guidance (CFG), whose quality--coverage tradeoff is well established, is the clearest case: higher guidance improves pass@1, but its advantage shrinks and reverses at larger k. Applying pass@k to few-step distilled models, we find the same tradeoff splits along training objectives: distribution-matching objectives concentrate the student's output distribution, boosting early-hit rates while eroding large-budget coverage, whereas consistency and trajectory-based objectives better preserve the teacher's coverage even at large k. We further show that this tradeoff extends to few-step causal video generation. Our findings reveal a previously overlooked cost of diffusion distillation: across both image and video generation, the choice of training objective fundamentally determines whether a few-step model inherits its teacher's distribution coverage or trades it away for single-draw quality.
Figures & tables
Fig. 1: SD3.5-Large CFG sweep experiment on GenEval2. (a) Best-of- k GenEval2 Soft-TIFA GM for CFG ∈{1,3.5,7} . The dotted line marks where CFG =1 overtakes CFG =7 . (b) Distribution of per-prompt success rates, where success is Soft-TIFA GM ≥0.3 over N=256 samples—a lenient cutoff, shared by all GenEval2 per-prompt diagnostics, that credits any mostly-correct sample (Appendix A.2 ).
Fig. 2: Teacher–student performance across sampling budgets. Pass@ k on GenEval and expected best-of- k on GenEval2, DPG-Bench, and GenAI-Bench for five distillation families ( N=256 per prompt). Solid and dashed curves denote the teacher and student, respectively; dotted lines mark the first teacher–student crossover.
Fig. 3: Distilled samples are more concentrated on matched-correct prompts. Joint top-2 PCA projections of CLIP embeddings from SD3.5-Large and SD3.5-Large-Turbo on three matched-correct prompts ( 100 samples per model, per-model 2σ ellipses). Each panel reports the mean pairwise cosine distance in the original CLIP space.
Fig. 4: Distillation objectives shape coverage across training and sampling budgets. (a) GenEval2 GM at k=1 across training checkpoints. (b) Best-of- k curves for the checkpoint from each run with the strongest k=1 performance.
Fig. 5: Coverage degradation is accompanied by diversity collapse. Along the same training trajectories, (a) reports mean pairwise CLIP distance within each prompt and (b) reports GenEval2 GM best-of- 64 . DMD2 contracts in representation space as its coverage declines, whereas MeanFlow preserves diversity while improving coverage.
Fig. 6: VBench-20 best-of- k curves for video forcing models. Each curve uses the exact subset-averaged estimator in Eq. 2 . Solid curves denote distribution-matching methods (Self-Forcing DMD and Causal-Forcing), dash-dot curves denote consistency distillation (Causal-CD), and dashed curves denote multi-step baselines. The four lower rows show Δm(k)=sWan(k)−sm(k) for AR-Diffusion, Causal-CD, Self-Forcing DMD, and Causal-Forcing, respectively. Positive differences favor Wan.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 7: Full CFG sweep. Pass@ k and best-of- k curves for SD3.5-Large (a–d) and FLUX.1-dev (e–f) under CFG ∈{1,3.5,7} . Dotted lines mark where CFG =1 first matches or overtakes CFG =7 .
Fig. 8: Per-prompt success-rate distributions for the CFG sweeps. Panel order follows Fig. 7 .
Fig. 9: Teacher–student crossover budgets by benchmark subscore. Rows are model families and columns are benchmark-specific tasks, atom counts, categories, or splits. Darker cells indicate later crossovers; gray cells marked NC do not cross within N=256 .
Fig. 10: Per-prompt diagnostics for SD3.5-Large and Turbo. Panels show exclusive wins over k , per-prompt success rates, and GenEval2 failed-atom counts.
Fig. 11: Per-prompt diagnostics for Sana and Sana-Sprint. Panels follow Fig. 10 .
Fig. 12: Per-prompt diagnostics for Qwen-Image and Lightning. Panels follow Fig. 10 .
Fig. 13: Per-prompt diagnostics for Z-Image and Z-Image-Turbo. Panels follow Fig. 10 .
Fig. 14: Per-prompt diagnostics for SDXL and LCM. Panels follow Fig. 10 ; LCM retains its advantage across the measured range.
Fig. 15: Best-of- k curves at iteration 4500. The top row shows DMD2 with embedded guidance 3.5 , 1 , and 7 ; the bottom row shows MeanFlow and ADD with guidance 3.5 . The dashed black curve is the shared 50-step FLUX.1-dev teacher.
Fig. 16: Per-prompt success rates in the controlled study. A sample counts as a success when its GenEval2 geometric-mean score is at least 0.3 , the cutoff shared with Figs. 1 and 8 .
standard
require ≥2
benchmark
comparison
k×
Δ256
k×
Δ256
GenAI-Bench
SD3.5-L / Turbo
8
+.063
16
+.059
Sana / Sprint
8
+.063
16
+.063
Qwen-Image / Lightning
4
+.083
8
+.072
Z-Image / Turbo
8
+.108
16
+.101
SDXL / LCM
256
+.004
—
−.009
Appendix
Tab. 1: Teacher–student crossovers and endpoint gaps survive FP-immune re-analysis. For each comparison: the budget at which the multi-step model first overtakes ( k× ) and the k=256 gap Δ256 (multi-step minus few-step), under the standard statistic and the FP-immune pass≥2@k of Eq. 6 . Success thresholds: VQAScore ≥0.8 (GenAI-Bench), Soft-TIFA GM ≥0.3 (GenEval2); GenEval is binary. Requiring two successes shifts each crossover by at most one budget doubling and preserves (GenEval: enlarges) every gap; SDXL/LCM is the coverage counter-example reported in the main text and remains one.
unsolved
FP
measured
evaluator
model
1−p^
inflation
Δ256
Soft-TIFA GM ≥0.3
Z-Image
.011
.001 [ .005 ]
+.156
( 1/3200 null passes)
Turbo
.168
.014 [ .077 ]
VQAScore ≥0.8
SD3.5-L
.082
0 [ .036 ]
+.063
( 0/2108 null passes)
Turbo
.144
0 [ .063 ]
Sana
.089
0 [ .039 ]
+.063
Appendix
Tab. 2: Null-pair calibration. Measured FP rate of each evaluator and the most it can add to each model’s pass@ 256 .
Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world scenarios. However, Distribution Matching Distillation (DMD), a leading paradigm, tends to degrade under limited NFE budgets, manifesting in video generation as layout instability, oversaturation, and broken motion dynamics. We trace this failure to a structural limitation: standard DMD is an intra-sample distribution-matching objective with coordinate-wise gradients, and thus imposes no explicit constraint on the relational geometry across batch elements or temporal frames, leaving the underlying copula largely unregulated. Combined with the mode-seeking tendency of its reverse-KL objective, this absence of relational guidance makes DMD prone to collapsing into local optima in the few-step regime. Motivated by this insight, we propose Copula-aware DMD (CoDMD), a lightweight relational regularizer that reuses score estimates already produced by the frozen teacher and the online fake model to construct pairwise relation matrices across samples and frames. These are matched through a supplementary distributional objective that requires no additional networks, datasets, or sampling trajectories. On the Wan-2.1-T2V model series at 1.3B & 14B scales, CoDMD distills 50-step teachers into 4-step students, achieving an approximate 25× speed-up while attaining VBench scores of 84.46 & 84.87, outperforming prior trajectory-based (rCM 82.81 & 84.05) and distribution-based (DMD 83.38 & 83.81) methods.
Few-step distillation accelerates diffusion models but must balance diversity and fidelity: trajectory-based distillation preserves mode coverage, while distribution matching sharpens samples but can reduce diversity. We show that this tension can be exploited in a noise-regime-dependent way: high-noise steps largely determine global modes, whereas low-noise steps refine local details. We propose CrossDistill, a trajectory-level hybrid distillation framework that splits the sampling trajectory at a crossover point, applies a trajectory-preserving objective on the high-noise interval and a distribution-matching objective on the low-noise interval, and couples the two stages through the crossover state. In contrast to loss-level mixing, and complementarily to training-time two-stage recipes, CrossDistill explicitly assigns complementary objectives along the noise axis, so that global branching is preserved before local statistics are sharpened. CrossDistill is a noise-level scheduling policy: PCM and DMD are plug-in instantiations, while the noise partition, crossover coupling, and objective ordering are the key design elements. Experiments on text-to-video diffusion models and qualitative image-to-video results show that CrossDistill expands the few-step quality-diversity frontier, retaining seed-level variation while achieving competitive visual fidelity.
Yuxi Liu, Haoyu Li, Yixiang Cai +8
Peking University, Melon Group · Alibaba Group · Tsinghua University
Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its successor DMD2 achieved strong generation quality and fast convergence. However, due to the nature of the reverse Kullback--Leibler (KL) objective, these methods exhibit two persistent failure modes: a substantial drop in sample diversity, and visibly over-saturated outputs that deviate from real-video appearance. In this work, we propose Data-Forcing Distillation (DFD), a simple post-training framework that restores diversity and fidelity in DMD with only a single-line of code change. At its core is the teacher score discrepancy to guide the student toward the real-data distribution, pulling it to missing modes (mitigating mode collapse) and away from problematic modes absent in real data (avoiding over-saturation). We provide an in-depth theoretical analysis of our framework and validate our approach on text-to-video, image-to-video, and autoregressive video generation. With only 100--300 steps of finetuning, DFD effectively restores diversity and fidelity on both Wan2.1-1.3B and Cosmos-Predict2.5-2B model, resolving the over-saturation artifacts with significantly better video dynamics and appearance, and even outperforms the teacher model.
Siyi Chen, Shaowei Liu, Yixuan Jia +4
University of Michigan · NVIDIA · University of Illinois Urbana-Champaign