Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.
Figures & tables
Figure 2: Qualitative comparisons over matched 4-NFE training. DMD and DMD2 develop unnatural textures after 500 iterations, whereas PDMD keeps a natural appearance.
Figure 3: DMD and PDMD updates. DMD moves the student along −d . The critic score scritic , the online-learning part of d , carries hidden error. PDMD projects out the component of d parallel to the known residual r to suppress critic error, because r is the critic error estimate. The two frames decode one real update target: the PDMD target contains fewer artifacts than the DMD target , while keeping 83% of the update norm.
Figure 3
Figure 6: Qualitative results on four-step Wan2.1-T2V-1.3B. Each example shows early and late frames from the same VBench prompt and seed across methods, following the AnyFlow evaluation protocol ( Gu et al., 2026 ) . In these examples, PDMD produces natural textures and less oversaturation than the matched DMD and DMD2 models. See Supp. Figure E5 for more visual results.
Video
Audio
User study (win − loss, %)
Method
NFE
Total
Quality
Dynamic
Semantic
PQ
CE
CU
IS
IB
DeSync ↓
Overall
Text
Visual
Motion
Audio
MiniMax-H3-33B ( MiniMax, 2026a )
50
82.41
82.22
66.67
83.20
6.567
4.188
6.213
5.15
0.229
0.797
−37.0
−7.5
−35.4
−11.9
−3.9
MiniMax-H3-33B ( MiniMax, 2026a )
4
79.48
79.54
44.44
79.23
6.056
3.167
5.580
3.35
0.129
0.932
+70.3
+16.5
+72.6
+44.7
+47.5
H3 Turbo LoRA ( LarryVrh, 2026 )
4
81.57
81.29
54.78
82.69
6.406
3.917
6.015
4.52
0.183
0.839
+34.6
+2.8
+38.0
+19.4
+9.3
rCM † ( Zheng et al., 2026 )
4
81.18
80.98
58.40
81.95
6.063
3.711
5.446
3.69
0.194
0.916
+43.7
+3.1
+51.2
+31.5
+17.1
AnyFlow † ( Gu et al., 2026 )
4
81.97
81.82
64.60
82.60
6.092
3.544
5.591
4.35
0.161
0.905
+59.2
−1.6
+66.4
+27.9
+25.8
Table 2: Quantitative 4-NFE MiniMax-H3-33B comparisons on VideoGen-Eval. † marks rows reporting the best reimplemented checkpoint. PDMD reaches the highest visual total and performs best across all audio metrics among the distilled models, and the user study prefers it to 4-NFE baselines on visual, motion, and audio quality. Text-alignment judgments are mostly ties (Supp. Table C4 ).
Figure 7: Qualitative results for 4-NFE MiniMax-H3 on VideoGen-Eval. PDMD shows fewer artifacts, more realistic textures, and better motion. Each prompt shows an earlier and a later frame of the same clip; Supp. Section E.4 shows more prompts.
Figure 8: Motion on 4-NFE MiniMax-H3 (VideoGen-Eval prompt 760). The compared baselines show translucent ghosting artifacts and severely blur the animals, while PDMD keeps them sharp.
Figure 9: Qualitative ablation of the projection direction on VideoGen-Eval (prompt 882). Results at 1,500 training iterations. In this example, projecting out the update component parallel to the student–critic endpoint residual yields less saturation than the other variants.
Figure 10: Quantitative ablation study: training dynamics of MiniMax-H3 on VideoGen-Eval. PDMD holds its total score; the others decline.
Figure 11: Single-step limitation.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
σ∈[0.15,1.1]
all levels
Deleted direction
γe
γs
γe−γs
γe bound
γe
γs
γe−γs
γe bound
Random direction
0.50
0.50
0.00
–
0.50
0.50
0.00
–
Critic score
0.52
0.49
0.03
–
0.52
0.49
0.03
–
Student–teacher residual
0.56
0.61
−0.05
–
0.55
0.59
−0.04
–
Student–critic residual ( PDMD )
0.68
0.57
0.11
0.25
0.61
0.56
0.04
0.21
Appendix
Table B1: Removal ratios of the four deleted directions on the ring of eight. Mean γe (critic error removed) and γs (ideal signal removed) over the same probes, noise draws, and snapshots as Figure B1 , along the PDMD run. Left: the noise band σ∈[0.15,1.1] in which the critic carries error. Right: all 12 noise levels. In the plane a uniformly random rank-one deletion removes 1/2 of any fixed vector’s energy on average, so 1/2 is the reference for both ratios. The last column of each block is the mean lower bound ∥e∥2/(∥e∥2+trΣ(q)) of Equation 10 on the same probes; it is a bound on γe for the student–critic endpoint residual and is listed in that row.
σ∈[0.15,1.1]
all levels
Deleted direction
ν↓
φ↑
ν↓
φ↑
Random direction
0.84
0.68
0.95
0.63
Critic score
0.81
0.71
0.88
0.64
Student–teacher residual
0.92
0.67
1.08
0.62
Student–critic residual ( PDMD )
0.72
0.74
0.89
0.66
Appendix
Table B2: Net effect of the four deletions on the ring of eight. Same probes, noise draws and snapshots as Table B1 , along the PDMD run. ν ( ↓ ) is the error ratio E∥d−d⋆∥2/E∥d−d⋆∥2 : below 1 the deletion moves the update closer to the ideal update, above 1 farther away. φ ( ↑ ) is the improved fraction, the share of probes on which the projected update is closer to d⋆ than the DMD update. Left: the noise band σ∈[0.15,1.1] in which the critic carries error. Right: all 12 noise levels. Best value per column in bold.
Figure B1: Direct measurement of the removal ratios on the ring of eight. Along the PDMD run of Section B.1 : the fraction of critic-error energy γe and of ideal-signal energy γs that the projection removes, against the noise level. The grey dashed line is the theoretical lower bound ∥e∥2/(∥e∥2+trΣ(q)) of Equation 10 on γe . Each curve is a mean over the snapshots from iteration 100 to 4,000. The dotted line marks the mean energy fraction 1/2 removed by a uniformly random rank-one deletion from a fixed vector in the plane.
Figure B2: Generated samples of the five 2D variants over training. Columns are the five variants of Section B.1 . Rows are iterations 0,1,000,…,4,000 of the same runs as Figure 5 . Point color encodes the angle of the latent z , visualizing how latent directions map to target modes. Each panel lists the energy distance ( ↓ , lower is better), the number of covered modes ( ↑ ), and the on-mode fraction ( ↑ , higher is better) at the shown iteration. The student–critic-endpoint-residual column ( PDMD ) covers all eight modes from iteration 2,000 onward and stays tight. The critic-score column covers the modes but does not concentrate. The teacher-residual column shows the instability near iteration 3,000 .
Student lr
Variant
Collapsed ↓
Mean ∣a∣↓
On-mode fraction
Mean energy ↓
2×10−3
DMD
14/20
0.355±0.219
0.970±0.045
1.274±0.856
Random direction ⊥
16/20
0.396±0.197
0.975±0.065
1.449±0.756
PDMD (ours)
3/20
0.138±0.162
0.973±0.039
0.386±0.689
1×10−3
DMD
19/20
0.477±0.104
0.913±0.231
1.899±0.618
Random direction ⊥
19/20
0.476±0.108
0.991±0.018
1.791±0.440
PDMD (ours)
12/20
0.331±0.215
0.925±0.224
1.317±1.046
Appendix
Table B3: Two-mode collapse counts over 20 consecutive seeds at two step sizes. Collapse is modes covered<2 at the end of training; ∣a∣ is the final imbalance; ∣a∣ , on-mode fraction, and energy distance are means ± standard deviations over seeds. The critic learning rate is 20× smaller than the student’s in both blocks. The random direction is control (b) from Section B.1 : the same rank-one deletion along an unstructured direction.
Figure B3: Two-mode collapse process under a lagging critic (seed 0). Rows are plain DMD and PDMD ; columns are training iterations. Each panel lists the iteration and the imbalance ∣a∣ ; ∣a∣=0 is the balanced solution and ∣a∣=0.5 is full collapse. Both variants start from a collapsed transient at iteration 400. Plain DMD stays locked on one mode for all 6,000 iterations. PDMD transports mass back between the modes (iteration 2,400 shows samples in transit) and settles near balance. The main text shows the collapsed DMD endpoint in Figure 5 .
Wan2.1-T2V-1.3B
MiniMax-H3-33B
Trainable parameters
all
LoRA, rank and scaling 128
Training data
42K captions, no real video
248K prompts, no real video
DMD2 discriminator positives
H3 teacher rollouts
H3 teacher rollouts
Clip
81 frames, 480×832
5 s, 544p, mixed aspect ratios
GPUs / global batch
64 H100 / 64
16 H100 / 16
Optimizer
AdamW (0,0.999) , wd 0.01
AdamW (0,0.9) , no wd
Appendix
Table C1: Training configurations of the video experiments. Our DMD, DMD2, and PDMD runs in Table 1 share the Wan2.1 column and the DMD † , DMD2 † and PDMD rows of Table 2 share the MiniMax-H3 column, except where a baseline paragraph says otherwise.
Stage 1
Stage 2 (the reported model)
rCM † : dCM, then dCM + DMD
Learning rate
5×10−6
5×10−5 (critic 2×10−5 )
Iterations
3,000
6,000, saved every 500
Initialization
LoRA from a full-parameter run
stage 1
Student:critic
no critic
1:4
Loss
LdCM
100LdCM+LDMD
Appendix
Table C2: Two-stage recipes of the rCM † and AnyFlow † ports on MiniMax-H3. Both follow their reference: a trajectory stage first, then a stage that adds DMD on top of it. Every other setting follows Table C1 . Iterations count every update, as in Section C.2 , so the student update count of a stage with a critic is a fraction of it. Stage 1 is listed at the checkpoint that initializes stage 2; training it further did not reach the reported model.
Text alignment
Visual quality
Motion quality
Baseline
Win
Tie
Loss
Net
Win
Tie
Loss
Net
Win
Tie
Loss
Net
Wan2.1-T2V-1.3B, 50×2 NFE
8.5
82.5
8.9
−0.4
36.7
44.3
19.1
+17.6
39.9
44.5
15.6
+24.3
DMD †
9.3
87.8
2.8
+6.5
58.1
37.4
4.5
+53.6
37.7
59.8
2.5
+35.2
DMD2 †
10.1
87.2
2.7
+7.5
49.9
39.7
10.4
+39.5
35.9
59.0
5.1
+30.9
ADV ( You et al., 2026 )
10.7
81.5
7.8
+2.9
46.6
41.4
12.0
+34.6
33.3
52.7
14.1
+19.2
rCM ( Zheng et al., 2026 )
5.7
86.6
7.8
−2.1
46.7
40.4
12.9
+33.8
29.6
58.4
11.9
+17.7
Appendix
Table C3: Vote distribution behind the Wan2.1 user study of Table 1 . Each row is PDMD against one baseline, on the three questions of that study. Win, tie and loss are shares of the judgments, in percent, and net is the win share minus the loss share. Text alignment draws ties on roughly 82 – 88% of judgments throughout; the visual and motion questions separate them. The grey row is the 50-step teacher.
Overall quality
Text alignment
Visual quality
Motion quality
Audio quality
Baseline
Win
Tie
Loss
Net
Win
Tie
Loss
Net
Win
Tie
Loss
Net
Win
Tie
Loss
Net
Win
Tie
Loss
Net
MiniMax-H3-33B, 50 NFE
14.2
34.6
51.2
−37.0
10.6
71.3
18.1
−7.5
8.0
48.6
43.4
−35.4
8.0
72.1
19.9
−11.9
7.5
81.1
11.4
−3.9
MiniMax-H3-33B, 4 NFE
79.3
11.6
9.0
+70.3
26.4
63.8
9.8
+16.6
78.3
16.0
5.7
+72.6
50.1
44.4
5.4
+44.7
49.6
48.3
2.1
+47.5
H3 Turbo LoRA
51.2
32.3
16.5
+34.7
13.4
76.0
10.6
+2.8
49.9
38.2
11.9
+38.0
26.9
65.6
7.5
+19.4
13.7
81.9
4.4
+9.3
DMD †
45.2
39.3
15.5
+29.7
15.2
73.1
11.6
+3.6
34.9
55.3
9.8
+25.1
19.9
72.4
7.8
+12.1
15.5
79.8
4.7
+10.8
DMD2 †
51.7
31.8
16.5
+35.2
19.1
70.8
10.1
+9.0
39.0
50.6
10.3
+28.7
27.6
64.1
8.3
+19.3
26.6
69.0
4.4
+22.2
Appendix
Table C4: Vote distribution behind the MiniMax-H3 user study of Table 2 . Each row is PDMD against one baseline, on the five questions of that study. Win, tie and loss are shares of the judgments, in percent, and net is the win share minus the loss share, computed from the displayed percentages. Text alignment draws ties on roughly 64 – 80% of judgments throughout. Grey rows are the undistilled base model at 50 and 4 sampling steps.
Figure D1: Effect of projection direction during 4-NFE MiniMax-H3 distillation. Six variants that differ only in which direction is removed from the DMD update d , one per column, in the order of Figure 9 . Each row is one of three SoRA prompts, given by its prompt id, and every row is a tight vertical triple: the same run at 500, 1,500 and 3,000 iterations, top to bottom, so reading downwards follows one training trajectory. Every frame is taken at 30% of the generated 1344×768 clip; prompt and seed are the same across variants. In these examples, all five control variants degrade in different ways. Retaining the student-critic endpoint-parallel component instead of removing it is the fastest failure: by 3,000 iterations it produces the same prompt-independent pebble texture for every prompt. PDMD continues to follow its prompt at 3,000 iterations. The five control variants show severe degradation on all three illustrated prompts.
Figure D2: Retained fraction of the DMD update norm on MiniMax-H3 at 4 NFE. dkept is the component the variant keeps, as a 100-iteration running mean. Left: video latents. Right: audio latents.
Figure D3: A larger student step and a partial projection on MiniMax-H3 at 4 NFE. Quality, semantic, and total scores of every 500-iteration checkpoint up to 5,000 iterations on the 387 VideoGen-Eval prompts. DMD is the corresponding variant in Figure 10 . The light green curve shows PDMD at the base learning rate, as far as it was scored, and the dark green curve shows PDMD trained with a 1.25× larger student learning rate ( Section D.4 ). The blue curve shows the run that restores half of the residual-parallel component removed by PDMD , using Equation D1 with β=0.5 ( Section D.3 ). Iteration 0 is the undistilled base model at 4 NFE. Horizontal lines show the teacher scores at 50 and 4 sampling steps.
Figure D4: PDMD with a 1.25× larger student step also does not degrade over 5,000 iterations. The first four columns are checkpoints of the single PDMD run of Section D.4 . The last column, set slightly apart, is the best DMD checkpoint over a learning-rate search, the DMD † row of Table 2 . Rows are labelled with the VideoGen-Eval prompt id; every panel is the frame at 80% of the clip, with the same prompt, seed and sampler throughout.
Video
Audio quality
Audio–video
Method
NFE
Student : critic
Total
Quality
Semantic
PQ
CE
CU
IS
IB
DeSync ↓
DMD †
4
1 : 1
82.37
82.24
82.89
6.321
3.674
5.900
4.32
0.148
0.891
DMD †
4
1 : 5
82.76
82.68
83.05
6.381
3.801
5.931
4.79
0.176
0.905
PDMD
4
1 : 1
82.90
83.14
81.92
6.538
4.123
6.209
5.08
0.191
0.808
PDMD ( Table 2 )
4
1 : 5
83.17
83.25
82.86
6.530
4.062
6.180
4.98
0.195
0.802
Appendix
Table D1: PDMD with one critic update per student update on MiniMax-H3. VideoGen-Eval scores on the 387 prompts (percentages) and audio metrics of the same clips, as in Table D2 . Student : critic is the ratio of student to critic updates; the best value per column is bold.
Video
Audio quality
Audio–video
Method
NFE
Total
Quality
Semantic
PQ
CE
CU
IS
IB
DeSync ↓
DMD †
2
81.58
81.87
80.39
4.529
2.380
3.781
1.88
0.085
1.022
PDMD (ours)
2
82.88
82.83
83.05
6.061
3.719
5.597
4.16
0.177
0.838
Appendix
Table D2: Best checkpoint of DMD and PDMD at 2 NFE on MiniMax-H3. VideoGen-Eval scores on the 387 prompts (percentages) and audio metrics of the same clips. PQ, CE, CU and IS are the audio-quality metrics of Table 2 ; IB and DeSync are the audio–video agreement metrics of Table 2 . Arrows give the preferred direction; the best value per column is bold.
Figure D5: Motion comparisons of DMD and PDMD at 2 NFE and 4 NFE. DMD † at 4 and 2 NFE and PDMD at 4 and 2 NFE, using the checkpoints reported in Tables 2 and D2 . VideoGen-Eval prompt 857 on MiniMax-H3, one row per model; twelve uniformly spaced frames of the 124-frame clip, left to right, labelled by frame index, with the same prompt and seed across rows. Against PDMD at 4 NFE, PDMD at 2 NFE shows slightly lower motion quality and less realistic object texture, but remains clearly better than the unprojected DMD † at 2 NFE.
Figure D6: Additional two-step generation results (part 1 of 4). DMD † at 2 NFE, PDMD at 2 NFE, and PDMD at 4 NFE on MiniMax-H3 for VideoGen-Eval prompts 736 and 735. Each row shows three uniformly sampled frames (0, 62, 123; zero-based) from a 124-frame clip, with the original aspect ratio preserved. † denotes our reimplementation.
Figure D7: Additional two-step generation results (part 2 of 4). DMD † at 2 NFE, PDMD at 2 NFE, and PDMD at 4 NFE on MiniMax-H3 for VideoGen-Eval prompts 733 and 743. Each row shows three uniformly sampled frames (0, 62, 123; zero-based) from a 124-frame clip, with the original aspect ratio preserved. † denotes our reimplementation.
Figure D8: Additional two-step generation results (part 3 of 4). DMD † at 2 NFE, PDMD at 2 NFE, and PDMD at 4 NFE on MiniMax-H3 for VideoGen-Eval prompts 744 and 753. Each row shows three uniformly sampled frames (0, 62, 123; zero-based) from a 124-frame clip, with the original aspect ratio preserved. † denotes our reimplementation.
Figure D9: Additional two-step generation results (part 4 of 4). DMD † at 2 NFE, PDMD at 2 NFE, and PDMD at 4 NFE on MiniMax-H3 for VideoGen-Eval prompts 760 and 769. Each row shows three uniformly sampled frames (0, 62, 123; zero-based) from a 124-frame clip, with the original aspect ratio preserved. † denotes our reimplementation.
Figure E1: One student step, before and after the projection. Results are from the training iteration-500 of the 4-NFE MiniMax-H3 PDMD run in Section 6 . Each row is one of the data-parallel ranks of the same student step, decoded three ways: the student endpoint xs , the DMD target xs−d/Z built from the raw numerator, and the PDMD target xs−d⊥/Z built from the projected one. Under the 125k-token packing each rank holds one sample; the rows are twelve of the sixteen ranks of that step, in rank order. Left: the frame at 30% of the clip. Right: the frame at 80%. The DMD target carries high-frequency speckle, but the PDMD target retains the same subject, colors, and layout while exhibiting less speckle. This comparison associates the removed component with visible artifacts.
Method
Saturation ( C∗/L∗ )
Chroma C∗
Wan2.1-T2V-1.3B
Wan2.1-T2V-1.3B, 50 steps
0.587
21.5
Wan2.1-T2V-1.3B, 4 steps
1.233
27.6
rCM
0.633
23.6
AnyFlow
0.607
25.9
DMD †
0.709
23.5
Appendix
Table E1: Color statistics of the models in Tables 1 and 2 . Wan2.1 values are means over the 4,720 clips of the VBench protocol in Section C.1 ; MiniMax-H3 values are means over the 387 VideoGen-Eval clips of Section C.2 .
Figure E2: Motion on four-step Wan2.1-T2V-1.3B, VBench prompt 119. Each column is one method from Table 1 ; the eighteen rows are uniformly spaced frames of the 81-frame clip, top to bottom, labelled by frame index, with the same prompt and seed across methods. Besides the visible color cast of the other methods, every baseline except ADV misreads the rider’s motion and turns his upper body by 180 degrees.
Figure E3: Motion on four-step MiniMax-H3, VideoGen-Eval prompt 857. Each row is one method from Table 2 ; the twelve frames are uniformly spaced over the 124-frame clip, left to right, labelled by frame index, with the same prompt and seed across methods. Turbo LoRA, rCM † and AnyFlow † blur the slowly moving bouquet into translucent ghosting. DMD † and DMD2 † exhibit shifted colors and smeared textures.
Figure E4: Motion on four-step MiniMax-H3, VideoGen-Eval prompt 895. Each row is one method from Table 2 ; the ten frames are uniformly spaced over the 124-frame clip, left to right, labelled by frame index, with the same prompt and seed across methods. Except for AnyFlow † , the distilled baselines render an almost completely static video. PDMD shows the baseball bat entering the frame, similar to the teacher.
Figure E5: Additional qualitative results for four-step Wan2.1-T2V-1.3B. Each block shows an early and a later frame from one video under the same layout, methods, protocol, and seed as Figure 6 ; † marks our implementations ( Section C.1 ).
Figure E6: Additional qualitative results for four-step Wan2.1-T2V-1.3B (continued). The remaining prompts use the same methods, frame positions, protocol, and seed as the preceding page.
Figure E7: Qualitative results on four-step MiniMax-H3 (part 1 of 2). The eight models of Table 2 on the first four of the 8 prompts, two frames of every clip per prompt with the earlier frame above. Every clip is 124 frames, and rows keep the benchmark’s own aspect ratio, so their heights differ. † marks our reimplementations.
Figure E8: Qualitative results on four-step MiniMax-H3 (part 2 of 2). Four more prompts, same models, same layout and the same two-frame rule as Figure E7 .
Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world scenarios. However, Distribution Matching Distillation (DMD), a leading paradigm, tends to degrade under limited NFE budgets, manifesting in video generation as layout instability, oversaturation, and broken motion dynamics. We trace this failure to a structural limitation: standard DMD is an intra-sample distribution-matching objective with coordinate-wise gradients, and thus imposes no explicit constraint on the relational geometry across batch elements or temporal frames, leaving the underlying copula largely unregulated. Combined with the mode-seeking tendency of its reverse-KL objective, this absence of relational guidance makes DMD prone to collapsing into local optima in the few-step regime. Motivated by this insight, we propose Copula-aware DMD (CoDMD), a lightweight relational regularizer that reuses score estimates already produced by the frozen teacher and the online fake model to construct pairwise relation matrices across samples and frames. These are matched through a supplementary distributional objective that requires no additional networks, datasets, or sampling trajectories. On the Wan-2.1-T2V model series at 1.3B & 14B scales, CoDMD distills 50-step teachers into 4-step students, achieving an approximate 25× speed-up while attaining VBench scores of 84.46 & 84.87, outperforming prior trajectory-based (rCM 82.81 & 84.05) and distribution-based (DMD 83.38 & 83.81) methods.
Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style video distillation faces two coupled challenges: the fake score must track a continuously evolving generator, making training costly when frequent updates are required, while reverse-KL-style matching can be mode-seeking and conservative for preserving strong motion dynamics. To address these issues, we propose \textbf{Score Gradient Matching Distillation (SGMD)}. SGMD adopts a fake-score perspective by directly optimizing the fake score toward the teacher, while using teacher stop-gradient Fisher as a stable distribution-matching objective. We provide a gradient analysis that motivates this objective choice under ideal tracking. Building on this, SGMD introduces a pair of dual potentials: negative-residual (NR) for outer-loop correction and residual-contraction (RC) for inner-loop tracking. Empirically, compared to DMD2, SGMD achieves an approximately ∼3× training speedup and substantially improves motion dynamics for 4-step distilled models while preserving temporal consistency. A human study confirms that SGMD is preferred in motion quality and overall preference, while visual quality and text alignment remain comparable. Code is available at https://github.com/ModelTC/LightX2V.
Zhuguanyu Wu, Ruihao Gong, Yang Yong +5
Beihang University · SenseTime Research · Hong Kong University of Science and Technology
Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its successor DMD2 achieved strong generation quality and fast convergence. However, due to the nature of the reverse Kullback--Leibler (KL) objective, these methods exhibit two persistent failure modes: a substantial drop in sample diversity, and visibly over-saturated outputs that deviate from real-video appearance. In this work, we propose Data-Forcing Distillation (DFD), a simple post-training framework that restores diversity and fidelity in DMD with only a single-line of code change. At its core is the teacher score discrepancy to guide the student toward the real-data distribution, pulling it to missing modes (mitigating mode collapse) and away from problematic modes absent in real data (avoiding over-saturation). We provide an in-depth theoretical analysis of our framework and validate our approach on text-to-video, image-to-video, and autoregressive video generation. With only 100--300 steps of finetuning, DFD effectively restores diversity and fidelity on both Wan2.1-1.3B and Cosmos-Predict2.5-2B model, resolving the over-saturation artifacts with significantly better video dynamics and appearance, and even outperforms the teacher model.
Siyi Chen, Shaowei Liu, Yixuan Jia +4
University of Michigan · NVIDIA · University of Illinois Urbana-Champaign