Distribution Matching Distillation (DMD) enables high-quality diffusion sampling in only a few steps, but its optimization dynamics remain dominated by coarse, low-frequency signals, delaying the recovery of fine-grained details. We identify a pronounced concentration of spectral amplitudes at low frequencies in the DMD directional error, where dominant low-frequency components overwhelm weaker mid- and high-frequency signals. To address this issue, we propose Spectral Amplitude Purification for Distribution Matching Distillation (SAP-DMD), a plug-and-play approach that adaptively modulates the amplitude spectrum of the DMD directional field. By suppressing the dominant tail of the amplitude spectrum, SAP-DMD reduces low-frequency dominance and promotes more effective recovery of fine structures and textures. Experiments on PixArt-α, SD3, and SD3.5 demonstrate that SAP-DMD accelerates training convergence and improves generation quality under both 2-step and 4-step sampling.
Figures & tables
Figure 1: Performance overview on SD3.5 Medium. (a) Training evolution of DMD2 and SAP-DMD. SAP-DMD recovers fine details earlier and improves DrawBench ( Saharia et al., 2022 ) evaluation metrics faster on all 200 prompts using 2-step generation without GAN loss. (b) 2-step samples generated by SAP-DMD, which show sharp structures and rich details across diverse prompts.
Figure 2: Frequency-domain diagnostics of the DMD directional error Δt on SD3.5 Medium. Statistics are computed at t=0.5 from a DMD2 checkpoint trained for 200 steps and averaged over 1,000 samples. (a) Radial amplitude profiles before and after purification; the dotted line denotes the mean channel-wise threshold Sˉ . (b) Spatial visualization of Δt before and after purification (top), with the corresponding log-amplitude distributions and channel-wise thresholds Sc (bottom). (c) Relative spectral energy shares of low-, mid-, and high-frequency components, defined by ∥u∥2<16 , 16≤∥u∥2<32 , and 32≤∥u∥2≤64 , respectively.
Method
NFE
GAN
MS-COCO
HPSv2.1
IR
PS
CS
Anime
Concept
Painting
Photo
Avg.
PixArt- α(512×512)
PixArt- α
25 × 2
0.7893
22.4067
30.47
30.87
29.04
28.69
28.43
29.26
YOSO
4
✓
0.7730
21.7805
27.87
30.43
30.44
30.38
27.42
29.67
DMD2 †
4
0.7945
22.2972
30.53
32.20
31.27
31.27
28.68
30.86
DMD2 †
4
✓
0.8047
22.4519
30.82
32.35
31.33
31.35
29.03
31.02
Table 1: Quantitative results of 4-step generation on MS-COCO and HPSv2.1. † : results from our re-implementation. ×2 indicates classifier-free guidance ( Ho and Salimans, 2022 ) requiring two NFEs per sampling step. GAN: with adversarial training enabled. Gray rows: multi-step base models. Best results are in bold and second best underlined.
Method
NFE
GAN
MS-COCO
HPSv2.1
IR
PS
CS
Anime
Concept
Painting
Photo
Avg.
SDXL
25 × 2
0.7586
22.4153
32.76
30.32
28.55
28.15
26.95
28.49
SD3 Medium
25 × 2
1.0026
22.5014
32.19
30.63
29.91
30.29
27.62
29.61
SD3.5 Medium
25 × 2
1.0208
22.4234
32.78
31.32
30.49
30.48
27.66
29.99
FLUX.1-dev
25
1.0653
23.0421
31.50
32.08
30.38
30.88
29.40
30.68
SDXL-Lightning
2
✓
0.6511
22.3895
31.28
30.86
29.53
29.49
27.23
29.28
Table 2: Cross-model comparison of 2-step generation on MS-COCO and HPSv2.1. † : results from our re-implementation. GAN: with adversarial training enabled. Gray rows: multi-step base models. Best results are in bold and second best underlined.
Table 5
Figure 3: Spectral evolution throughout SAP-DMD training on SD3.5 Medium. Statistics are averaged over 256 samples at t=0.5 with a checkpoint interval of 100 steps. (a) Low-frequency dominance RLF , defined as the ratio of mean amplitudes over ∥u∥2<16 and 16≤∥u∥2≤64 , remains consistently reduced after purification. (b) Relative amplitude suppression, measured as 1−AˉSAP(∥u∥2)/Aˉ(∥u∥2) , where Aˉ denotes the sample-averaged radial amplitude.
Figure 4: Qualitative comparison of 4-step (top) and 2-step (bottom) generation.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Radial amplitude profiles across representative diffusion timesteps on SD3.5 Medium. The DMD directional field consistently exhibits pronounced low-frequency amplitude concentration, while SAP-DMD suppresses dominant amplitudes and largely preserves weaker components.
Figure 6: Relative spectral energy shares across representative diffusion timesteps on SD3.5 Medium. Low-frequency energy becomes increasingly dominant in the raw DMD directional field as t increases. Across all timesteps, SAP-DMD substantially reduces this relative low-frequency dominance and increases the energy shares of mid- and high-frequency components.
Figure 7: Standardized log-amplitude distributions across representative diffusion timesteps on SD3.5 Medium. Statistics are computed at each timestep from a DMD2 trained for 200 steps and averaged over 128 samples, with each distribution standardized independently before aggregation. The dashed curve denotes a standard Gaussian for reference. The log-amplitude distributions are substantially more symmetric than the raw amplitude distributions and remain bell-shaped across timesteps, supporting the log-domain k -sigma boundary used by SAP-DMD.
Figure 8: Relative contributions of different source frequency bands to the parameter gradient on SD3.5. SAP-DMD consistently reduces the low-frequency contribution and increases the high-frequency contribution after backpropagation through the generator Jacobian. t=0.98 is used because the generator Jacobian vanishes at the terminal timestep t=1.0 .
Batch
TTUR
GAN
MS-COCO
HPSv2.1
A100
size
IR
PS
CS
Anime
Concept
Painting
Photo
Avg.
hours
4
1
1.0867
22.4414
31.87
34.16
33.17
33.35
29.99
32.67
10.56
4
5
1.0761
22.4248
31.92
34.01
32.96
33.09
29.78
32.46
28.82
4
1
✓
1.1104
22.5945
32.14
33.58
32.92
33.04
29.97
32.38
+11.09
32
1
1.1245
22.5483
32.00
34.28
33.29
33.32
30.50
32.85
90.34
32
1
✓
1.1444
22.6080
31.98
34.45
33.48
33.63
30.77
33.08
+86.39
Appendix
Table 5: Sensitivity of 2-step SAP-DMD on SD3.5 Medium to the training recipe. GAN: with GAN loss in the second training stage; + : additional A100 hours of this stage.
Method
MS-COCO
HPSv2.1
IR
PS
CS
Anime
Concept
Painting
Photo
Avg.
DMD2
1.0693±.0246
22.5196±.0359
31.4974±.0727
33.71±.40
33.11±.31
33.33±.30
30.17±.29
32.58±.32
SAP-DMD
1.1125±.0009
22.6240±.0227
31.6910±.1253
34.50±.10
33.80±.10
33.98±.13
30.85±.24
33.28±.13
Appendix
Table 6: Results of 4-step DMD2 and SAP-DMD on SD3.5 Medium without GAN loss over three matched training seeds, reported as mean ± standard deviation.
Backbone
DMD2
SAP-DMD
Δ
PixArt- α
0.568 ± 0.007
0.566 ± 0.006
−0.002
SD3
0.598 ± 0.009
0.617 ± 0.007
+0.019
SD3.5
0.656 ± 0.008
0.672 ± 0.005
+0.016
Appendix
Table 7: Sample diversity measured by average pairwise LPIPS (higher is more diverse) for 4-step generation. Results are reported as mean ± standard error over 128 prompts.
Figure 9: Early-stage high-frequency recovery on SD3.5, SD3, and PixArt- α . We report the normalized AUC of the high-frequency log-power discrepancy relative to paired 25-step teacher outputs over the first 25% of training. Bars show the mean over 128 prompts, and error bars denote 95% bootstrap confidence intervals.
Figure 10: Fine-detail recovery during training on PixArt- α with 4-step generation. Columns: training steps; both methods share the prompt and initial noise.
Figure 11: Fine-detail recovery during training on SD3 with 4-step generation. Columns: training steps; both methods share the prompt and initial noise.
Figure 12: Fine-detail recovery during training on SD3.5 with 4-step generation. Columns: training steps; both methods share the prompt and initial noise.
Step distillation has become a leading technique for accelerating diffusion models, among which Distribution Matching Distillation (DMD) and Consistency Distillation are two representative paradigms. While consistency methods enforce self-consistency along the full PF-ODE trajectory to steer it toward the clean data manifold, vanilla DMD relies on sparse supervision at a few predefined discrete timesteps. This restricted discrete-time formulation and mode-seeking nature of the reverse KL divergence tends to exhibit visual artifacts and over-smoothed outputs, often necessitating complex auxiliary modules -- such as GANs or reward models -- to restore visual fidelity. In this work, we introduce Continuous-Time Distribution Matching (CDM), migrating the DMD framework from discrete anchoring to continuous optimization for the first time. CDM achieves this through two continuous-time designs. First, we replace the fixed discrete schedule with a dynamic continuous schedule of random length, so that distribution matching is enforced at arbitrary points along sampling trajectories rather than only at a few fixed anchors. Second, we propose a continuous-time alignment objective that performs active off-trajectory matching on latents extrapolated via the student's velocity field, improving generalization and preserving fine visual details. Extensive experiments on different architectures, including SD3-Medium and Longcat-Image, demonstrate that CDM provides highly competitive visual fidelity for few-step image generation without relying on complex auxiliary objectives. Code is available at https://github.com/byliutao/cdm.
Tao Liu, Hao Yan, Mengting Chen +8
1VCIP, College of Computer Science, Nankai University · 2Alibaba Group · College of Artificial Intelligence, Jilin University
Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse-KL matching objective: Simplex-DMD uses continuous token relaxations and pathwise gradients, while Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. We develop both methods for multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, for sequences of 1,024 tokens, Simplex-DMD achieves a generative perplexity of 45.6 at a unigram entropy of 5.44 nats in just 4 NFEs, a 49% reduction relative to the strongest evaluated diffusion baseline at matched entropy and sampling budget. Reinforce-DMD improves the frontier at larger budgets, reaching a generative perplexity of 14.9 at an entropy of 5.00 nats with 256 NFEs, a 20% reduction under the same comparison protocol.
Paul Le Van Kiem, Dario Shariatian, Umut Simsekli +1
Inria, PSL Research University · Cohere · CMAP, Ecole Polytechnique
Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.
Zimo Wang, Junkun Yuan, Angtian Wang +9
University of California, San Diego · ByteDance Inc.