Researchers are exploring effective one-step generative model continuously, and, Drifting Models (Deng et al., 2026), demonstrate great potential in one-step generation recently. There are works that reveal the connection between Diffusion & Flow Style Generative Models (DFSGMs) (Ho et al., 2020; Song et al., 2020a;b; Lipman et al., 2022; Liu et al., 2022) and Drifting Models (Li & Zhu, 2026; Lai et al., 2026; Turan et al., 2026). But no one has yet established a precise correspondence between the Drifting Model and the widely used distillation method- Distribution Matching Distillation (DMD/DMD2) (Yin et al., 2024b;a) to the best of our knowledge, even though their optimization objective formulas are virtually identical. In this paper, we prove that by converting the velocity-field / noise-field from the pre-trained DFSGMs into the attraction force field in Drifting Models and estimating the repulsion force field from the generative distribution, training the Drifting Model is naturally equivalent to the Distribution Matching Distillation. With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).
Figures & tables
Method
Signal Scale ( αt )
Noise Scale ( σt )
Flow Matching
t
1−t
VP Diffusion
αt∈[0,1]
1−αt2
VE Diffusion
1
σt∈[0,σmax]
Table 1: Comparison of noise perturbation schedules across different generative frameworks.
Method
steps
FID ( ↓ )
DMD
328k
2.7238
DMD
410k
2.6432
MBDMD
328k
2.5728
MBDMD
410k
2.4697
Table 2 : Comparison of FID of MBDMD and vanilla DMD.
Method
K
Sampling
Best FID ( ↓ )
Iso-FLOP best FID ( ↓ )
DMD
1
N/A
2.6432
2.6432
DMD
3
i.i.d
2.5274
2.5702
MBDMD
3
strat
2.4697
2.5680
MBDMD
5
strat
2.4819
2.5109
Table 3 : Comparison of FIDs with different K and sampling method of MBDMD and vanilla DMD.
Method
Seg 1 avg. FID ( ↓ )
Seg 2 avg. FID ( ↓ )
Seg 3 avg. FID ( ↓ )
Seg 4 avg. FID ( ↓ )
Seg 5 avg. FID ( ↓ )
Seg 6 avg. FID ( ↓ )
Seg 7 avg. FID ( ↓ )
Seg 8 avg. FID ( ↓ )
MBDMD K=3
25.531
4.285
3.644
3.378
3.244
3.110
3.124
3.036
MBDMD K=5
27.350
4.405
3.652
3.421
3.215
3.141
3.114
3.055
MBDMD K=7
29.064
4.517
3.744
3.433
3.225
3.200
3.096
3.063
Table 4 : Comparison of mean FIDs in every sub-interval of MBDMD with different K .
Method
steps
FID ( ↓ )
DMD
157k
1.5643
DMD
200k
1.5518
MBDMD
157k
1.5304
MBDMD
200k
1.5304
Table 5 : Comparison of FID of MBDMD and vanilla DMD with GAN loss.
Method
Fwd Pass ( ↓ )
FID ( ↓ )
Patch FID ( ↓ )
Clip ( ↑ )
DMD2
1
19.01
26.98
0.336
MBDMD
1
18.81
26.86
0.335
Table 6 : Comparison of FID of MBDMD and vanilla DMD.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 1 : A 2D toy experiments to show that Fake Score Model can be replaced by the repulsion kernel function calculation. In this figure, (a) is the ground-truth; (b) is the results generated by pure kernel calculation for both attraction and repulsion; (c) shows the result by transferring velocity from pretrained Flow Matching Model to calculate attraction force and use kernel calculation for repulsion; (d) is the result from DMD which using velocity from pretrained Flow Matching Model to calculate attraction force and using a online updating Fake Score Model to calculate repulsion force.
Iterative generative models such as Flow Matching and Diffusion models have demonstrated strong test-time scaling behavior, where additional inference computation can improve generation quality. In contrast, Drift Models offer efficient one-step generation, but their direct generation paradigm limits such flexibility. In this work, we propose Drift Flow Matching (DFM), a framework that connects drifting generative modeling with flow-based iterative generation. DFM preserves the efficiency of direct transport maps while enabling generation to be refined through multiple inference steps when desired. This bridges the gap between one-step Drift Models and multi-step Flow Matching methods, and provides a novel generative paradigm that can adapt sampling computation to different quality--efficiency requirements. Extensive experiments across different tasks and datasets demonstrate the effectiveness and generality of the proposed framework.
Chenrui Ma, Xi Xiao, Lin Zhao +3
University of California, Irvine · University of Virginia · University of Alabama at Birmingham +1
Flow and diffusion models suffer from slow inference due to computationally expensive numerical integration. Distillation provides a promising way for a student model to learn from a teacher's dynamics, enabling one-step or few-step generation. However, existing methods often depend on curated distillation datasets, costly teacher rollouts, or auxiliary proxy networks, which complicate model training and scaling. In this work, we propose Consistent Distribution Matching, a simulation-free and data-free distillation method for accelerating diffusion and flow models while preserving strong generative capacity. Our key insight is to unify sample generation and score estimation with one student network. Thus, our framework uses only two models, a frozen teacher and a trainable student, and optimizes one objective. We prove that minimizing our objective indicates Wasserstein convergence of the student flow-map pushforwards to the teacher marginals. On ImageNet 256×256, our method attains an FID of 2.04 with a single function evaluation (1-NFE) and a 4-NFE FID of 1.37 within 40 epochs of training, surpassing the state-of-the-art distillation baselines without data. Our code code and model are available at https://consistentdmd.github.io/.
Yuxiang Fu, Qi Yan, Zike Wu +4
University of British Columbia · Vector Institute for AI · Canada CIFAR AI Chair
Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existing distillation methods often rely on multiple auxiliary networks, carefully designed training stages, or complex optimization pipelines. In this work, we revisit the recently proposed Drifting Model objective and show that a single drifting loss can be directly used to simplify one step distillation. A key observation is that the pretrained diffusion teacher itself already provides a strong representation space. Unlike the original Drifting Model, which relies on an additional pretrained feature extractor, we use intermediate hidden states of the pretrained teacher model as the feature representation. This removes the need for training or introducing an extra representation network while preserving a semantically meaningful feature geometry for drifting. Furthermore, we introduce a lightweight mode coverage loss to mitigate mode collapse during distillation and encourage the student generator to cover diverse teacher-supported regions. Extensive experiments on ImageNet and SDXL demonstrate that our method achieves efficient one step generation with competitive image quality and diversity, achieving FID scores of 1.58 on ImageNet-64×64 and 18.4 on SDXL, while substantially simplifying the overall distillation framework.
Yuan Zhang, Chenyi Li, Guoqing Ma +7
JD Explore Academy, China · The Hong Kong University of Science and Technology · Tsinghua University +1