Researchers are exploring effective one-step generative model continuously, and, Drifting Models (Deng et al., 2026), demonstrate great potential in one-step generation recently. There are works that reveal the connection between Diffusion & Flow Style Generative Models (DFSGMs) (Ho et al., 2020; Song et al., 2020a;b; Lipman et al., 2022; Liu et al., 2022) and Drifting Models (Li & Zhu, 2026; Lai et al., 2026; Turan et al., 2026). But no one has yet established a precise correspondence between the Drifting Model and the widely used distillation method- Distribution Matching Distillation (DMD/DMD2) (Yin et al., 2024b;a) to the best of our knowledge, even though their optimization objective formulas are virtually identical. In this paper, we prove that by converting the velocity-field / noise-field from the pre-trained DFSGMs into the attraction force field in Drifting Models and estimating the repulsion force field from the generative distribution, training the Drifting Model is naturally equivalent to the Distribution Matching Distillation. With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).
Figures & tables
Method
Signal Scale ( αt )
Noise Scale ( σt )
Flow Matching
t
1−t
VP Diffusion
αt∈[0,1]
1−αt2
VE Diffusion
1
σt∈[0,σmax]
Table 1: Comparison of noise perturbation schedules across different generative frameworks.
Method
steps
FID ( ↓ )
DMD
328k
2.7238
DMD
410k
2.6432
MBDMD
328k
2.5728
MBDMD
410k
2.4697
Table 2 : Comparison of FID of MBDMD and vanilla DMD.
Method
K
Sampling
Best FID ( ↓ )
Iso-FLOP best FID ( ↓ )
DMD
1
N/A
2.6432
2.6432
DMD
3
i.i.d
2.5274
2.5702
MBDMD
3
strat
2.4697
2.5680
MBDMD
5
strat
2.4819
2.5109
Table 3 : Comparison of FIDs with different K and sampling method of MBDMD and vanilla DMD.
Method
Seg 1 avg. FID ( ↓ )
Seg 2 avg. FID ( ↓ )
Seg 3 avg. FID ( ↓ )
Seg 4 avg. FID ( ↓ )
Seg 5 avg. FID ( ↓ )
Seg 6 avg. FID ( ↓ )
Seg 7 avg. FID ( ↓ )
Seg 8 avg. FID ( ↓ )
MBDMD K=3
25.531
4.285
3.644
3.378
3.244
3.110
3.124
3.036
MBDMD K=5
27.350
4.405
3.652
3.421
3.215
3.141
3.114
3.055
MBDMD K=7
29.064
4.517
3.744
3.433
3.225
3.200
3.096
3.063
Table 4 : Comparison of mean FIDs in every sub-interval of MBDMD with different K .
Method
steps
FID ( ↓ )
DMD
157k
1.5643
DMD
200k
1.5518
MBDMD
157k
1.5304
MBDMD
200k
1.5304
Table 5 : Comparison of FID of MBDMD and vanilla DMD with GAN loss.
Method
Fwd Pass ( ↓ )
FID ( ↓ )
Patch FID ( ↓ )
Clip ( ↑ )
DMD2
1
19.01
26.98
0.336
MBDMD
1
18.81
26.86
0.335
Table 6 : Comparison of FID of MBDMD and vanilla DMD.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 1 : A 2D toy experiments to show that Fake Score Model can be replaced by the repulsion kernel function calculation. In this figure, (a) is the ground-truth; (b) is the results generated by pure kernel calculation for both attraction and repulsion; (c) shows the result by transferring velocity from pretrained Flow Matching Model to calculate attraction force and use kernel calculation for repulsion; (d) is the result from DMD which using velocity from pretrained Flow Matching Model to calculate attraction force and using a online updating Fake Score Model to calculate repulsion force.