Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA2. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA2 student performs better than the 25-step teacher and evaluated few-step distillers.
Figures & tables
Figure 1 : Qualitative results of four-step text-to-image generation of DMA 2 ( 512×512 ).
Figure 2 : Overview of the proposed DMA 2 framework. DM-Band restricts teacher matching to a diagnostically selected high-noise range. For real-data guidance, DINO-Adv and AF-Loss share a frozen DINOv2 encoder and provide local patch guidance and a parameter-free semantic distribution-field target, respectively. Only Gθ , the fake-score critic, and the adversarial heads are trained; all auxiliary components are removed at inference.
Figure 3 : Separability diagnostic (C2ST). Each panel plots a held-out classifier’s real/fake separability ( y -axis: ROC-AUC over 1,536 matched real/generated pairs, 0.5 is chance) against the DeCo matching level t ( x -axis; larger t is less noise). The panels vary the setup: (a) the RGB input transform (raw texture, high-pass, low-pass, 32×32 patch shuffle); (b) the generation stage tg∈{0,.25,.5,.75} ; (c) the input representation (RGB, the SDXL latent, frozen DINOv2). RGB separability changes sharply with t and is high only near clean, whereas the SDXL latent stays flat and DINOv2 stays separable into the high-noise band.
Figure 4 : Eq. 1 terms across tm ; Δ exhibits broader coherent patterns at high noise and sparse fine-scale patterns near clean. Each panel is independently normalized for spatial visualization; colors are not comparable in magnitude across cells.
DPG-Bench ↑
U(.02,.98)
U(.02,.35)
Pixel (DeCo)
82.55
82.96
Latent (SDXL)
67.64
63.23
Table 5
DPG-Bench ↑
GenEval ↑
Train cost ↓
Adversary
1
4
1
4
s/step
#learn. [-1pt] (adv. path)
DMD2 Feat. GAN
78.35
82.12
0.6960
0.7709
1.82
560M
Frozen-DINOv2
81.26
82.55
0.7723
0.8155
0.72
1.58M
Table 6
Figure 5 : Anchor-Field Loss (AF-Loss). The generated image’s normalized final-layer DINOv2 CLS feature z=ϕ(x^) is attracted toward a rolling bank of real features Br and repelled from a bank of generated features Bf , giving a displacement V(z) . AF-Loss converts this displacement into a stopped target whose gradient flows only through z→x^→Gθ ; the field itself has no learnable parameters. The banks are detached FIFO queues capped at 256 features.
Teacher DMD2 Decoupled DMA 2 (ours)
Teacher DMD2 Decoupled DMA 2 (ours)
“A set of four green plastic food containers displayed against a stark white background, each captured from a distinct angle to showcase the varying perspectives …”
“An up-close image showcasing the intricate interior of a walnut, split cleanly down the middle to reveal its textured, brain-like halves …”
“Three sleek, dark wooden boats are resting along the banks of a tranquil, azure blue lake, their oars tucked neatly inside …”
“A sprawling field blanketed with vibrant wildflowers, where a tall giraffe and a striped zebra stand side by side under acacia trees …”
Teacher DMD2 Decoupled DMA 2 (ours)
Teacher DMD2 Decoupled DMA 2 (ours)
“A cat on a leather chair next to remotes.”
“A large dog sitting on top of a roof.”
Figure 6 : Qualitative comparison at four steps. Top block: long, densely specified DPG-Bench prompts (abbreviated with “…”); bottom block: short COCO captions. DMA 2 preserves object count and composition and recovers fine structure, across both long and short text.
Model
NFE
GenEval ↑
DPG ↑
VQA ↑
IS ↑
CLIP ↑
Rec. ↑
NIQE ↓
DeCo ( Ma et al., 2026 ) (official)
25
0.8216
81.580
0.7014
35.58
0.3200
0.3110
4.064
Flash ( Chadebec et al., 2025 )
4
0.7814
76.703
0.6900
29.94
0.3178
0.1486
5.688
DMD2 ( Yin et al., 2024a )
4
0.7709
82.115
0.7058
36.76
0.3170
0.4349
4.168
TDM ( Luo et al., 2025 )
4
0.7924
82.367
0.7071
36.48
0.3165
0.4447
3.871
Decoupled ( Liu et al., 2026b )
4
0.8100
83.095
0.7111
36.51
0.3195
0.3516
3.953
DMDR ( Jiang et al., 2026 )
4
0.8063
83.160
0.7023
34.22
0.3198
0.4414
3.710
Table 1 : Main comparison. All few-step methods are re-implemented on the same teacher; bold/underline mark the best/second-best students at each step count.
Module
Metric
DM-Band
DINO-Adv
AF-Loss
GenEval ↑
DPG ↑
VQA ↑
IS ↑
Rec. ↑
NIQE ↓
0.8037
81.115
0.6930
34.24
0.2802
5.174
✓
0.8155
82.550
0.7045
39.38
0.4297
3.337
✓
✓
0.8193
82.959
0.7077
40.21
0.4318
3.034
✓
0.8090
83.194
0.7079
39.46
0.3161
3.574
✓
✓
0.8178
83.427
0.7113
39.52
0.4324
3.307
Table 2 : Component ablation. Bold/underline mark the best/second-best results.
Figure 7 : Hyperparameter sensitivity in DPG, GenEval, and IS; green bands mark defaults. These are OFAT sweeps around an anchor that is not the full model, so their default and zero-weight points are not directly comparable to the Table 2 component ablation (anchor protocol in Appendix D ).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Time (s/step) ↓
DMD2 ( Yin et al., 2024a )
1.82
Decoupled DMD ( Liu et al., 2026b )
1.82
DMDR ( Jiang et al., 2026 )
0.90
DMA 2 (ours)
0.72
Appendix
Table 3 : Same-pipeline training speed. All methods use the same hardware, global batch size, and end-to-end step-timing protocol. Lower is better.
Model
NFE
GenEval ↑
DPG ↑
VQA ↑
IS ↑
CLIP ↑
Rec. ↑
NIQE ↓
PixelGen-XXL teacher
25
0.7965
79.360
0.6804
37.24
0.3141
0.3328
4.064
DMA 2 -PixelGen (ours)
4
0.8112
81.708
0.6897
40.46
0.3184
0.4892
3.088
DMA 2 -PixelGen (ours)
1
0.8072
80.898
0.6961
39.48
0.3232
0.3872
3.275
Appendix
Table 4 : Extension to a PixelGen-XXL teacher: DMA 2 few-step students vs. the original PixelGen teacher. Metrics as in Table 1 ; ↑ / ↓ give direction, bold is best and underline second best per column. The four-step student uses the strictly-rerun schedule-matched replica.
DMD matching range
U(.02,.20)
U(.02,.35)
U(.02,.70)
U(.15,.70)
U(.60,.90)
U(.02,.98)
U(.02,{.25,.50,.75,.98})
narrow
best
one-sided
wide
low-noise
full range
per-stage
GenEval ↑
0.8095
0.8193
0.8113
0.7924
0.7341
0.8155
0.8164
DPG-Bench ↑
82.720
82.959
82.426
82.038
80.330
82.550
82.018
CLIP ↑
0.3201
0.3225
0.3184
0.3213
0.3194
0.3214
0.3211
IS ↑
38.79
40.21
39.78
39.30
36.21
39.38
38.71
Recall ↑
0.4182
0.4318
0.4198
0.4062
0.3773
0.4297
0.4027
Appendix
Table 5 : Noise selection: matching-band position, width, and per-stage nesting. Four-step DINO-Adv students ( 20 k, λGAN=.01 , AF-Loss disabled) differ only in the DMD matching range; ↑ / ↓ give direction and bold is best per row. The high-noise band U(.02,.35) leads on nearly every alignment metric.
Objective
GenEval ↑
DPG ↑
VQA ↑
IS ↑
CLIP ↑
Rec. ↑
NIQE ↓
AF-Loss (CLS field, ours)
0.8090
83.194
0.7079
39.46
0.3246
0.3161
3.574
Direct DINOv2-CLS perceptual
0.8156
82.077
0.7028
37.85
0.3217
0.3146
4.675
DINOv2-CLS MMD (distribution matching)
0.8077
81.877
0.6949
35.77
0.3206
0.3330
4.433
Appendix
Table 6 : AF-Loss versus two alternative feature-space objectives on the same normalized last-layer CLS features, encoder, four-step protocol, and loss weight ( .05 ): direct DINOv2-CLS perceptual matching, and a matched multi-scale DINOv2-CLS MMD (distribution matching) against the same 256 -feature real bank. Only the feature-space objective changes; bold is best per column.
AF feature
DPG ↑
DINOv2 {2,5,8,11} spatial patch features (per-layer mean-pool, concat)
81.970
DINOv2 {2,5} spatial patch-feature mean+std
81.430
DINOv2 {2,5} spatial patch-feature quantile
81.760
DINOv2 normalized final-layer CLS feature (ours)
83.194
Appendix
Table 7: Anchor-Field feature ablation. All variants share the same field form at λAF=.05 ; the normalized final-layer CLS feature is our default.
Anchor-Field term
DPG ↑
Attract only (real bank Br )
82.910
Repel only (fake bank Bf )
80.170
Attract + repel (ours)
83.194
Appendix
Table 8: Anchor-Field attract/repel ablation, at the default feature and λAF=.05 .
Teacher DMD2 Decoupled DMA 2 (ours)
Teacher DMD2 Decoupled DMA 2 (ours)
“An intricate oil painting that captures two rabbits standing upright in a pose reminiscent of the iconic American Gothic portrait, in early 20th-century rural clothing …”
“A rider atop a chestnut horse in the middle of a spacious pasture enclosed by a wooden fence, dotted with patches of green grass …”
“A frisky golden retriever with a shiny, shaggy coat stands next to a life-sized penguin statue in the midst of a bustling public park …”
“A vibrant yellow rabbit, its fur almost glowing with cheerfulness, bounds energetically across a sprawling meadow, its sizeable red-framed glasses slipping comically …”
Teacher DMD2 Decoupled DMA 2 (ours)
Teacher DMD2 Decoupled DMA 2 (ours)
“A polar bear walking over rocks in its enclosure.”
“A teddy bear wearing a green robe.”
Appendix
Figure 8 : Additional four-step qualitative comparisons under identical prompts (columns: 25-step teacher, DMD2, Decoupled, and DMA 2 ). Top block: long DPG-Bench prompts (abbreviated with “…”); bottom block: short COCO captions.