Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial post-training for pixel diffusion. Across two pixel backbones, the method jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality. We further investigate why it works. Frequency-band and power-law analyses show that the original models systematically underproduce natural-image high-frequency content, while adversarial post-training restores this missing spectral power. In contrast, perceptual loss also increases high-frequency content but sacrifices distribution fidelity and prompt alignment. Nearest-neighbor, recall, and matched no-GAN SFT controls further rule out memorization, mode dropping, and additional optimization as simple explanations. Finally, we examine the boundary of this effect. Under the tested latent diffusion configurations, the same procedure does not produce comparable joint gains and adds almost no decoded high-frequency power. These results identify direct output access to the image statistics being corrected as a key factor governing when adversarial post-training succeeds.
Figures & tables
Figure 1 : Each pair shows w/o GAN (left) and w/ GAN (right), where w/ GAN means adding the adversarial loss during post-training: eight 5122 pairs across diverse prompts and styles (top) and two 1K pairs with zoom-ins (bottom).
w/o GAN
w GAN
w/o GAN
w GAN
w/o GAN
w GAN
w/o GAN
w GAN
Figure 2 : Same-prompt, same-seed comparison with and without GAN fine-tuning.
variant
DPG Score ↑
FID ↓
IS ↑
CMMD ↓
Rec ↑
pFID ↓
TOPIQ ↑
MUSIQ ↑
MANIQA ↑
DeCo (pixel)
no-GAN SFT
81.6
33.27
39.35
0.836
0.361
27.91
0.711
75.5
0.636
+GAN
83.3
28.59
39.77
0.736
0.406
24.38
0.768
76.5
0.712
PixelGen (pixel)
no-GAN SFT
78.4
33.94
38.58
0.762
0.319
30.81
0.755
75.7
0.645
+GAN
80.8
33.20
38.76
0.725
0.403
30.51
0.799
77.0
0.723
Table 1 : GAN vs. no-GAN on two pixel backbones. DPG Score is evaluated on DPG-Bench ( Hu et al., 2024 ) ; all other metrics use COCO-30k ( Lin et al., 2014 ) . Bold = better within each model.
Figure 3 : Pixel radial-profile band share (%) over 30,000 images/model; gray = no-GAN, green =+ GAN.
model
HF log- Δ
α w/o GAN
α+ GAN
DeCo
+0.34
2.59
2.24
real
–
≈2.19
Table 2: Pixel spectral statistics on COCO-30k; natural α≈2.19 .
variant
mean NN ↓
maximum NN ↓
no-GAN
0.586
0.943
+ GAN
0.586
0.929
Δ
+0.0001
−0.014
Table 3: DINOv2 nearest-neighbor similarity to ∼60 k-image training set.
Figure 4 : DPG Score ( Hu et al., 2024 ) trajectories initialized from the official, pre-SFT checkpoints. The GAN curves use an image-only DINOv2 discriminator. (a) DeCo; (b) PixelGen.
Saturation
Contrast
Colorfulness
Sharpness (Lap. var)
2-D HF spectral-energy ratio (%)
original
0.532
0.240
0.260
0.0076
4.3
+ perceptual
0.511
0.210
0.221
0.022
7.9
+ GAN
0.544
0.239
0.256
0.052
13.9
Table 4 : Image statistics on DPG-Bench ( Hu et al., 2024 ) . Perceptual supervision reduces color statistics; GAN preserves them while adding sharpness and 2-D HF spectral energy.
Original
Perceptual
+ GAN
Original
Perceptual
+ GAN
Original
Perceptual
+ GAN
Figure 5 : PixelGen outputs (original, perceptual, and + GAN; full image + zoom).
Figure 6 : Image-quality / preference win-rate vs. no-GAN (1,065 DPG-Bench matched pairs ( Hu et al., 2024 ) ). The pixel diffusion’s GAN (DeCo) is preferred across metrics, well above the latent diffusion’s GAN (SANA).
variant
DPG Score ↑
FID ↓
IS ↑
CLIP ↑
CMMD ↓
Rec ↑
pFID ↓
TOPIQ ↑
MUSIQ ↑
MANIQA ↑
DeCo (pixel)
no-GAN SFT
81.6
33.27
39.35
0.318
0.836
0.361
27.91
0.711
75.5
0.636
+GAN
83.3
28.59
39.77
0.319
0.736
0.406
24.38
0.768
76.5
0.712
SANA (latent)
no-GAN SFT (L1)
83.6
36.58
39.20
0.320
0.872
0.340
32.20
0.760
75.8
0.631
+GAN self-GAN (DiT)
83.4
36.24
38.75
0.321
0.853
0.325
33.29
0.727
76.2
0.614
+GAN feat-PatchGAN (L2)
83.5
38.41
36.41
0.320
0.915
0.334
35.18
0.764
76.0
0.646
+GAN dec-RGB PatchGAN (B1)
79.7
40.95
29.75
0.317
0.910
0.219
44.29
0.741
74.8
0.639
Table 5 : Pixel–latent comparison. DPG Score uses DPG-Bench ( Hu et al., 2024 ) ; other metrics use COCO-30k. Bold marks the better matched result.
Figure 7 : Frequency response of the self-GAN (DiT) variants in Table 5 .
model
HF log- Δ
α w/o GAN
α+ GAN
SANA
+0.09
2.54
2.80
PixArt
−0.14
2.72
2.86
real
–
≈2.19
Table 6 : Spectral statistics of the self-GAN (DiT) variants in Table 5 .
Figure 8 : The frozen PixArt VAE attenuates decoded-HF response by 3.5–11× relative to the pixel identity map.
Figure 9 : Noise-gate rationale. Across 1,000 prompts, structure stabilizes by t≈0.35−0.55 while detail remains incomplete; per-image residuals show the same coarse-to-fine pattern.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
discriminator
DPG Score ↑
FID ↓
IS ↑
CLIP ↑
CMMD ↓
Rec ↑
pFID ↓
TOPIQ ↑
MUSIQ ↑
MANIQA ↑
no-GAN SFT
81.6
33.27
39.35
0.318
0.836
0.361
27.91
0.711
75.5
0.636
no prior (from scratch)
PatchGAN ( t≥0.53 )
81.7
33.50
39.09
0.318
0.843
0.365
28.07
0.711
75.5
0.633
PatchGAN ( t≥0.20 )
81.9
32.51
39.09
0.318
0.843
0.388
26.44
0.719
75.9
0.666
StyleGAN
82.1
32.93
38.84
0.318
0.834
0.365
26.73
0.695
75.5
0.603
with prior (frozen feat.)
DINOv2
83.4
31.02
39.72
0.318
0.825
0.447
26.70
0.787
76.6
0.729
DINOv2-Large
83.2
30.70
40.15
0.318
0.822
0.447
26.81
0.789
76.6
0.730
Appendix
Table 7 : DeCo discriminator sweep. DPG Score uses DPG-Bench ( Hu et al., 2024 ) ; other metrics use COCO-30k. Gate t≥0.35 unless shown; DINOv2-text uses t≥0.53 . Shading is metric-wise.
gate
DPG Score ↑
FID ↓
IS ↑
CLIP ↑
CMMD ↓
Rec ↑
pFID ↓
TOPIQ ↑
MUSIQ ↑
MANIQA ↑
t≥0.0 (full)
82.6
31.63
37.91
0.319
0.755
0.427
26.12
0.767
76.56
0.727
t≥0.20
83.1
31.16
38.93
0.318
0.798
0.427
26.16
0.778
76.63
0.724
t≥0.35
83.4
31.02
39.72
0.318
0.825
0.447
26.70
0.787
76.58
0.729
t≥0.53
82.9
31.20
41.40
0.319
0.799
0.450
26.55
0.781
76.36
0.715
t≥0.80
82.3
31.99
40.38
0.319
0.811
0.412
26.98
0.766
76.01
0.694
Appendix
Table 8 : DeCo noise-gate sweep (fixed DINOv2, w=0.1 ). DPG Score is evaluated on DPG-Bench ( Hu et al., 2024 ) ; all other metrics use COCO-30k. Different thresholds trade prompt alignment, distribution fidelity, coverage, and no-reference quality; neither t≥0.35 nor t≥0.53 is uniformly superior.
weight
DPG Score ↑
FID ↓
IS ↑
CLIP ↑
CMMD ↓
Rec ↑
pFID ↓
TOPIQ ↑
MUSIQ ↑
MANIQA ↑
w=0.01
82.4
31.71
40.59
0.318
0.798
0.420
26.32
0.769
76.44
0.715
w=0.05
82.9
30.93
40.31
0.318
0.805
0.447
26.05
0.777
76.51
0.721
w=0.10
83.4
31.02
39.72
0.318
0.825
0.447
26.70
0.787
76.58
0.729
w=0.50
83.0
30.73
38.65
0.318
0.846
0.436
27.80
0.798
76.94
0.736
w=1.00
82.9
30.73
38.08
0.318
0.866
0.424
28.97
0.804
76.98
0.749
Appendix
Table 9 : DeCo GAN-weight sweep (DINOv2, gate t≥0.35 ). DPG Score is evaluated on DPG-Bench ( Hu et al., 2024 ) ; all other metrics use COCO-30k.
variant
DPG Score ↑
FID ↓
IS ↑
CMMD ↓
Rec ↑
pFID ↓
TOPIQ ↑
MUSIQ ↑
MANIQA ↑
DeCo
no-GAN SFT
81.6
33.27
39.35
0.836
0.361
27.91
0.711
75.5
0.636
perceptual (LPIPS+DINO)
81.4
34.14
38.95
0.826
0.367
28.64
0.749
76.2
0.661
+ GAN (DINOv2-text)
83.3
28.59
39.77
0.736
0.406
24.38
0.768
76.5
0.712
PixelGen
perceptual (native)
78.4
33.94
38.58
0.762
0.319
30.81
0.755
75.7
0.645
+GAN
80.8
33.20
38.76
0.725
0.403
30.51
0.799
77.0
0.723
Appendix
Table 10 : GAN vs. perceptual fine-tuning on two pixel backbones. DPG Score is evaluated on DPG-Bench ( Hu et al., 2024 ) ; all other metrics use COCO-30k. The perceptual arm raises no-reference sharpness but trades away distribution metrics; the GAN improves both. The DeCo GAN row uses the main DINOv2-text configuration. Bold marks the favorable value per column within each model.
PixArt no-GAN
PixArt +GAN
DeCo no-GAN
DeCo +GAN
PixArt no-GAN
PixArt +GAN
DeCo no-GAN
DeCo +GAN
full image
center zoom
Appendix
Figure 10 : Same-prompt output-access comparison (three prompts; full image + center zoom). PixArt no-GAN vs. + GAN are nearly identical, whereas DeCo + GAN visibly gains fine detail.
Figure 11 : Additional 512 px comparisons. Each adjacent pair: left without GAN, right with our GAN fine-tuning (DeCo, same prompt and seed). Best viewed zoomed in.
Figure 12 : Guidance and sampler order do not recover the GAN’s high frequency. No-GAN model swept over CFG and sampler order (orders 1 and 2 ); the + GAN model (CFG 4 , order 2 ) is the green dashed reference. Left: decoded high-frequency band share; right: MANIQA. Both no-GAN curves stay far below + GAN.
model
CFG
order
HF %
mid %
TOPIQ ↑
MANIQA ↑
no-GAN
3
1
0.010
0.190
0.502
0.392
no-GAN
3
2
0.014
0.221
0.520
0.404
no-GAN
4
1
0.012
0.204
0.510
0.398
no-GAN
4
2
0.015
0.234
0.516
0.403
no-GAN
5
1
0.013
0.215
0.512
0.401
no-GAN
5
2
0.017
0.245
0.505
0.398
Appendix
Table 11 : No-GAN CFG/order sweep vs. the + GAN model (DeCo, 300 fixed prompts/seeds). HF/mid bands are the decoded radial-power shares, and TOPIQ/MANIQA are no-reference quality metrics. Neither higher CFG nor order- 2 sampling approaches the + GAN high-frequency or quality.
method
HF log- Δ
α
FID ↓
TOPIQ ↑
MANIQA ↑
no-GAN
0.00
2.59
33.27
0.711
0.636
+ unsharp (HF-matched)
+0.35
2.38
32.7
0.719
0.673
+ unsharp (slope-matched)
+0.55
2.23
32.3
0.691
0.669
+ GAN (main, Table 1 )
+0.34
2.24
28.59
0.768
0.712
Appendix
Table 12: Spectrum-matched sharpening control (COCO-30k). A plain unsharp-mask filter applied to the no-GAN outputs, tuned to match the + GAN high-frequency spectrum (HF log- Δ and slope α ), reproduces the GAN’s spectral signature but recovers only a small fraction of its FID gain and does not reach its no-reference quality. The no-GAN and + GAN reference rows match Tables 1 and 2 .
Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both global structure and local texture in the same high-dimensional space. While recent work improves pixel diffusion through alternative prediction targets, training objectives, and architectures, these advances typically require training a new model from scratch. We show there is a cheaper, complementary strategy: \textbf{a frozen, pretrained pixel diffusion model can guide itself}. Our key observation is that intermediate layers of a pretrained pixel diffusion transformer can be decoded into coarse predictions that capture the main low-frequency structure, while the final layers progressively refine local, high-frequency details. We therefore attach a lightweight prediction head to an intermediate layer, keep the backbone frozen, and use the discrepancy between the intermediate and final predictions as a self-guidance direction during sampling. To train this head, we further find that real images are not necessary. Instead, model-generated samples suffice and even outperform real images for training the head, especially in enhancing the high-frequency components that pixel diffusion tends to underfit. Across multiple pixel diffusion models on ImageNet, our \textbf{Synthetic Self-Guidance (SSG)} consistently improves generation while adapter training requires less than 1% of full-model training compute: it reduces FID by over 50% across the evaluated JiT variants without classifier-free guidance (CFG) and further improves strong baselines with CFG, e.g., JiT-H/16 from 1.86 to 1.67 and PixelREPA-H/16 from 1.81 to 1.59. Our code is available at https://github.com/zfu006/SSG.
Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which usually denoises in a compact and structured latent space, pixel diffusion needs to learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose representation grounding that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-256x256, PixelDiT2 achieves an FID of 1.46 after 600 epochs; at 512x512 resolution, PixelDiT2 achieves an FID of 1.48 after 680 epochs.
Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA2. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA2 student performs better than the 25-step teacher and evaluated few-step distillers.