Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a 256→512→1024 curriculum, after first ablating the prediction target and representation alignment at 2562 to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on 4× DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at 10242. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.
Figures & tables
Fig. 2: x - versus v -prediction at 2562 , absolute scores at matched steps (EMA, 25-step FlowDPM++, CFG 4.5). Following JiT [ 5 ] , the x -prediction arm also uses its logit-normal timestep distribution (0.8,0.8) instead of (0.0,1.0) , so the contrast does not isolate the prediction target.
Fig. 3: Representation-alignment ablations at 2562 , absolute scores at matched steps (same protocol as Fig. 2 ). All arms share the same model and data and differ only in the alignment objective. Each y-axis is broken to fit Self-Flow on the same scale.
Stage
Steps
Batch
Samples
256
0–367K
1024
375.8M
512 multi-AR
367K–530K
1024
166.9M
1024 multi-AR
530K–625K
512
48.6M
SFT
625K–665K
512
20.5M
Tab. 1: Training stages. Samples = steps × global batch.
Model
GenEval
DPG
LongText
OneIG
FLUX.1-dev 12B [ 33 ]
0.66
83.8
0.607
0.434
i1 3B ∗ [ 51 ]
0.84
86.7
0.922
–
Z-Image 6B ∗ [ 39 ]
0.84
88.1
0.935
0.546
Qwen-Image 20B [ 3 ]
0.87
88.3
0.943
0.539
Iris-3B
0.798
86.52
0.857
0.540
Tab. 2: Official benchmarks at 10242 . References are copied from the respective papers; FLUX.1-dev from Qwen-Image [ 3 ] (GenEval, DPG), X-Omni [ 49 ] (LongText) and OneIG-Bench [ 50 ] (OneIG, English overall). ∗ Uses prompt rewriting.
Fig. 5: FLUX.2 Klein conversion probes after only 1K steps at 10242 , one column per input initialization and trunk LR, with the median loss over steps 900–1000 (PiT head, global batch 8; “Gain ridge” is the gain-matched ridge). These are early diagnostic samples used to compare initializations, not final results; see Fig. 6 for the converted model.
Fig. 6: Samples from PiT 30K, the converted FLUX.2 Klein used downstream, at 10242 (CFG 4, 100 steps), generated directly in pixel space.
Fig. 7: Depth on four Hypersim test frames from the three arms of Tab. 4 , each trained with the same direct-regression recipe for 10K steps. Predictions are fitted to the ground truth with a scale and shift in log depth and shown as inverse depth on a shared per-row colour scale; black marks invalid ground truth.
NYUv2
KITTI
ETH3D
ScanNet
DIODE
Mean
Hypersim
Parent
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
SEE ↓
Klein latent
.050
.967
.086
.921
.058
.971
.060
.954
.107
.924
.072
.947
.528
Klein pixel
.064
.953
.087
.919
.065
.960
.075
.931
.117
.906
.081
.934
.526
Iris-3B
.057
.961
.083
.927
.064
.965
.069
.939
.082
.935
.071
.946
.496
Tab. 4: Depth after 10K steps of the same direct-regression recipe, zero-shot on five benchmarks. AbsRel ↓ and δ1↑ per benchmark, their mean over the five benchmarks, and SEE at kernel 3 on Hypersim ↓ .
Parent
Hypersim
KITTI
ETH3D
DIODE
Klein latent
0.97
0.99
0.99
0.98
Klein pixel
1.31
1.44
1.36
1.34
Iris-3B
2.14
1.61
2.21
2.25
Tab. 5: Patch-grid ratio of the raw depth predictions of Tab. 4 : mean absolute second difference of the prediction on pixels adjacent to 16×16 patch borders over its mean elsewhere, both image axes, 60 images per benchmark. 1 means no grid ↓ . NYUv2 and ScanNet are omitted because the latent model, which has no patch grid, already scores well below one there.
Fig. 8: 4× restoration of two DIV2K validation images with native Real-ESRGAN degradations (the Tab. 6 pairs). Left: the full Iris-3B output. Right: the marked regions for each method, cropped from full-resolution outputs. HYPIR-SD2 is the public release. All methods use an empty prompt and one forward pass. We chose the images and regions after viewing the outputs.
Model
PSNR ↑
SSIM ↑
LPIPS ↓
NIQE ↓
MUSIQ ↑
DeQA ↑
Klein latent
20.99
0.557
0.275
3.26
67.11
4.06
Klein pixel
20.87
0.504
0.274
4.11
67.03
4.04
Iris-3B
20.68
0.481
0.292
3.94
68.01
3.87
R-ESRGAN [ 25 ]
21.76
0.579
0.386
3.79
60.37
3.85
StableSR [ 20 ]
21.91
0.557
0.407
4.10
53.91
3.59
DiffBIR v2 [ 21 ]
21.29
0.476
0.471
3.39
66.21
3.88
Tab. 6: DIV2K 4× restoration on validation images with native 4× Real-ESRGAN degradations; every method receives the same saved input pairs and is scored with the same PyIQA metric settings and DeQA-Score revision. PSNR and SSIM are computed on RGB. Public methods run at their released settings, except that all text conditioning (including tag and caption generators) is replaced by an empty prompt. Grouped as ours, GAN, multi-step diffusion, one-step diffusion.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 9: Depth on the five zero-shot benchmarks after 10K steps of the same direct-regression recipe.
Fig. 10: 4× restoration of further DIV2K validation images with native Real-ESRGAN degradations (the Tab. 6 pairs).
Fig. 11: 4× restoration of further DIV2K validation images (continued).
Fig. 12: 4× restoration of real-world low-quality images from RealLQ250 [ 55 ] (no ground truth).
Fig. 13: 4× restoration of four real-world example inputs from the public HYPIR repository [ 12 ] (no ground truth).