Representation alignment (REPA) accelerates diffusion transformer training, but its alignment targets are almost exclusively semantic encoders such as DINOv2 and CLIP. Recent analysis points to spatial structure, not global semantics, as the carrier of the alignment effect, yet dense-prediction foundation models trained to predict that structure remain overlooked as REPA targets. In pixel-space diffusion, SAM2, Depth Anything v2, and Metric3D v2 each outperform the DINOv2-only GenEval baseline, with the two geometric teachers leading the segmentation teacher. A flat sum of all four teachers, however, lands below the best single geometric teacher, as semantic and geometric gradients compete for one denoiser projection. We introduce PixelDense, which routes DINOv2 and SAM2 through a semantic projection stream, routes Depth Anything v2 and Metric3D v2 through a geometric projection stream, and adds a weight-space orthogonality penalty that keeps the two streams in disjoint subspaces. All four teachers are frozen during training and dropped at inference. Applied to PixelGen and DeCo with a single recipe, PixelDense improves GenEval, DPG-Bench, and HPS v2.1, raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093, and beats every single-teacher and unfactored multi-teacher variant. In partial-noise reconstruction, independent panoptic, depth, and surface-normal probes show up to 53.1% PQ gain and 36.0% depth AbsRel reduction at τ=0.5 across COCO and Flickr30K. From random initialization, PixelDense also reaches the baseline's peak GenEval 1.23x faster. In SDEdit editing on PIE-Bench, PixelDense keeps more of the source background and layout at every edit strength, raising background PSNR by up to 2.2 dB.
Figures & tables
Figure 2 : PixelDense factors dense-prediction supervision into a semantic stream and a geometric stream. The pixel diffusion denoiser fθ takes the noisy image xt and a Qwen3 text embedding y , and reads its block- 8 token state into two parallel projections, a Ds=768 semantic projection Wsem aligned to frozen DINOv2 and SAM2 features, and a Dg=1024 geometric projection Wgeo aligned to frozen Depth Anything v2 and Metric3D v2 features. A weight-space orthogonality penalty Lorth on the rows of Wsem and Wgeo keeps the two streams from collapsing into a shared subspace of the denoiser feature. Only the denoiser, the projections, and the per-teacher heads receive gradients. The four dense teachers are frozen during training and dropped at inference, so the deployed model is the original three-channel RGB generator with its original sampler.
Model
Type
#Params
GenEval ↑
DPG ↑
HPS ↑
Sing. Obj.
Two Obj.
Count.
Colors
Attr.
Pos.
Overall
SDXL [ 40 ]
latent
2.6B
98.00
74.00
39.00
85.00
23.00
15.00
0.55
74.7
–
SD3 [ 6 ]
latent
8B
98.00
84.00
66.00
74.00
43.00
40.00
0.68
–
–
PixelDiT (1024 res) [ 62 ]
pixel
1.3B
100.00
94.00
70.00
90.00
65.00
53.00
0.78
83.7
0.250
PixelFlow [ 5 ]
pixel
0.9B
–
–
–
–
–
–
0.60
77.9
–
PixelGen [ 35 ]
pixel
1.1B
99.00
88.00
59.00
90.00
70.00
70.00
0.79
-
0.281
Table 1 : PixelDense lifts compositional alignment over the matched PixelGen fine-tune at no inference-time cost. GenEval [ 9 ] , DPG-Bench [ 17 ] , and HPS v2.1 [ 54 ] on the released PixelGen-XXL [ 35 ] backbone at 512×512 . Our rows use the 10,000 -step, effective-batch- 256 , 2× H200 recipe of Section 4.1 . Higher is better for every metric. Bold marks PixelDense entries that improve over the matched PixelGen fine-tune.
Figure 3 : Qualitative geometry preservation comparison between PixelGen and PixelDense. We show noised/reconstructed images and pseudo labels for segmentation, depth, and normals. PixelGen shows more ghosting and structural distortions, while PixelDense better preserves GT geometry.
Dataset
Reference
τ
OneFormer PQ ↑
OneFormer mIoU ↑
Depth AbsRel ↓
Normal err. ( ∘ ) ↓
PixelGen
PixelDense
PixelGen
PixelDense
PixelGen
PixelDense
PixelGen
PixelDense
COCO val
Real panoptic
0.5
23.23
31.43
37.67
45.47
–
–
–
–
0.7
15.70
21.31
30.64
36.59
–
–
–
–
0.9
7.99
9.64
22.19
24.26
–
–
–
–
Pseudo orig.
0.5
29.53
41.16
38.70
47.99
1.58/1.06
1.06/0.68
21.80
17.84
0.7
18.93
26.51
30.73
37.32
2.09/1.54
1.63/1.12
26.90
23.48
Table 2 : Partial-noise reconstruction shows stronger geometry preservation for PixelDense. PixelGen and PixelDense reconstruct x^0 from xt=(1−τ)x0+τε at three noise levels. OneFormer PQ/mIoU are averaged over Swin-L and DiNAT-L COCO panoptic models; Depth cells report mean/median AbsRel after per-image affine fitting; normal cells report mean angular error in degrees. Pseudo references are obtained by applying the predictor to the original image after the shared resize and center crop. Higher is better for PQ and mIoU; lower is better for depth and normals.
Table 5
Figure 4 : PixelDense converges faster on GenEval from random initialization.
Model
GenEval ↑
DPG ↑
Sing. Obj.
Two Obj.
Count.
Colors
Attr.
Pos.
Overall
DeCo T2I, released checkpoint [ 34 ]
100.00
92.00
72.00
91.00
79.00
80.00
0.8600
81.4
DeCo + fine-tune
99.38
94.44
73.75
93.62
79.00
77.00
0.8620
81.4
PixelDense
100.00
93.43
74.38
93.62
80.00
80.00
0.8690
81.8
Table 5 : The PixelDense recipe transfers to the DeCo backbone without retuning. The teachers, losses, hyperparameters, and checkpoint-selection rule used on PixelGen-XXL are applied unchanged to DeCo [ 34 ] at 512×512 . Higher is better for every metric, and bold marks PixelDense entries that improve over the matched DeCo fine-tune.
Metric
Model
Noise level τ
0.3
0.4
0.5
0.6
0.7
0.8
0.9
Background, outside the edit mask
LPIPS ↓
PixelGen
0.22
0.24
0.26
0.28
0.30
0.32
0.35
PixelDense
0.19
0.21
0.23
0.24
0.26
0.29
0.32
PSNR (dB) ↑
PixelGen
26.07
24.72
23.27
21.76
20.17
18.19
15.57
PixelDense
28.00
26.74
25.39
23.92
22.30
20.23
17.10
Table 6 : PixelDense keeps more of the source image in SDEdit editing on PIE-Bench. Only the checkpoint differs between the two models. Background metrics are computed outside the edit mask, and FID is computed against the source images. Bold marks PixelDense entries that improve over PixelGen.
Figure 5 : Qualitative comparison between PixelGen and our PixelDense on text-to-image generation. PixelDense better preserves object counts, spatial relations, color binding, and scene geometry than PixelGen, yielding stronger prompt alignment and visual coherence.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Teacher
Backbone
Tap
Feature dim.
Group
DINOv2 [ 37 ]
ViT-B/14 patch tokens
last layer
768
semantic
SAM2 [ 43 ]
Hiera-L image encoder
last layer
256
semantic
Depth Anything v2 [ 60 ]
ViT-L encoder
layer 17
1024
geometric
Metric3D v2 [ 16 ]
ViT-L encoder
layer 17
1024
geometric
Appendix
Table 7 : Frozen dense-prediction teachers used by PixelDense. All teachers operate on the clean image r at 512×512 , are frozen during training, and are dropped at inference. Feature dimension is the per-token width before normalization. Group is the semantic-versus-geometric assignment used by Section 3.3 .
Figure 6 : Semantic and geometric teachers light up complementary spatial structure on the same image. The leftmost column shows the RGB input, and the next four columns show the per-patch RMS magnitude of spatially normalized features from DINOv2, SAM2, Depth Anything v2, and Metric3D v2, with yellow marking high response. DINOv2 fires on object regions and SAM2 fires on instance boundaries, while Depth Anything v2 traces depth-ordered silhouettes and Metric3D v2 traces surface-normal layout.
Run
Pearson ↑
Δcos↓
Pearson −Δcos↑
Mean red. ↑
Sem. red.
Geo. red.
Sem./Geo. gap ↓
Reduction CV ↓
Naive sum, equal weight
0.9057
0.7430
0.1627
0.7213
0.7573
0.6854
0.0719
0.0973
PixelDense
0.9083
0.5966
0.3116
0.7340
0.7337
0.7343
0.0006
0.0629
Appendix
Table 8 : PixelDense preserves long-term co-improvement across teachers while breaking the short-term lockstep that the naive sum imposes. Run-level coupling and balance diagnostics on the four cosine alignment losses logged every 50 optimizer steps for 10,000 steps, giving 200 points per run. Pearson is the off-diagonal mean of the 4×4 cross-loss correlation on those 200 points. Δcos is the off-diagonal mean of the cosine on the 199 consecutive step-to-step loss deltas. Pearson −Δcos isolates long-term co-improvement from short-term lockstep. Mean reduction is the average per-teacher relative loss reduction between the first and last ten logged points. Sem. and Geo. reduction restrict that average to the two semantic teachers and the two geometric teachers, and Sem./Geo. gap is their absolute difference. Reduction CV is the across-teacher coefficient of variation of the four reductions. Higher Pearson, higher Pearson −Δcos , higher mean reduction, lower Sem./Geo. gap, and lower CV are better.
DINOv2
SAM2
DA2
M3D
Naive sum, equal weight, off-diagonal mean 0.7430
DINOv2
1.0000
0.7287
0.7120
0.6950
SAM2
0.7287
1.0000
0.6799
0.6671
DA2
0.7120
0.6799
1.0000
0.9756
M3D
0.6950
0.6671
0.9756
1.0000
PixelDense, off-diagonal mean 0.5966
Appendix
Table 9 : Step-to-step loss-delta cosine matrices on the four cosine alignment losses, naive sum versus PixelDense. Each off-diagonal entry is the cosine of step-to-step deltas of the two named per-teacher losses, computed on the 199 consecutive deltas of the 200 logged points per run. The top sub-block reports the naive equal-weight sum, and the bottom sub-block reports PixelDense. The within-geometric Depth Anything v2 against Metric3D pair stays near 0.95 under both objectives, reflecting the shared 3D structure of the two geometric teachers. The other five off-diagonal entries collapse only under PixelDense, with the largest drop on the within-semantic DINOv2 against SAM2 pair.
Run
DINOv2
SAM2
DA2
M3D
Mean
Naive sum, equal weight
83.10%
68.35%
72.81%
64.26%
72.13%
PixelDense
78.49%
68.26%
77.49%
69.37%
73.40%
Appendix
Table 10 : Per-teacher relative loss reduction over 10,000 training steps, naive sum versus PixelDense. The reduction is computed between the first ten and last ten logged points of the 200 -point training log per run. Higher values indicate that more of a teacher’s signal has been absorbed. Under the naive sum, Metric3D is the slowest teacher in the bank at 64.26% and DINOv2 is the fastest at 83.10% . PixelDense narrows the spread to within 10 points across the four teachers and raises the mean from 72.13% to 73.40% .
Dataset
Reference
τ
PQ All↑
PQ Th↑
PQ St↑
mIoU ↑
PixelGen
PixelDense
PixelGen
PixelDense
PixelGen
PixelDense
PixelGen
PixelDense
COCO val
Real panoptic
0.5
23.25
31.54
26.32
35.96
18.62
24.86
37.53
45.36
0.7
15.69
21.31
17.16
23.93
13.48
17.37
30.61
36.49
0.9
8.02
9.72
7.92
9.96
8.17
9.36
22.16
24.32
Pseudo orig.
0.5
29.72
41.33
31.90
44.39
26.43
36.71
38.62
47.79
0.7
19.01
26.65
19.95
28.14
17.59
24.39
30.76
37.38
Appendix
Table 11 : OneFormer Swin-L panoptic scores per dataset, reference type, and noise fraction. Reconstructions of x^0 are scored against the indicated reference at 512×512 . PQ is reported on All, Things, and Stuff. Higher is better, and bold marks the PixelDense column.
Dataset
Reference
τ
PQ All↑
PQ Th↑
PQ St↑
mIoU ↑
PixelGen
PixelDense
PixelGen
PixelDense
PixelGen
PixelDense
PixelGen
PixelDense
COCO val
Real panoptic
0.5
23.21
31.32
26.06
35.68
18.91
24.75
37.80
45.59
0.7
15.70
21.31
17.07
23.63
13.65
17.80
30.67
36.70
0.9
7.97
9.57
7.90
9.67
8.08
9.42
22.22
24.20
Pseudo orig.
0.5
29.34
40.99
31.23
43.66
26.49
36.95
38.77
48.19
0.7
18.86
26.37
19.66
27.57
17.65
24.55
30.69
37.26
Appendix
Table 12 : OneFormer DiNAT-L panoptic scores per dataset, reference type, and noise fraction. Same protocol as Table 11 , with DiNAT-L replacing Swin-L. The two backbones agree to within 0.005 on PQ and mIoU per cell.
Dataset
τ
Mean AbsRel ↓
Median AbsRel ↓
PixelGen
PixelDense
PixelGen
PixelDense
COCO val
0.5
1.5805
1.0639
1.0559
0.6752
0.7
2.0919
1.6286
1.5393
1.1179
0.9
2.8269
2.6105
2.2362
2.0290
Flickr30K
0.5
1.6061
1.0257
1.1443
0.6770
0.7
2.2186
1.6639
1.7021
1.2144
Appendix
Table 13 : Marigold v1.1 depth AbsRel against pseudo references, mean and median. Mean and median per-image AbsRel after closed-form per-image scale and shift fitting, lower is better. The median AbsRel agrees with the mean on every cell, so the comparison is not driven by a few badly aligned outliers.
Dataset
τ
Mean ( ∘ ) ↓
Median per-image ( ∘ ) ↓
Mean of per-image medians ( ∘ ) ↓
PixelGen
PixelDense
PixelGen
PixelDense
PixelGen
PixelDense
COCO val
0.5
21.80
17.84
20.79
16.48
15.25
11.76
0.7
26.90
23.48
26.61
22.80
19.97
16.50
0.9
33.20
31.94
33.70
32.36
26.50
24.93
Flickr30K
0.5
22.45
18.25
21.03
16.88
16.27
12.26
0.7
28.12
24.62
27.55
23.59
21.88
17.84
Appendix
Table 14 : Marigold v1.1 surface-normal angular error against pseudo references, mean and per-image-median statistics. Mean pixelwise angular error in degrees, median per-image angular error, and mean of per-image medians. Lower is better. Per-image medians track the means on every cell, so the comparison is not driven by tail pixels.
Figure 7 : Additional PixelDense samples on the PixelGen-XXL backbone at 512×512 . Object-centered photographs, painterly portraits, atmospheric landscapes, and stylized fantasy scenes. The dense teachers are dropped at inference, and sampling matches the protocol of Section 4.1 .
Figure 8 : Additional partial-noise reconstruction comparisons between PixelGen and PixelDense. Each row shows the noised image at the indicated τ , the two reconstructions, and the panoptic, depth, and surface-normal probes from OneFormer [ 18 ] and Marigold [ 20 ] . PixelDense reconstructions track object boundaries, depth ordering, and surface orientation more faithfully than PixelGen at every τ .
Model
Anime
Concept-art
Paintings
Photo
Average
PixelGen-XXL T2I, released checkpoint
0.292
0.278
0.269
0.284
0.281
DeCo T2I, released checkpoint
0.285
0.267
0.259
0.277
0.272
PixelGen + DINOv2-only fine-tune
0.290
0.277
0.268
0.285
0.280
PixelGen + PixelDense, ours
0.292
0.280
0.272
0.284
0.282
Appendix
Table 15 : HPS v2.1 per-style breakdown across the four prompt categories of Wu et al. [54] . Each split contains 800 prompts, and the rightmost column averages the four splits. Higher is better, and bold marks PixelDense entries that improve over the matched PixelGen fine-tune.
Role
Asset
License
Pixel backbone
PixelGen-XXL [ 35 ]
Apache 2.0 use
Pixel backbone
DeCo-T2I [ 34 ]
Apache 2.0 use
Semantic teacher
DINOv2 [ 37 ]
Apache 2.0
Semantic teacher
SAM2 [ 43 ]
Apache 2.0
Geometric teacher
Depth Anything v2 [ 60 ]
Apache 2.0
Geometric teacher
Metric3D v2 [ 16 ]
BSD 2-Clause
Appendix
Table 16 : Public assets used in this paper, with the license under which each is released. All assets are used without modification. The training corpus and evaluation prompt sets are used through their official releases.