In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement residual-stream burden and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet 2562. Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.
Figures & tables
Hierarchical design
Decoupling via long skip
Decoupling via cross-attention
ADM ( Dhariwal and Nichol, 2021 ) , SiD / SiD2 ( Hoogeboom et al., 2023 ; Hoogeboom et al., 2025 ) , PixelFlow ( Chen et al., 2025b )
PixNerd ( Wang et al., 2026a ) , DiP ( Chen et al., 2026b ) , DeCo ( Ma et al., 2026 ) , PixelDiT ( Yu et al., 2026 )
RIN ( Jabri et al., 2023 ) , HyperDiT ( He et al., 2026 ) , DuSPiT ( Bai et al., 2026 )
Table 1: A taxonomy of pixel-space diffusion architectures. Selected works. More details are in Appendix A .
Figure 1: Toy experiment. Each bar gives the number of directions needed to explain 90% of the hidden-state variance. The hidden states under x -prediction are low-dimensional, whereas those under v -prediction are not, illustrating the burden.
Figure 2: A few patch directions preserve recognizable structure. Each reconstruction keeps the indicated number of principal directions within every 16×16 patch. Eight directions preserve the shape of the mountain, and additional directions recover texture.
Prediction target
Clean
Velocity
Clean
Data space
Pixel
Pixel
Whitened
FID ↓
10.19
139.83
132.95
Effective rank
4.80
367.37
768
Stable rank
1.41
5.94
768
90% variance rank
8
668
692
Table 2: Whitening degrades x -prediction.
Figure 3: Selective filtering depends on target and patch geometry. Top: squared overlaps between the leading 64 input directions and embedding-Gram eigenvectors; a bright diagonal indicates alignment. Bottom: embedding gain on the leading and trailing 64 directions, normalized over all 768. Whitened axes retain the original clean-PCA order.
Figure 4: Class readability follows the burden. Linear-probe Top-1 accuracy along the residual stream for the plain DiT in pixel and whitened space, three decoupled v -prediction models (Section 4.3 ) and the shared workspace of SiHC (Section 5.2 ). Pixel x -prediction and the decoupled models are markedly more class-readable than pixel v -prediction and both whitened models. Columns fix the clean coefficient t , and shading shows the standard deviation across noise draws.
Figure 5: Semantic endpoints are far more concentrated than the velocity target. Target spectra provide a reference for the prediction demand. A faster rise means that fewer directions carry the variance. All three B-size models train successfully under v -prediction, reaching FID 6.33 – 13.31 compared with 139.83 for the plain DiT (Table 20 ).
Figure 6: Decoupled models filter their input like x -prediction. The top row shows squared overlaps between clean-patch principal directions and embedding-Gram eigenvectors, where a bright diagonal indicates alignment. The bottom row shows the share of input gain on the leading and trailing 64 clean-patch directions.
Figure 7: Persistent state and workspace in three residual designs, read from bottom to top. (a) A plain residual state is also the layer workspace. (b) HC carries several states but reads one width- C workspace for each update. (c) SiHC assigns each state a local input and prediction, tracked by color, and connects the states to shared computation through feature-wise read/write maps.
Prediction
Plain FID ↓
mHC FID ↓
Pixel, x
10.19
9.78
Pixel, v
139.83
25.36
DINOv2-B, x
5.50
3.93
DINOv2-B, v
16.56
3.98
Table 3: HC helps v -prediction most. All models are B-size and use Heun-50 with CFG 2.9 . Each pixel patch or DINOv2-B token has 768 coordinates.
Design
What changes
State region
Streams
FID ↓
Plain DiT
16×16
1
139.83
mHC, copied input
Four persistent states
16×16
4
25.36
mHC, spatial input
Local input and prediction per state
8×8
4
13.25
+ identity carry, scalar maps
Static maps without state mixing
8×8
4
11.50
+ feature-wise maps (SiHC)
Channel-specific read and write
8×8
4
10.10
+ 4×4 subpatches
More states, smaller regions
4×4
16
8.07
Table 4: Each design choice improves v -prediction. Parameters and GFLOPs remain nearly unchanged across the progression.
Model
Prediction
Params (M)
GFLOPs
FID ↓
IS ↑
JiT-H/16 ( Li and He, 2026 )
x
953
364
1.86
303.4
JiT-G/16 ( Li and He, 2026 )
x
2000
767
1.82
292.6
PixelREPA-H/16 ( Shin et al., 2026 )
x
953
364
1.81
317.2
PixNerd-XL/16 † ( Wang et al., 2026a )
v
700
268
1.89
309.0
DiP-XL/16 † ( Chen et al., 2026b )
v
631
255
1.75
273.0
DeCo-XL/16 † ( Ma et al., 2026 )
v
682
245
1.69
304.0
Table 5: Pixel-space generation on ImageNet- 2562 . We compare with a selection of recent pixel-space models under a common sampling setting: all rows use Heun-50, and our evaluations follow JiT’s FID protocol ( Li and He, 2026 ) . Some of these models achieve better results with their own sampling configurations, for example DeCo and PixelDiT. These results and further details are given in Appendix A.5 . † Uses REPA; –: not reported.
Appendix figures & tables41 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Noisy-input path
Connection to backbone
Prediction computation
Analytic output paths
JiT [ Li and He, 2026 ]
Noisy input enters both the patch embedding and an analytic output skip.
Final patch features predict the clean image.
Linear patch head, then analytic conversion to velocity.
AsymFlow [ Chen et al., 2026a ]
An analytic input path recovers velocity outside a chosen noise subspace.
A JiT backbone predicts full-dimensional data minus low-rank noise.
Patch head with projection-based analytic velocity recovery.
Decoders conditioned on final backbone features
DiP [ Chen et al., 2026b ]
Noisy RGB patches enter local U-Nets.
Context tokens enter U-Net bottlenecks.
Patch-local convolutions with down/up blocks and skips.
DeCo [ Ma et al., 2026 ]
Noisy RGB and positions form pixel queries.
Upsampled coarse features supply AdaLN-zero conditioning.
Pixel-wise residual MLPs without attention.
Appendix
Table 6: Prediction interfaces in pixel-space models. Analytic skips, conditioned decoders and cross-attention provide different paths from noisy inputs to predictions. SiHC connects persistent local states to shared computation through linear read/write maps.
Model
Params (M)
GFLOPs
FID ↓
IS ↑
Latent diffusion with VAE
DiT-XL/2 [ Peebles and Xie, 2023 ]
675
238
2.27
278.2
SiT-XL/2 [ Ma et al., 2024 ]
675
238
2.06
277.5
SiT-XL/2 + REPA [ Yu et al., 2025 ]
675
238
1.42
305.7
LightningDiT-XL/2 [ Yao et al., 2025 ]
675
238
1.35
295.3
DDT-XL/2 [ Wang et al., 2026c ]
675
238
1.26
310.6
Appendix
Table 7: Generation quality and model cost on ImageNet- 2562 . Reported results span pixel space and latent spaces based on VAEs or RAEs. GFLOPs count one denoiser forward pass; training and sampling schedules follow each source. REPA is indicated explicitly for the pixel models that use it. –: not reported. ∗ Computed from the published or released configuration (Appendix E.6 ).
Figure 8: Corruption spreads patch variance across directions. Left: variance along clean-PCA directions. Right: cumulative variance. Decreasing the clean coefficient t flattens the spectrum and slows the cumulative rise, so more directions are needed to explain the same fraction of patch variance.
Figure 9: Whitening broadens both prediction targets. Left: target covariance eigenvalues. Right: cumulative variance. The whitened spectra assign equal variance to every direction, so their normalized cumulative curves coincide. Raw clean patches concentrate variance in a small leading subspace.
Figure 10: Cumulative gain shows which patch directions enter the stream. Each curve sums embedding gain in clean-PCA order. A faster rise means that more gain is assigned to the leading clean directions.
Figure 11: Leading patch directions preserve recognizable structure across images. Each row is one image. Columns show the original and reconstructions retaining 8, 16, 32, 64 and 128 principal directions within each 16×16 patch. More directions recover fine texture.
Figure 12: Video tubelets have concentrated spectra. Left: cumulative variance, with the marker at r90 . Right: covariance eigenvalues normalized by total variance.
Paper t
0.1
0.3
0.5
0.7
0.75
0.9
Affine clean r90
1
2
3
5
5
7
Affine velocity r90
674
656
638
603
588
488
Appendix
Table 8: Affine clean prediction remains concentrated across noise levels. Each entry gives the number of directions explaining 90% of predicted variance. Compare clean and velocity at each t .
Figure 13: Target and embedding spectra across coordinates. Cumulative target covariance (left) and embedding-Gram eigenvalues (right), each sorted independently and normalized by its sum. Pixel x -prediction has the most concentrated target and embedding spectra; whitening removes this concentration.
Gain share (%)
Pixel, v
Pixel, x
DeCo
PixelDiT
DiP
HyperDiT
Top 256
41.7
95.3
96.9
98.0
86.5
97.9
Bottom 256
28.0
0.2
0.2
0.3
5.4
0.3
Appendix
Table 9: Embedding gain assigned to the leading and trailing 256 clean-PCA directions, as a percentage of total gain over all 768 directions. Pixel x is the plain DiT with x -prediction, and all other models use v -prediction.
Model
k=16
k=32
k=64
k=128
k=256
Pixel, x
0.903
0.925
0.921
0.909
0.949
Pixel, v
0.076
0.095
0.268
0.530
0.652
DeCo-B
0.948
0.946
0.930
0.927
0.929
PixelDiT-B
0.920
0.877
0.885
0.891
0.935
DiP-B
0.924
0.929
0.884
0.726
0.770
HyperDiT-B
0.925
0.867
0.874
0.891
0.912
Appendix
Table 10: Embeddings under x -prediction and decoupled designs align with the leading clean-patch subspace. Higher mean squared cosine indicates stronger overlap. Columns expand the compared subspaces from k=16 to k=256 .
Figure 14: DINO targets are broad under both prediction choices. Left: covariance eigenvalues of RAE-normalized DINOv2-B tokens and the corresponding velocity target. Right: cumulative variance, with markers at r90 . Adding unit Gaussian variance broadens an already distributed clean spectrum.
Target
ER
SR
r90
r95
r99
DINO, x
579.6
66.84
563
649
735
DINO, v
708.3
121.31
651
708
755
Appendix
Table 11: DINO target ranks. Statistics use the 768-dimensional spectra in Figure 14 ; definitions follow Appendix B .
Figure 15: Semantic endpoints concentrate variance relative to the velocity target. Panels fix t=0.1,0.5,0.9 . Solid curves show the three semantic endpoints, with clean and velocity target spectra as dashed and dotted references. Direction rank uses a logarithmic axis to make the leading subspace visible.
Figure 16: Decoupled v -prediction models concentrate both main-stream interfaces. (a) Lower r90 means fewer directions explain 90% of output-feature variance. The dashed lines mark raw-target references. (b) Faster-rising curves indicate a more concentrated embedding spectrum. (c) Gain along clean-PCA directions reveals which input variations enter the stream: pixel v -prediction retains a broader range than x -prediction and the decoupled models.
Setting
Value
Covariance fitting set
100,000 ImageNet training images, 100 per class
Selection / flip seed
20260806 / 20260807
Preprocessing
ADM center crop to 2562 , RGB in [−1,1] ; one deterministic horizontal-flip draw per selected image
Patch coordinates
Non-overlapping 16×16 RGB; channel, patch row, patch column order
Observations
100,000×256=25.6 M patches of dimension 768
Centering
One global 768-dimensional mean over all patches
Appendix
Table 12: Patch covariance and endpoint measurement protocols. Each covariance treats image-position pairs as observations.
Figure 17: The learned input skip approaches the clean-to-velocity coefficient during training. Both surfaces show the clean coefficient t , training epoch and the learned skip coefficient. (a) −α(t) , with the analytic reference 1/(1−t) drawn at the final recorded epoch. (b) −(1−t)α(t) , with the corresponding reference at one. Solid blue lines show the final learned curves; dashed orange lines show the references. The surfaces use 246 recorded training snapshots, with denser sampling early in training and shape-preserving interpolation between the 20 measured time points. (c) FID at all ten evaluated checkpoints, using 50,000 samples, EMA weights and Heun-50 sampling; markers denote measurements.
Epoch
Clean FID
Clean IS
Velocity FID
Velocity IS
40
298.46
2.60
238.82
2.85
80
166.89
5.82
243.77
2.29
120
148.95
7.64
204.79
2.55
160
140.39
8.61
209.37
3.01
200
132.95
9.54
212.54
3.35
Appendix
Table 13: Training history in whitened space. Lower FID and higher IS are better.
Setting
Value
Weights
Raw epoch 200 / step 250,200
Images
2,048 ImageNet training images, uniform sampling without replacement
Corruption
128 Gaussian draws per image at each t∈{0.25,0.50,0.75}
Pairing
Identical image indices and deterministic Gaussian draws across models/times
Input / conditioning
Pixel or whitened training coordinates (Appendix B.9 ); null class 1000; no image flips
Observation
Attention/MLP input before normalization and adaptive modulation
Appendix
Table 14: Normalized sensitivity protocol. The spatial token grid is preserved when computing variance.
Figure 18: Image-to-noise ratio and class readability through depth. Top: variation between images relative to variation from noise ( B/W ). Bottom: class readability measured by linear-probe Top-1 accuracy. Curves cover the plain DiT in pixel and whitened space, the decoupled models and SiHC. Columns fix the clean coefficient t .
Figure 19: Image and noise variation evolve differently through depth. Columns fix the clean coefficient t . The top and middle rows show noise variance W/D and between-image variance B/D , with D the state dimension. The bottom row shows their ratio B/W . Larger ratios indicate stronger image differences relative to corruption. Vertical axes are logarithmic, with a shared range within each row.
t=0.25
t=0.50
t=0.75
Model
Sˉ
B/D
Rˉ
Sˉ
B/D
Rˉ
Sˉ
B/D
Rˉ
Pixel, v
0.164
0.081
1.04
0.133
0.142
1.89
0.071
0.221
3.90
Pixel, x
0.116
0.137
1.50
0.045
0.192
4.39
0.014
0.234
17.69
DeCo
0.090
0.098
1.41
0.035
0.149
4.44
0.010
0.196
19.07
PixelDiT
0.086
0.197
2.78
0.032
0.224
7.87
0.010
0.241
28.83
DiP
0.083
0.121
1.72
0.049
0.162
3.56
0.036
0.182
11.24
Appendix
Table 15: Depth-averaged image and noise variation. Each entry averages a per-observation statistic. In particular, Rˉ averages the individual ratios rather than dividing the averaged variances. Complete curves appear in Figure 19 .
Figure 20: SiHC’s late workspace is less sensitive to corruption than pixel v -prediction. Columns fix t . Top: noise-induced variance per coordinate ( W/D ). Bottom: between-image variation relative to noise ( B/W ). Vertical ranges differ between panels.
Setting
Value
Dataset / ambient dimension
8,192 points / 512; dataset seed 42
Construction
Keep coordinates 0 and 2 of the Swiss roll, project with a Gaussian-QR 512×2 basis, standardize each ambient coordinate
Standardization
Subtract empirical mean, divide by empirical standard deviation plus 10−6
Table 17: Pixel-space training settings. B/L/H scaling uses 200 epochs without representation alignment. The extended H and XL runs use 600 epochs.
SiHC-B
SiHC-L
SiHC-XL
SiHC-H
Depth
12
24
40
32
Hidden dimension
768
1,024
1,024
1,280
Attention heads
12
16
16
16
First block with context
5
9
9
9
REPA
No
No
Yes
No
Parameters (M)
131.0
459.5
762.8
953.8
Appendix
Table 18: Configurations of SiHC models. Block indices are one-based. Parameter counts and GFLOPs cover the denoiser, excluding the REPA encoder and projector.
Experiment
Solver
EMA
CFG
Interval
Pixel B-size / ablations
Heun50
0.9999
2.9
[0.1,1]
SiHC-L
Heun50
0.9999
2.4
[0.1,1]
SiHC-H
Heun50
0.9999
2.1
[0.1,0.9]
SiHC-XL + REPA
Heun50
0.9999
2.4
[0.1,0.9]
Whitened x / v
Heun50
0.9999
2.9
[0.1,1]
DINO/RAE plain / mHC
Heun50
0.9999
2.9
[0.1,1]
Appendix
Table 19: Generation protocols for our experiments. All rows use 50,000 samples. Guidance intervals are in the paper’s clean-time convention.
Training epochs
Model
40
80
120
160
200
DeCo-B
63.84
19.42
12.64
9.54
8.07
DiP-B
82.13
27.74
19.54
15.89
13.31
PixelDiT-B
67.44
14.83
9.44
7.35
6.33
Appendix
Table 20: Generation trajectories of B-size decoupled models. FID on 50,000 ImageNet- 2562 samples at the indicated training epochs, using Heun50 and CFG 2.9. All runs use seed 0.
Configuration
b
S
Read/write
Carry
FID
mHC copied input
16
4
mHC
mixed
25.36
mHC spatial input
8
4
mHC
mixed
13.25
Static scalar access
8
4
scalar
identity
11.50
SiHC, spatial input
8
4
feature-wise
identity
10.10
SiHC, direct input
4
16
feature-wise
identity
8.07
+ 32 in-context class tokens
4
16
feature-wise
identity
6.17
Appendix
Table 21: B-size ablation configurations. All models use v -prediction. b is local input side length in pixels, and S is the number of streams per computational patch.
Model
Depth × width
Params
GFLOPs
FID ↓
IS ↑
DiT-only configurations reported by DiP
DiT-only
26×1152
629M
221.2∗
5.28
243.8
DiT-only
32×1152
772M
272.0∗
4.91
251.7
DiT-only
26×1280
776M
272.0∗
4.28
249.6
DiT-only
26×1536
1.1B
389.3∗
2.83
285.6
SiHC, 200 epochs, without REPA
Appendix
Table 22: Scaling pixel-space Transformers on ImageNet- 2562 . The upper group reproduces DiP’s DiT-only results with its Euler-100 protocol [ Chen et al., 2026b ] . The lower group uses 200 training epochs without REPA and Heun-50, with CFG 2.9/2.4/2.1 for B/L/H. GFLOPs count one denoiser forward with two FLOPs per multiply-add. ∗ Analytical estimates from the published configurations and released backbone (Appendix E.6 ).
Figure 21: Scaling SiHC without representation alignment. Left: B/L/H at 200 epochs, with 32 in-context class tokens and sublayerwise read/write. Right: continued training of SiHC-H, including all measured points from epoch 200 onward. Generation quality levels off around FID 2.06 .
Figure 22: Two execution granularities of SiHC. Blue boxes are persistent spatial states; orange boxes operate on the compact workspace. Blockwise SiHC writes the combined attention/MLP update once. Sublayerwise SiHC writes after attention and reads again for the MLP. Both use identity carry and the same spatial input/output assignment.
Figure 23: Fusion removes intermediate wide-state materialization. Blue tensors hold the S persistent streams; orange tensors have workspace width C . All base reads are computed from the same stage input. The triangular recurrence restores the effect of earlier writes before each nonlinear update. Only the final operation reconstructs the wide state.
Operation
Computation / implementation
Output shape
Coefficient contraction
Reduce AℓPj over streams for j<ℓ
T×T×C
All base reads
Read X0 with every Aℓ ; separate outputs feed compiled consumers
T tensors Q×C
Active triangular sum
At update ℓ , accumulate the previous ℓ updates with Γℓj
Q×C
Nonlinear update
Attention/MLP, normalization and conditioning on the compact workspace
Q×C
Final accumulation
Reduce all writes, add X0 , store the final spatial states
Q×S×C
Appendix
Table 23: Fused interface operations. T counts residual updates in one stage and Q=BbatchM .
Training
Inference
Size
Routing
ms
GiB
ms
GiB
Blockwise
B
Torch
123.42
26.82
35.90
3.05
B
Fused
97.05
18.90
26.54
3.01
L
Torch
365.24
70.83
102.17
5.05
L
Fused
292.77
50.33
76.76
4.99
Appendix
Table 24: Compiled SiHC with and without fused routing. Batch size 128 on H100 80GB. Training measures an optimizer step with two EMAs, and inference measures one denoiser forward. Memory is peak allocated GiB. OOM denotes failure at the same requested batch size.
Setting
Hidden-state probes
Frozen backbone
Raw epoch 200; null-class conditioning
Data
All 1,281,167 training and 50,000 validation images
Image preprocessing
ADM center crop; no random horizontal flips
Noise coefficient
α=0.25,0.50,0.75 , with paper time t=1−α
Feature / head
Per-token RMS ( 10−6 ) then GAP; affine 768→1000
Training
40 epochs; batch size 4,096
Appendix
Table 25: Linear-probe fitting and evaluation. Every boundary and noise level has an independent affine classifier.
Model
25% noise
50% noise
75% noise
Pixel, x
25.32 / 46.45
27.54 / 49.64
22.82 / 43.91
Pixel, v
5.38 / 14.51
4.56 / 12.68
3.47 / 10.17
Pixel SiHC, v
25.51 / 46.68
27.81 / 50.32
22.88 / 43.98
DINO, x
80.15 / 95.47
80.09 / 95.41
80.06 / 95.36
DINO, v
81.09 / 95.74
81.07 / 95.81
80.93 / 95.77
Appendix
Table 26: State organization and prediction space shape class readability. Entries give Top-1 / Top-5 accuracy (%) at the last measured boundary. Columns increase input noise, and larger values indicate more linearly readable class information.
Figure 24: SiHC recovers clean-like class readability while predicting velocity. Panels fix the input-noise level. Linear-probe Top-1 accuracy measures class information through depth: SiHC approaches pixel x -prediction and remains well above pixel v -prediction. Shading denotes standard deviation across noise draws.
Figure 25: Separate pixel paths support stronger semantic representations. Columns fix the clean coefficient t . Top: between-image variation relative to noise ( B/W ). Bottom: linear-probe Top-1 accuracy through depth. Compare the decoupled v -prediction models with pixel x - and v -prediction. Shading denotes standard deviation across noise draws.
Figure 26: DINO-space layer inputs retain high class readability under both prediction targets. Linear-probe accuracy follows block boundaries for DINO (orange) and pixel (blue) features. Solid lines use x -prediction and dashed lines use v -prediction. All curves use 50% input noise. Shading denotes standard deviation across noise draws. The input control is reported in Table 27 .
Matched input noise
Top-1 (%)
Top-5 (%)
0%
79.372
95.036
25%
79.872±.037
95.267±.011
50%
80.485±.072
95.545±.011
75%
80.820±.088
95.847±.042
Appendix
Table 27: Noisy DINO inputs retain strong class information. Rows increase input noise. Top-1 and Top-5 measure linear-probe accuracy, and ± denotes standard deviation across noise draws.
Leveraging representation encoders for generative modeling offers a path for efficient, high-fidelity synthesis. However, standard diffusion transformers fail to converge on these representations directly. While recent work attributes this to a capacity bottleneck proposing computationally expensive width scaling of diffusion transformers we demonstrate that the failure is fundamentally geometric. We identify Geometric Interference as the root cause: standard Euclidean flow matching forces probability paths through the low-density interior of the hyperspherical feature space of representation encoders, rather than following the manifold surface. To resolve this, we propose Riemannian Flow Matching with Jacobi Regularization (RJF). By constraining the generative process to the manifold geodesics and correcting for curvature-induced error propagation, RJF enables standard Diffusion Transformer architectures to converge without width scaling. Our method RJF enables the standard DiT-B architecture (131M parameters) to converge effectively, achieving an FID of 3.37 where prior methods fail to converge. Code: https://github.com/amandpkr/RJF
Flow samplers consume velocity, but the neural network can predict the clean endpoint and convert it to velocity through a fixed affine readout. We study this choice with JLT, a latent Transformer in a frozen variational autoencoder (VAE) representation. For squared error, the optimal clean and velocity predictors are algebraically equivalent; a finite Transformer assigns different computation to its learned output under the two interfaces. A local Gaussian analysis identifies a known residual response supplied by the readout and isotropic target variance added by velocity prediction. Measured FLUX.2 channel spectra support this geometric distinction: 90% of target variance occupies 83 of 128 clean directions versus 109 velocity directions. Under a matched velocity objective, clean prediction improves ImageNet FID-50K from 6.56 to 2.70 at Base scale and from 2.12 to 1.47 at Large scale, with lower FID at every measured Large checkpoint. Scaling clean prediction to 951M parameters reaches FID-50K 1.19 and IS 271.96. In addition, an objective ablation at Base scale shows that direct clean regression reaches FID-50K 2.38 without time-dependent error weighting. These results show how moving known computation outside the network changes learning under algebraically equivalent flow interfaces. Code: https://github.com/akatsuki-neo/JLT/blob/main/README.md
Funing Fu, Tenghui Wang, Guanyu Zhou +2
Independent Researcher · Wuhan University of Technology · Hangzhou Jiyi Artificial Intelligence Co., Ltd.
Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which usually denoises in a compact and structured latent space, pixel diffusion needs to learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose representation grounding that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-256x256, PixelDiT2 achieves an FID of 1.46 after 600 epochs; at 512x512 resolution, PixelDiT2 achieves an FID of 1.48 after 680 epochs. Project page: https://pixeldit.github.io/pixeldit2/