Representation alignment has become an effective way to accelerate diffusion training, but its benefits do not transfer reliably to pixel-space clean-image prediction. In JiT, we find that auxiliary feature alignment can improve access to semantic features while reducing access to image variation needed for clean-image prediction, creating a mismatch between the auxiliary objective and the denoising task. This suggests a different principle: auxiliary supervision should improve the prediction target itself rather than impose a separate representation target. We introduce JAx (Just Align x), a prediction-supervision method that aligns clean-image predictions across noise levels. JAx couples a noisier student observation with a cleaner observation through a Markov degradation that preserves the original JiT input distribution. Under this coupling, the oracle prediction from the cleaner state has the same conditional mean as the optimal JiT target, while its conditional target covariance is no greater. Thus, oracle prediction alignment preserves the population JiT objective up to a constant while providing a lower-variance training target. To make this construction practical with an imperfect EMA teacher, JAx combines ground-truth supervision with a reliability-gated coupling band that selects nearby teacher states based on prediction risk. On ImageNet 256x256, JAx consistently improves FID and accelerates convergence across JiT-B/16, L/16, and H/16, without an external encoder or changes to the architecture or sampling procedure. Gradient diagnostics further show reduced minibatch gradient variance, while ablations demonstrate that the gains cannot be explained by time reweighting alone. These results show that prediction-space supervision provides a simple and principled alternative to representation alignment for pixel-space generative models.
Figures & tables
Figure 1: REPA initially improves FID on JiT-S/4 at ImageNet 32×32 , but the advantage reverses later in training.
Figure 2: Linear-probe observations at epoch 200. (A) REPA increases DINO-feature decodability in both JiT and SiT, while feature-complementary residual decodability decreases in JiT and increases in SiT. (B) JiT exhibits an alignment–repair detour in linear decodability: its residual deficit narrows toward the output. SiT shows gains at all measured layers.
Figure 3: Overview of JA x . An EMA teacher predicts the clean image from a cleaner observation, while the online student receives a Markov-coupled noisier observation. Training combines direct JiT supervision with prediction alignment, without modifying the architecture or sampling procedure.
Figure 4: Training convergence across JiT scales. Dashed and solid lines denote JiT and JA x , respectively. JiT-L/16 and H/16 baselines use the reported 200- and 600-epoch results.
Method
External representations
FID ↓
JiT
None
4.37
JiT + SRA 2
VAE
4.48
JiT + SRA
None
4.38
JiT + Self-Flow
None
4.31
JiT + REPA
DINOv2
5.12
JA x (ours)
None
4.17
Table 1: Auxiliary supervision on JiT-B/16 after 200 epochs. External representations are supplied by a separate pretrained encoder. Bold and underlining mark the best FID overall and without external representations, respectively.
Figure 5: Minibatch gradient variance relative to JiT at epoch 200 and epoch 600, after normalizing the combined JA x gradient by 1+λ . JA x has lower gradient variance across training stages.
Figure 6: Curated samples generated by JA x on ImageNet 256×256 using the standard JiT-H architecture and sampling procedure.
Method
Δℓ
FID ↓
Baseline and controls
Vanilla JiT
–
4.37
Reweighting only
–
4.39
Fixed gap
2
4.48
Conditional sampling
>0
4.43
Reliability-gated bands
Table 2: Ablations on JiT-B/16 at 200 epochs. RCB denotes the reliability-gated coupling band. Prediction-supervision variants use Markov coupling unless noted. FID uses 50K samples.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Target / coupling
ρ
E[Y∣Zt=z]
Var(Y∣Zt)
Clean image X
–
0.6z
0.80
Oracle, independent noise
0
0.3z
0.45
Oracle, shared noise
1
0.9z
0.05
Oracle, Markov coupling
0.5
0.6z
0.30
Appendix
Table 3: Analytical target statistics for X∼N(0,1) , t=1/3 , and s=1/2 . The required conditional mean is mt(z)=0.6z . These values are exact Gaussian calculations.
Figure 7: Comparison of auxiliary supervision strategies for generative modeling. (a) Auxiliary self-supervised learning, such as masked reconstruction. (b) Hidden-feature alignment with a frozen pretrained encoder. (c) SRA: alignment between shallow student and deeper EMA-teacher features. (d) Self-Flow: feature alignment with dual-timestep noise scheduling. (e) JA x : direct clean-prediction alignment under reliability-gated Markov coupling, without an additional trainable alignment head. All methods retain the base generative objective. Dashed arrows indicate EMA updates, and teacher targets are stop-gradient. Inputs are schematic; conditioning, latent encoding, and ground-truth target connections are omitted for clarity.
Configuration
JA x -B/16
JA x -L/16
JA x -H/16
Architecture
Depth
12
24
32
Hidden dimension
768
1024
1280
Attention heads
12
16
16
Image size
256×256
Patch size
16×16
Appendix
Table 4: Implementation details of JA x on ImageNet 256×256 . The backbone architectures and base training configurations follow JiT. The alignment weight is shared across model scales. Training-teacher EMA settings are listed separately from the EMA candidates used for generation evaluation.
Scale
FID ↓
IS ↑
JiT
JA x
JiT
JA x
B/16
3.66
3.41
275.1
271.82
L/16
2.36
2.26
298.5
284.93
H/16
1.86
1.82
303.4
302.8
Appendix
Table 5: Results on ImageNet 256×256 after 600 epochs, with unchanged JiT architectures and sampling procedures. Bold marks the lower FID at each scale.
Method
External components
Params (M)
Epochs
GFLOPs
FID ↓
IS ↑
Latent-space diffusion
DiT-XL/2 ( Peebles and Xie, 2023 )
VAE
675+49
1400
238
2.27
278.2
SiT-XL/2 ( Ma et al., 2024 )
VAE
675+49
1400
238
2.06
277.5
SiT-XL/2 + REPA ( Yu et al., 2025 )
VAE, DINOv2
675+49
200
238
1.96
264.0
SiT-XL/2 + REPA ( Yu et al., 2025 )
VAE, DINOv2
675+49
800
238
1.42
305.7
Pixel-space diffusion: existing systems
Appendix
Table 6: System-level comparison on ImageNet 256×256 . Parameters include the generator and tokenizer decoder used at inference, excluding training-only components. GFLOPs measure one generator forward pass excluding tokenizer computation. JA x improves FID over JiT at each evaluated scale without changing inference computation. The specific REPA encoder used by DeCo is unspecified.
Figure 8: Uncurated samples from Class 084 (peacock)
Figure 9: Uncurated samples from Class 141 (redshank)
Figure 10: Uncurated samples from Class 213 (Irish setter)
Figure 11: Uncurated samples from Class 352 (impala)
Figure 12: Uncurated samples from Class 388 (giant panda)
Figure 13: Uncurated samples from Class 406 (altar)
Figure 14: Uncurated samples from Class 829 (streetcar)
Figure 15: Uncurated samples from Class 927 (trifle)
Figure 16: Uncurated samples from Class 932 (pretzel)
Figure 17: Uncurated samples from Class 937 (broccoli)
Figure 18: Uncurated samples from Class 941 (acorn squash)
Flow samplers consume velocity, but the neural network can predict the clean endpoint and convert it to velocity through a fixed affine readout. We study this choice with JLT, a latent Transformer in a frozen variational autoencoder (VAE) representation. For squared error, the optimal clean and velocity predictors are algebraically equivalent; a finite Transformer assigns different computation to its learned output under the two interfaces. A local Gaussian analysis identifies a known residual response supplied by the readout and isotropic target variance added by velocity prediction. Measured FLUX.2 channel spectra support this geometric distinction: 90% of target variance occupies 83 of 128 clean directions versus 109 velocity directions. Under a matched velocity objective, clean prediction improves ImageNet FID-50K from 6.56 to 2.70 at Base scale and from 2.12 to 1.47 at Large scale, with lower FID at every measured Large checkpoint. Scaling clean prediction to 951M parameters reaches FID-50K 1.19 and IS 271.96. In addition, an objective ablation at Base scale shows that direct clean regression reaches FID-50K 2.38 without time-dependent error weighting. These results show how moving known computation outside the network changes learning under algebraically equivalent flow interfaces. Code: https://github.com/akatsuki-neo/JLT/blob/main/README.md
Funing Fu, Tenghui Wang, Guanyu Zhou +2
Independent Researcher · Wuhan University of Technology · Hangzhou Jiyi Artificial Intelligence Co., Ltd.
Representation alignment with pretrained vision models has recently shown strong potential for accelerating diffusion transformer training. By aligning intermediate diffusion features with clean-image representations from self-supervised vision encoders, existing methods improve convergence and generation quality. However, such alignment also introduces a non-trivial constraint: diffusion models operate on noisy inputs whose usable information varies across timesteps, while the reference features are extracted from clean images. In this paper, we revisit this mismatch from a token-level perspective. We find that, under full-token representation alignment, tokens with large alignment-gradient norms exhibit a stable spatial preference, suggesting that the alignment objective does not affect all tokens uniformly and may encourage the model to rely on the complete set of clean-image tokens. To address this issue, we propose MaskAlign, a token-subset representation alignment method that applies alignment to randomly sampled token subsets during training. By exposing the model to different token subsets across iterations, MaskAlign reduces the dependence of representation alignment on the complete token set and encourages alignment behavior that is more stable under token-subset perturbations. To mitigate the information loss caused by directly dropping tokens, we further introduce a lightweight pre-mask token mixing block that shares information across tokens before masking.
Lianyu Pang, Tianlin Pan, Cheng Da +5
The Hong Kong University of Science and Technology · Kuaishou Technology · University of Chinese Academy of Sciences
Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4% using the class token and 10.1% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.