Representation alignment has become an effective way to accelerate diffusion training, but its benefits do not transfer reliably to pixel-space clean-image prediction. In JiT, we find that auxiliary feature alignment can improve access to semantic features while reducing access to image variation needed for clean-image prediction, creating a mismatch between the auxiliary objective and the denoising task. This suggests a different principle: auxiliary supervision should improve the prediction target itself rather than impose a separate representation target. We introduce JAx (Just Align x), a prediction-supervision method that aligns clean-image predictions across noise levels. JAx couples a noisier student observation with a cleaner observation through a Markov degradation that preserves the original JiT input distribution. Under this coupling, the oracle prediction from the cleaner state has the same conditional mean as the optimal JiT target, while its conditional target covariance is no greater. Thus, oracle prediction alignment preserves the population JiT objective up to a constant while providing a lower-variance training target. To make this construction practical with an imperfect EMA teacher, JAx combines ground-truth supervision with a reliability-gated coupling band that selects nearby teacher states based on prediction risk. On ImageNet 256x256, JAx consistently improves FID and accelerates convergence across JiT-B/16, L/16, and H/16, without an external encoder or changes to the architecture or sampling procedure. Gradient diagnostics further show reduced minibatch gradient variance, while ablations demonstrate that the gains cannot be explained by time reweighting alone. These results show that prediction-space supervision provides a simple and principled alternative to representation alignment for pixel-space generative models.
Figures & tables
Figure 1: REPA initially improves FID on JiT-S/4 at ImageNet 32×32 , but the advantage reverses later in training.
Figure 2: Linear-probe observations at epoch 200. (A) REPA increases DINO-feature decodability in both JiT and SiT, while feature-complementary residual decodability decreases in JiT and increases in SiT. (B) JiT exhibits an alignment–repair detour in linear decodability: its residual deficit narrows toward the output. SiT shows gains at all measured layers.
Figure 3: Overview of JA x . An EMA teacher predicts the clean image from a cleaner observation, while the online student receives a Markov-coupled noisier observation. Training combines direct JiT supervision with prediction alignment, without modifying the architecture or sampling procedure.
Figure 4: Training convergence across JiT scales. Dashed and solid lines denote JiT and JA x , respectively. JiT-L/16 and H/16 baselines use the reported 200- and 600-epoch results.
Method
External representations
FID ↓
JiT
None
4.37
JiT + SRA 2
VAE
4.48
JiT + SRA
None
4.38
JiT + Self-Flow
None
4.31
JiT + REPA
DINOv2
5.12
JA x (ours)
None
4.17
Table 1: Auxiliary supervision on JiT-B/16 after 200 epochs. External representations are supplied by a separate pretrained encoder. Bold and underlining mark the best FID overall and without external representations, respectively.
Figure 5: Minibatch gradient variance relative to JiT at epoch 200 and epoch 600, after normalizing the combined JA x gradient by 1+λ . JA x has lower gradient variance across training stages.
Figure 6: Curated samples generated by JA x on ImageNet 256×256 using the standard JiT-H architecture and sampling procedure.
Method
Δℓ
FID ↓
Baseline and controls
Vanilla JiT
–
4.37
Reweighting only
–
4.39
Fixed gap
2
4.48
Conditional sampling
>0
4.43
Reliability-gated bands
Table 2: Ablations on JiT-B/16 at 200 epochs. RCB denotes the reliability-gated coupling band. Prediction-supervision variants use Markov coupling unless noted. FID uses 50K samples.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Target / coupling
ρ
E[Y∣Zt=z]
Var(Y∣Zt)
Clean image X
–
0.6z
0.80
Oracle, independent noise
0
0.3z
0.45
Oracle, shared noise
1
0.9z
0.05
Oracle, Markov coupling
0.5
0.6z
0.30
Appendix
Table 3: Analytical target statistics for X∼N(0,1) , t=1/3 , and s=1/2 . The required conditional mean is mt(z)=0.6z . These values are exact Gaussian calculations.
Figure 7: Comparison of auxiliary supervision strategies for generative modeling. (a) Auxiliary self-supervised learning, such as masked reconstruction. (b) Hidden-feature alignment with a frozen pretrained encoder. (c) SRA: alignment between shallow student and deeper EMA-teacher features. (d) Self-Flow: feature alignment with dual-timestep noise scheduling. (e) JA x : direct clean-prediction alignment under reliability-gated Markov coupling, without an additional trainable alignment head. All methods retain the base generative objective. Dashed arrows indicate EMA updates, and teacher targets are stop-gradient. Inputs are schematic; conditioning, latent encoding, and ground-truth target connections are omitted for clarity.
Configuration
JA x -B/16
JA x -L/16
JA x -H/16
Architecture
Depth
12
24
32
Hidden dimension
768
1024
1280
Attention heads
12
16
16
Image size
256×256
Patch size
16×16
Appendix
Table 4: Implementation details of JA x on ImageNet 256×256 . The backbone architectures and base training configurations follow JiT. The alignment weight is shared across model scales. Training-teacher EMA settings are listed separately from the EMA candidates used for generation evaluation.
Scale
FID ↓
IS ↑
JiT
JA x
JiT
JA x
B/16
3.66
3.41
275.1
271.82
L/16
2.36
2.26
298.5
284.93
H/16
1.86
1.82
303.4
302.8
Appendix
Table 5: Results on ImageNet 256×256 after 600 epochs, with unchanged JiT architectures and sampling procedures. Bold marks the lower FID at each scale.
Method
External components
Params (M)
Epochs
GFLOPs
FID ↓
IS ↑
Latent-space diffusion
DiT-XL/2 ( Peebles and Xie, 2023 )
VAE
675+49
1400
238
2.27
278.2
SiT-XL/2 ( Ma et al., 2024 )
VAE
675+49
1400
238
2.06
277.5
SiT-XL/2 + REPA ( Yu et al., 2025 )
VAE, DINOv2
675+49
200
238
1.96
264.0
SiT-XL/2 + REPA ( Yu et al., 2025 )
VAE, DINOv2
675+49
800
238
1.42
305.7
Pixel-space diffusion: existing systems
Appendix
Table 6: System-level comparison on ImageNet 256×256 . Parameters include the generator and tokenizer decoder used at inference, excluding training-only components. GFLOPs measure one generator forward pass excluding tokenizer computation. JA x improves FID over JiT at each evaluated scale without changing inference computation. The specific REPA encoder used by DeCo is unspecified.
Figure 8: Uncurated samples from Class 084 (peacock)
Figure 9: Uncurated samples from Class 141 (redshank)
Figure 10: Uncurated samples from Class 213 (Irish setter)
Figure 11: Uncurated samples from Class 352 (impala)
Figure 12: Uncurated samples from Class 388 (giant panda)
Figure 13: Uncurated samples from Class 406 (altar)
Figure 14: Uncurated samples from Class 829 (streetcar)
Figure 15: Uncurated samples from Class 927 (trifle)
Figure 16: Uncurated samples from Class 932 (pretzel)
Figure 17: Uncurated samples from Class 937 (broccoli)
Figure 18: Uncurated samples from Class 941 (acorn squash)