In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet 256×256 study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
Figures & tables
Figure 1: Multi-embedding prediction. From the condition and the noisy image, the model predicts the next embeddings in the sequence, those of the clean image. Supervision is applied in embedding space; the predicted embeddings then condition a diffusion generator.
Figure 2: Multi-token vs. multi-embedding prediction. Both predict several items at once. (a) Tokens are discrete, so hidden states are decoded into tokens by output heads. (b) Embeddings are continuous, so the outputs of the model are the predictions themselves. In both, the same context predicts all targets in one forward pass.
Loss
FID ↓
Acc. (%) ↑
Cos. sim.
37.82
83.1
MSE
35.23
82.6
InfoNCE
31.36
81.7
Table 1: Loss. Acc. is ImageNet-1K top-1 accuracy of the fine-tuned NEPA model.
Method
Tokenizer
Epochs
FID ↓
sFID ↓
IS ↑
Pre. ↑
Rec. ↑
Other tokenizers
REPA + EQ-VAE ( Kouzelis et al., 2025a )
EQ-VAE
200
1.70
5.13
283.0
0.79
0.62
LightningDiT-XL/1 ( Yao et al., 2025 )
VA-VAE
800
1.35
4.15
295.3
0.79
0.65
LightningDiT + IG ( Zhou et al., 2026 )
VA-VAE
680
1.19
4.11
269.0
0.79
0.66
DiT-XL + CMuon ( Chen et al., 2026 )
VA-VAE
200
1.18
–
–
–
–
REPA-E ( Leng et al., 2025 )
E2E-VAE
800
1.12
4.09
302.9
0.79
0.66
Table 9: Class-conditional generation on ImageNet 256×256 with guidance. Top (gray): methods with modified or alternative tokenizers. Bottom: methods in the standard SD-VAE latent space. Metrics and training epochs on ImageNet-1K are taken from each paper; for ours, epochs are those of the NEPA model + the generator. “–” means not reported. ‡ AutoGuidance.
Scale
Interval
ODE-96
SDE-250
1.0
[0,1]
8.77
7.59
1.4
[0,1]
2.58
2.32
2.4
[0.3,1]
1.72
1.46
3.6
[0.4,1]
1.57
1.32
Table 10: Guidance. Applied only for t in the interval. Shaded: setting of Table 9 .
Figure 7: Qualitative results of NEPA-DiT-XL on ImageNet 256×256 .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
NEPA-B
NEPA-L
NEPA-XL
Architecture
Input dim.
32×32×4
32×32×4
32×32×4
Patch size
4
4
4
Num. layers
12
24
28
Hidden dim.
768
1,024
1,152
Num. heads
12
16
16
Appendix
Table 11: Hyperparameter setup of the NEPA model.
NEPA-DiT-B
NEPA-DiT-L
NEPA-DiT-XL
NEPA-DiT-XL + REPA
Architecture
Input dim.
32×32×4
32×32×4
32×32×4
32×32×4
Patch size
2
2
2
2
Num. layers
4+8
8+16
10+18
10+18
Hidden dim.
768
1,024
1,152
1,152
Num. heads
12
16
16
16
Appendix
Table 12: Hyperparameter setup of the generator. Num. layers counts two-stream + single-stream MM-DiT blocks. The last column is the model in Section 4.4 .
Params
GFLOPs / step
Model
Generator
NEPA
Generator
NEPA
TFLOPs / image
NEPA-DiT-B
135M
711M
59
92
–
NEPA-DiT-L
458M
711M
210
92
–
NEPA-DiT-XL, ODE-96
683M
711M
309
92
62
NEPA-DiT-XL, SDE-250
683M
711M
309
92
160
SiT-XL/2 + REPA, SDE-250
675M
–
229
–
91
Appendix
Table 13: Parameters and sampling FLOPs. GFLOPs per sampling step for one image; the NEPA model runs on the 64 image tokens with the condition cached. Totals include guidance on t∈[0.4,1] and VAE decoding. SiT-XL/2 is counted in the same way.
Figure 8: Samples of NEPA-DiT-XL on ImageNet 256×256 .
Figure 9: Samples of NEPA-DiT-XL on ImageNet 256×256 (continued).
Figure 10: Samples of NEPA-DiT-XL on ImageNet 256×256 (continued).
Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to p(x1∣xt) rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~\citet{Biroli2024}. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet-2562, achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at https://github.com/CompVis/iDDM.
Tommaso Martorella, Alexandre Galashov, Felix Krause +4
Munich Center for Machine Learning (MCML) · Google DeepMind
Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While methods improve the FID and related metrics, it is increasingly unclear whether they reflect real progress in generative modeling. The natural alternative, i.e., text-to-image (T2I) generation, is perceived as too costly or inconvenient to train and evaluate and is often skipped. We argue that this perception no longer holds. We introduce NanoGen, a unified DiT training and evaluation framework. NanoGen matches state-of-the-art DiT baselines on ImageNet and, with 12 lines of configuration change, also trains competitive text-to-image models. It currently supports RAE, VAE, pixel-space, and MeanFlow diffusion methods under both ImageNet and T2I setups. Under NanoGen, training T2I requires comparable compute to ImageNet. After training 21 latent diffusion models with NanoGen, we observe that method ranking shows no strong correlation between ImageNet and T2I generation: Pearson correlation is between -0.377 and -0.580 across three metrics. This suggests that a method which improves class-conditional ImageNet FID may show no corresponding improvement on T2I, clearly indicating the necessity of evaluating DiTs on both tasks. To this end, we summarize ImageNet and text-to-image results, which yields DiffusionBench, a holistic benchmark for DiT research. We recommend reporting DiffusionBench in place of ImageNet alone: methods that improve DiffusionBench are more likely to reflect broader progress.
Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. Therefore, attention emerges as a universal machine learning principle, based on a general training objective. We also show similarities in basic functional principle of auto-encoders and attention-based models. These equivalences allows us to interchange these designs based on practical requirements. As an example, we can reformulate the diffusion framework to reduce the lengthy training process and computation-intensive image generation. Using this approach, a simplified algorithm is proposed for image generation which is based on attention mechanism. Results show that the attention-based implementation achieves comparable performance with significantly less effort and computational resources.
Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3
School of Electrical Engineering, Iran University of Science & Technology, Tehran, IRAN · Independent researcher