In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet 256×256 study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
Figures & tables
Figure 1: Multi-embedding prediction. From the condition and the noisy image, the model predicts the next embeddings in the sequence, those of the clean image. Supervision is applied in embedding space; the predicted embeddings then condition a diffusion generator.
Figure 2: Multi-token vs. multi-embedding prediction. Both predict several items at once. (a) Tokens are discrete, so hidden states are decoded into tokens by output heads. (b) Embeddings are continuous, so the outputs of the model are the predictions themselves. In both, the same context predicts all targets in one forward pass.
Loss
FID ↓
Acc. (%) ↑
Cos. sim.
37.82
83.1
MSE
35.23
82.6
InfoNCE
31.36
81.7
Table 1: Loss. Acc. is ImageNet-1K top-1 accuracy of the fine-tuned NEPA model.
Method
Tokenizer
Epochs
FID ↓
sFID ↓
IS ↑
Pre. ↑
Rec. ↑
Other tokenizers
REPA + EQ-VAE ( Kouzelis et al., 2025a )
EQ-VAE
200
1.70
5.13
283.0
0.79
0.62
LightningDiT-XL/1 ( Yao et al., 2025 )
VA-VAE
800
1.35
4.15
295.3
0.79
0.65
LightningDiT + IG ( Zhou et al., 2026 )
VA-VAE
680
1.19
4.11
269.0
0.79
0.66
DiT-XL + CMuon ( Chen et al., 2026 )
VA-VAE
200
1.18
–
–
–
–
REPA-E ( Leng et al., 2025 )
E2E-VAE
800
1.12
4.09
302.9
0.79
0.66
Table 9: Class-conditional generation on ImageNet 256×256 with guidance. Top (gray): methods with modified or alternative tokenizers. Bottom: methods in the standard SD-VAE latent space. Metrics and training epochs on ImageNet-1K are taken from each paper; for ours, epochs are those of the NEPA model + the generator. “–” means not reported. ‡ AutoGuidance.
Scale
Interval
ODE-96
SDE-250
1.0
[0,1]
8.77
7.59
1.4
[0,1]
2.58
2.32
2.4
[0.3,1]
1.72
1.46
3.6
[0.4,1]
1.57
1.32
Table 10: Guidance. Applied only for t in the interval. Shaded: setting of Table 9 .
Figure 7: Qualitative results of NEPA-DiT-XL on ImageNet 256×256 .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
NEPA-B
NEPA-L
NEPA-XL
Architecture
Input dim.
32×32×4
32×32×4
32×32×4
Patch size
4
4
4
Num. layers
12
24
28
Hidden dim.
768
1,024
1,152
Num. heads
12
16
16
Appendix
Table 11: Hyperparameter setup of the NEPA model.
NEPA-DiT-B
NEPA-DiT-L
NEPA-DiT-XL
NEPA-DiT-XL + REPA
Architecture
Input dim.
32×32×4
32×32×4
32×32×4
32×32×4
Patch size
2
2
2
2
Num. layers
4+8
8+16
10+18
10+18
Hidden dim.
768
1,024
1,152
1,152
Num. heads
12
16
16
16
Appendix
Table 12: Hyperparameter setup of the generator. Num. layers counts two-stream + single-stream MM-DiT blocks. The last column is the model in Section 4.4 .
Params
GFLOPs / step
Model
Generator
NEPA
Generator
NEPA
TFLOPs / image
NEPA-DiT-B
135M
711M
59
92
–
NEPA-DiT-L
458M
711M
210
92
–
NEPA-DiT-XL, ODE-96
683M
711M
309
92
62
NEPA-DiT-XL, SDE-250
683M
711M
309
92
160
SiT-XL/2 + REPA, SDE-250
675M
–
229
–
91
Appendix
Table 13: Parameters and sampling FLOPs. GFLOPs per sampling step for one image; the NEPA model runs on the 64 image tokens with the condition cached. Totals include guidance on t∈[0.4,1] and VAE decoding. SiT-XL/2 is counted in the same way.
Figure 8: Samples of NEPA-DiT-XL on ImageNet 256×256 .
Figure 9: Samples of NEPA-DiT-XL on ImageNet 256×256 (continued).
Figure 10: Samples of NEPA-DiT-XL on ImageNet 256×256 (continued).