cs.CVOct 1, 2026

Embedding Prediction Helps Image Generation

Authors: Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu

Organizations: University of Michigan · Carnegie Mellon University

Abstract

In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet 256×256256\times256 study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Improved Distributional Diffusion Models

    Sep 29, 2026Tommaso Martorella, Alexandre Galashov, Felix Krause +4Iterative Denoising Methods

  2. DiffusionBench: On Holistic Evaluation of Diffusion Transformers

    Jun 23, 2026Xingjian Leng, Jaskirat Singh, Zhanhao Liang +5Diffusion TransformersDiffusion Models

  3. Denoising Diffusion Generative Models Secretly Calculate Attentions

    Sep 1, 2026Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3Denoising Diffusion Probabilistic ModelImage Generation