cs.LGOct 4, 2026

Diffusion Transformers are Provably Optimal In-context Generators

Authors: Guoji Fu, Tomoya Wakayama, Ryotaro Kawata, Atsushi Nitanda, Wee Sun Lee, Taiji Suzuki

Organizations: National University of Singapore · RIKEN AIP · The University of Tokyo · A*STAR CFAR Nanyang Technological University

Abstract

Generative foundation models are attracting interest for their ability to produce desired outputs from demonstrations given at inference time, without updating parameters. However, since a few demonstrations cannot uniquely identify the intended task, the challenge is how to learn and sample from an output distribution that reflects this task uncertainty. In this work, we theoretically analyze how a Diffusion Transformer (DiT), pretrained across diverse tasks, learns and generates predictive distributions for a new query from demonstrations. We first show that the natural target to generate from finite demonstrations is not an output derived from estimating a single task, but rather a predictive distribution that captures the task uncertainty remaining after observing the demonstrations. We then prove that a DiT can learn this predictive distribution through score estimation, using attention to aggregate information from demonstrations and diffusion to generate samples. Owing to this property, with sufficient pretraining resources and diffusion sampling steps, the resulting DiT achieves the minimax optimal rate over a Hölder class of test-time tasks. These results imply that DiT acts as a statistically grounded in-context generator capable of generating distributions adapted to new tasks while retaining the uncertainty inherent in finite demonstrations.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling

    Date pendingArman Adibi, Alireza Jafari, Mohammad Ghavamzadeh +1Accelerated SamplingTransformer Architectures

  2. Learning to Read the Contextual Tokens in Diffusion Transformers

    Oct 5, 2026Omer Dahary, Etai Sella, Hadar Averbuch-Elor +2Diffusion Transformers

  3. Embedding Prediction Helps Image Generation

    Oct 1, 2026Sihan Xu, Ji Xie, Zilin Wang +2Diffusion TransformersImage Generation