stat.MLOct 6, 2026

Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size

Authors: Dongsun Yoon, Saptarshi Chakraborty

Organizations: Department of Statistics, University of Michigan

Abstract

Discrete diffusion models have emerged as a practically successful framework for generative modeling on discrete product spaces, yet their statistical generalization properties remain poorly understood. Discrete real-world data such as text or biological sequences often concentrate on a small fraction of the astronomically large ambient space because of semantic or physical constraints, but existing bounds fail to capture this distributional structure and instead scale with the size of the ambient space, giving rise to almost vacuous error bounds. We address this gap for uniform discrete diffusion, one of the two dominant discrete diffusion paradigms alongside masking diffusion, by deriving statistical guarantees governed by the effective support size sn(P0)s_n(P_0), a sample-size-dependent measure of distributional complexity. Given nn independent and identically distributed (i.i.d.) samples from an unknown data distribution P0P_0 on [K]d[K]^d, we show that, with appropriate choices of network size and hyperparameters, the expected total variation (TV) loss scales as O(sn(P0)/n)O(\sqrt{s_n(P_0)/n}), while the expected Kullback--Leibler (KL) divergence is bounded by O(1nsn(P0)log⁡(eKd/sn(P0))log⁡n)O(\frac{1}{n}s_n(P_0)\log(eK^d/s_n(P_0))\log n). Furthermore, we show that the TV rate is minimax optimal and that the KL rate is minimax optimal up to a factor of log⁡n\log n. Together, these upper and lower bounds show that uniform discrete diffusion successfully avoids the curse of dimensionality for distributions with small effective support size: the TV error rate depends on the ambient state-space size only through sn(P0)s_n(P_0), while the corresponding KL rate incurs only an additional logarithmic dependence on the ambient state-space size.

Figures & tables

Explore similar work

CardsList
  1. Efficient Sampling with Discrete Diffusion Models: Sharp and Adaptive Guarantees

    Feb 16, 2026Daniil Dmitriev, Zhihan Huang, Yuting WeiScore-Based Diffusion ModelDiffusion Dynamics

  2. Vocabulary-size-independent Convergence of Discrete Diffusion Models: adjoint equations induce the right space

    May 17, 2026Kelvin Kan, Xingjian Li, Benjamin J. Zhang +3Diffusion ModelsGenerative Models

  3. When Diffusion Model Can Ignore Dimension: An Entropy-Based Theory

    May 8, 2026Ahmad Aghapour, Erhan BayraktarDiffusion SamplingHigh-Dimensional