math.PROct 26, 2025

Beyond the Semicircle: Free Diffusion Models with Prescribed Equilibria

Authors: Swagatam Das

Abstract

A growing class of machine-learning objects -- covariance and Gram matrices, kernel and attention matrices, MIMO channel matrices, density operators -- are naturally spectra rather than coordinate vectors. Building a denoising diffusion model for such data by corrupting eigenvalues coordinatewise is not merely elegant: it converges to the wrong limit, because eigenvalues repel rather than move independently. Free probability theory supplies the correct forward process, with free convolution replacing classical convolution and Voiculescu's conjugate variable replacing the score, but constant-coefficient free diffusions have a hidden limitation of their own: their only possible equilibrium is the semicircular law, whatever the target distribution looks like. We show that state-dependent free volatility removes this restriction. For any sufficiently regular compactly supported target law, we explicitly construct, through a closed-form Hilbert-transform drift, a free Fokker--Planck flow having that law as a stationary spectral distribution. Under an additional convex-potential condition, the construction admits a globally relaxing diffusion interpretation, reverse-time dynamics, and a trainable score. The generalization is genuine rather than cosmetic: no change of spectral variable reduces it to the constant-coefficient case, and it is a Wasserstein gradient flow only for an explicitly characterized family of coefficients that contains no bounded non-constant member. We instantiate the construction on the Marchenko--Pastur law, the limiting spectrum of isotropic sample covariance matrices, and validate it numerically: the free-versus-coordinatewise gap, an exact dimension-independent score driving reverse-time matrix dynamics at several matrix sizes, the designed Marchenko--Pastur equilibrium, and end-to-end generation with a learned score.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 17, 2026cs.LG

Vocabulary-size-independent Convergence of Discrete Diffusion Models: adjoint equations induce the right space

Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology. Existing convergence theory, however, exhibits fundamental limitations. KL-based analyses diverge under singular priors such as the masked distribution, while bounds in total variation (TV) depend on the vocabulary size SS and become vacuous for modern language tasks, where vocabularies contain hundreds of thousands of tokens. We develop a unified adjoint-equation-based framework that establishes vocabulary-size-independent convergence guarantees in any integral probability metric (IPM). To the best of our knowledge, our bounds are the first to be entirely free of SS and applicable to both masked and uniform priors. Importantly, our results can extend existing step complexity guarantees to any IPM. Also, our theory relies only on a single standard rate-matrix regularity assumption and applies to general priors. Five novel techniques drive our improvements: 1. working in the space of observables via adjoint equations rather than directly with probability measures; 2. a regularity analysis that yields bounds on any IPM; 3. a coupling argument that removes SS-dependence under uniform transitions; and 4. score-marginal cancellation and 5. exit-routing techniques that remove SS-dependence under masked transitions. Our framework thus sharply departs from prior analyses and avoids the shortcomings of pathspace-KL and existing TV-based approaches. Beyond convergence bounds, our framework provides a versatile toolkit for further theoretical study of discrete diffusion models, including principled choices of loss functions and vocabulary-size-independent step complexity.
Aug 4, 2026cs.LG

Simulation-free and finite-time diffusion model

The performance of generative diffusion models is determined by the choice of the reference diffusion process connecting the empirical and prior distributions. Conventional approaches typically trade off simulation-free training against finite-time generation. We propose a framework for designing the reference process that achieves both simultaneously. The key idea is to prescribe tractable time-dependent conditional distributions and then construct the reference process realizing them as its marginals. This framework reveals that score matching is not fundamental to diffusion-model training but instead emerges naturally through reversal of the reference process. We further show that conditional flow matching arises as the small-noise limit of the proposed framework.
May 7, 2026cs.LG

A Unified Measure-Theoretic View of Diffusion, Score-Based, and Flow Matching Generative Models

We survey continuous-time generative modeling methods based on transporting a simple reference distribution to a data distribution via stochastic or deterministic dynamics. We present a unified framework in which diffusion models, score-based generative models, and flow matching are instances of learning a time-dependent vector field that induces a family of marginals (ρt)t∈[0,1](ρ_t)_{t \in [0,1]} governed by continuity and Fokker-Planck equations. Such a unified theory is timely because these methods are converging methodologically, yet fragmented notation and competing derivations continue to obscure their shared structure and the practical tradeoffs governing sampling, stability, and computation. Within this framework, we (i) derive reverse-time sampling for diffusion and score-based models as controlled stochastic dynamics, (ii) show that the probability flow ODE yields identical marginals and connects diffusion to likelihood-based normalizing flows, and (iii) interpret flow matching as direct regression of the velocity field under a chosen interpolation, clarifying when it coincides with or differs from score-based training. We compare objectives, sampling schemes, and discretization errors under unified notation, discuss connections to Schrodinger bridges and entropic optimal transport, and summarize theoretical guarantees and open problems on approximation, stability, and scalability.