There has been a proliferation of sampling algorithms based on Wasserstein gradient flows (WGF) and forward-only diffusion processes (FODP), often accompanied by theoretical guarantees of exponentially fast convergence to the target distribution. These guarantees are frequently interpreted as evidence that such methods can efficiently sample complex multimodal distributions, often supported by empirical results. In this work, we argue that this interpretation is fundamentally misleading. By invoking the Jordan-Kinderlehrer-Otto (JKO) scheme and Otto calculus, we establish that the canonical WGF sampling dynamics and overdamped forward diffusion share the same density evolution and therefore inherit the same metastability and slow-mixing phenomena long understood in nonequilibrium statistical physics. We analyze this family of samplers using two complementary tools -- spectral analysis and mean first-passage time (MFPT) analysis -- and show that well-separated multimodality can induce exponentially long mixing times associated with small spectral gaps and rare inter-mode transitions. For the commonly adopted log-linear annealing schedule studied here, we find that introducing intermediate distributions does not remove the exponential scaling of the total transport time. The limitation is structural rather than implementation-specific: purely local, gradient-driven transport mechanisms can require exponentially long times to transport probability mass across well-separated modes. We argue that this represents a fundamental limitation of WGF- and FODP-based sampling in their standard forms, and motivates future development of fundamentally nonlocal mechanisms for efficient multimodal sampling.
Figures & tables
Figure 1: Mixing times predicted by the spectral (discrete markers) and MFPT (continuous line) analyses. Spectral: ∣ℜ(λ1)∣−1 ; MFPT: Eq. ( 13 ) with Eq. ( 46 ). Insets illustrate the probability densities of the bimodal Gaussian mixtures.
Figure 2: Performance comparison of different WGF-based samplers and i.i.d. samples of the bimodal Gaussian mixture in D=1 .
Figure 3: Mixing times when Na−1 intermediate scales are used in annealing, measured by the description in Sec. 3.3 . Na=1 corresponds to no annealing. The complementary estimation using λ1 presented in Sec. 2.2 is shown for comparison, showing an accurate estimation of the scaling of mixing times for 2d≳0.87 .
Figure 4: Performance comparison of FODP with different terminal time T and i.i.d. samples of the bimodal Gaussian mixture in D=1 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
VWGF quantity
Symbol / Name
Value
Training particle count
Ntrain
512000
Evaluation particle count
Neval
100
Batch size
–
512
Number of training epochs
–
40
Learning rate
–
0.001
Network architecture — layers
–
2 (D=1), 3 (D=10)
Appendix
Table 1: VWGF settings used for the appendix comparison.
Figure 5: Performance comparison of different WGF-based samplers and i.i.d. samples of the bimodal Gaussian mixture in D=10 .
Sampling from unnormalized multimodal distributions with limited density evaluations remains a fundamental challenge in machine learning and natural sciences. Successful approaches construct a bridge between a tractable reference and the target distribution. Parallel Tempering (PT) serves as the gold standard, while recent diffusion-based approaches offer a continuous alternative at the cost of neural training. In this work, we introduce Conditional Diffusion Sampling (CDS), a framework that combines these two paradigms. To this end, we derive Conditional Interpolants, a class of stochastic processes whose transport dynamics are governed by an exact, closed-form stochastic differential equation (SDE), requiring no neural approximation. Although these dynamics require sampling from a non-trivial initialization distribution, we show both theoretically and empirically that the cost of this initialization diminishes for sufficiently short diffusion times. CDS leverages this by a two-stage procedure: (1) PT is used to efficiently sample the initial distribution, and then (2) samples are transported via the transport SDE. This combination couples the robust global exploration of PT with efficient local transport. Experiments suggest that CDS has the potential to achieve a superior trade-off between sample quality and density evaluation cost compared to state-of-the-art samplers.
Francisco M. Castro-Macías, Pablo Morales-Álvarez, Saifuddin Syed +3
University of Granada · Work done while visiting the University of Cambridge. · University of British Columbia +2
We consider the problem of sampling from a probability distribution π. It is well known that this can be written as an optimisation problem over the space of probability distributions in which we aim to minimise the Kullback--Leibler divergence from π. We consider the effect of replacing π with a sequence of moving targets (πt)t≥0 defined via geometric tempering on the Wasserstein and Fisher--Rao gradient flows. We show that convergence occurs exponentially in continuous time, providing novel bounds in both cases. We also consider popular time discretisations and explore their convergence properties. We show that in the Fisher--Rao case, replacing the target distribution with a geometric mixture of initial and target distribution never leads to a convergence speed up both in continuous time and in discrete time. Finally, we explore the gradient flow structure of tempered dynamics and derive novel adaptive tempering schedules.
Francesca Romana Crucinio, Sahani Pathiraja
ESOMAS, University of Turin, Italy · Collegio Carlo Alberto, Turin, Italy · School of Mathematics & Statistics, UNSW Sydney, Australia
We study the convergence of Wasserstein-Fisher-Rao (WFR) gradient flows for sampling from probability distributions known up to a normalisation constant. By combining Wasserstein transport with Fisher-Rao birth-death dynamics, WFR flows balance exploration and selection. These flows have been recognised as a promising mechanism to accelerate convergence beyond Langevin dynamics. We show that for a class of strongly log-concave target distributions satisfying additional curvature conditions, WFR flows preserve strong log-concavity, in contrast to Wasserstein flows which enjoy this property only in the Gaussian setting. Exploiting this result, we derive explicit non-asymptotic convergence rates for the symmetrised Kullback-Leibler divergence, without requiring a warm-start as required in current estimates. In particular, we show that the convergence rate decomposes additively into Wasserstein and Fisher-Rao contributions, thereby confirming a recent conjecture within this setting. These results provide refined convergence guarantees and further develop the theoretical foundations of WFR gradient flows for sampling and Bayesian inference.
Francesca Romana Crucinio, Sahani Pathiraja
ESOMAS, University of Turin, Italy & Collegio Carlo Alberto, Turin, Italy · School of Mathematics & Statistics, UNSW Sydney, Australia