There has been a proliferation of sampling algorithms based on Wasserstein gradient flows (WGF) and forward-only diffusion processes (FODP), often accompanied by theoretical guarantees of exponentially fast convergence to the target distribution. These guarantees are frequently interpreted as evidence that such methods can efficiently sample complex multimodal distributions, often supported by empirical results. In this work, we argue that this interpretation is fundamentally misleading. By invoking the Jordan-Kinderlehrer-Otto (JKO) scheme and Otto calculus, we establish that the canonical WGF sampling dynamics and overdamped forward diffusion share the same density evolution and therefore inherit the same metastability and slow-mixing phenomena long understood in nonequilibrium statistical physics. We analyze this family of samplers using two complementary tools -- spectral analysis and mean first-passage time (MFPT) analysis -- and show that well-separated multimodality can induce exponentially long mixing times associated with small spectral gaps and rare inter-mode transitions. For the commonly adopted log-linear annealing schedule studied here, we find that introducing intermediate distributions does not remove the exponential scaling of the total transport time. The limitation is structural rather than implementation-specific: purely local, gradient-driven transport mechanisms can require exponentially long times to transport probability mass across well-separated modes. We argue that this represents a fundamental limitation of WGF- and FODP-based sampling in their standard forms, and motivates future development of fundamentally nonlocal mechanisms for efficient multimodal sampling.
Figures & tables
Figure 1: Mixing times predicted by the spectral (discrete markers) and MFPT (continuous line) analyses. Spectral: ∣ℜ(λ1)∣−1 ; MFPT: Eq. ( 13 ) with Eq. ( 46 ). Insets illustrate the probability densities of the bimodal Gaussian mixtures.
Figure 2: Performance comparison of different WGF-based samplers and i.i.d. samples of the bimodal Gaussian mixture in D=1 .
Figure 3: Mixing times when Na−1 intermediate scales are used in annealing, measured by the description in Sec. 3.3 . Na=1 corresponds to no annealing. The complementary estimation using λ1 presented in Sec. 2.2 is shown for comparison, showing an accurate estimation of the scaling of mixing times for 2d≳0.87 .
Figure 4: Performance comparison of FODP with different terminal time T and i.i.d. samples of the bimodal Gaussian mixture in D=1 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
VWGF quantity
Symbol / Name
Value
Training particle count
Ntrain
512000
Evaluation particle count
Neval
100
Batch size
–
512
Number of training epochs
–
40
Learning rate
–
0.001
Network architecture — layers
–
2 (D=1), 3 (D=10)
Appendix
Table 1: VWGF settings used for the appendix comparison.
Figure 5: Performance comparison of different WGF-based samplers and i.i.d. samples of the bimodal Gaussian mixture in D=10 .