Some diffusion posterior samplers construct Gaussian-tilted intermediate distributions along the reverse process. We observe that these targets can be pulled back to clean-space posteriors with weaker conditioning, with samples transported analytically to the corresponding noisy-space target through a Gaussian bridge. For the sequential Monte Carlo (SMC) sampler MCGDiff, the effective observation variance of this pulled-back problem is up to twice the diffusion-noise variance. We exploit this structure to initialize MCGDiff directly at an intermediate time: an approximate solver samples the softened clean-space posterior, the Gaussian bridge maps these samples to the tilted target, and only the remaining SMC suffix is run. This trades asymptotic consistency for finite-particle performance. With moment-matching posterior sampling (MMPS) as the solver, the hybrid improves sliced Wasserstein distance by roughly 2× at matched particle count on a structured Gaussian-mixture inverse problem, and by more than an order of magnitude when the posterior-relevant mode is rare under the prior. A prior-initialization control, which retains the bridge but drops the clean-space conditioning, shows that on MCGDiff's standard Gaussian-mixture benchmark most of the improvement is insensitive to the conditioning. Conditioning the initialization gives a further consistent gain on the structured problem, and becomes decisive on a rare-mode problem, where resampling cannot repopulate a mode absent from the initial population.
Figures & tables
Figure 1: Effect of initialization time t⋆ on the standard GMM benchmark. (a) Final posterior SWD for oracle and MMPS initialization. (b) MMPS approximation error before and after the exact bridge to vt⋆ . Generated with 1000 MMPS denoising steps and averaged over 80 seeds.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Posterior samples on the standard GMM benchmark for MCGDiff, direct MMPS, MMPS initialization, and prior initialization. MMPS and Prior initialization give qualitatively comparable samples on this relatively simple benchmark. Teal: exact posterior samples; Red: approximate method samples.
Figure 3: Posterior samples on the rare-mode GMM benchmark for MCGDiff, direct MMPS, MMPS initialization, and two separate prior initialization examples. Unless marked otherwise, initialized methods use t⋆=0.3 and ESS-thresholded resampling at 0.7N . Prior initialization ⋆ instead uses t⋆=0.4 with resampling at every step. This variant illustrates a failure mode of prior-based initialization: it collapses to the wrong posterior mode, whereas the ESS-thresholded prior initialization recovers the correct mode but remains less accurate than MMPS initialization. Teal: exact posterior samples; Red: approximate method samples.
Method
N
t⋆
GMM
Structured
Rare ( ρ=0.002 )
Rare ( ρ=0.01 )
Baselines
MCGDiff
1k
–
2.307±0.45
1.271±0.20
5.818±0.40
3.052±0.67
MCGDiff (ESS 0.7)
1k
–
1.856±0.40
0.762±0.12
2.836±0.66
0.280±0.14
MCGDiff
2k
–
1.897±0.31
0.838±0.12
–
–
MCGDiff (ESS 0.7)
2k
–
1.519±0.31
0.548±0.09
–
–
MCGDiff
16k
–
1.354±0.23
0.382±0.05
–
–
Appendix
Table 1: Sliced Wasserstein distance (lower is better) for the GMM benchmarks using MMPS initialization with 100 denoising steps. Values are mean ± 95% confidence interval over 80 random seeds. Initialization methods use the ESS <0.7 MCGDiff suffix. Column N gives the number of MCGDiff particles used. Among approximate initialization methods, the best result is bold and the second best is underlined. “Exact posterior floor” is the SWD calculated from comparing two batches of 1000 samples of exact posterior samples. Oracle init values for t⋆=1 are included for higher particles N for the standard GMM benchmark to provide a fairer baseline, removing the bias due to the benchmark’s inaccurate terminal initialization (see Appendix A.3 ).
Method
t⋆
GMM
Structured
Baselines
MCGDiff (1k)
–
2.615±0.46
1.271±0.20
MCGDiff (1k, ESS 0.7)
–
1.784±0.29
0.762±0.12
Oracle init
0.25
0.444±0.07
0.320±0.06
0.30
0.466±0.08
0.322±0.04
Appendix
Table 2: Sliced Wasserstein distance (lower is better) for the GMM benchmarks using MMPS with 1000 denoising steps. Values are mean ± 95% confidence interval over 80 random seeds. Methods use the ESS <0.7 MCGDiff suffix unless marked otherwise. The number of MCGDiff particles used for all relevant methods is 1,000. Among approximate methods, the best result is bold and the second best is underlined. Data in this table generated with different random seeds to corresponding data in Figure 1 .
Pretrained diffusion models are powerful priors for inverse problems, but posterior sampling under nonlinear, non-differentiable forward models remain hard. We introduce diffusion waltz, an MCMC method using SDEdit-style noising-denoising as a proposal, corrected via Metropolis-Hastings for exact posterior sampling without prior evaluation. We further propose injecting observations into the proposal while preserving exactness, using a gradient-free ensemble Kalman update. On a non-differentiable Navier-Stokes initial condition recovery task, diffusion waltz outperforms existing baselines across different noise and nonlinearity regimes.
Sequential Monte Carlo (SMC) methods are a natural tool for post-hoc conditioning of pretrained generative models, but in many applications the mutation kernels used by the particle system are biased approximations of an ideal Feynman--Kac flow. This paper develops a non-asymptotic error analysis for such SMC samplers. Under forward-smoothing forgetting conditions, we decompose the total error into a kernel bias, measuring the effect of replacing the ideal transition kernels by approximate ones, and a finite-particle Monte Carlo error. Our approach relies on extending local Doeblin-type conditions and Lyapunov drift arguments for Markov kernels to conditional distributions, thereby enabling a principled control of the bias. We then instantiate this general framework for conditional sampling with score-based diffusion models, and derive the first non-asymptotic error bound that jointly controls initialization error, time discretization, and score approximation in the reverse diffusion dynamics as well as finite-particle Monte Carlo error.
Stanislas Strasman, Gabriel Victorino Cardoso, Sylvain Le Corff +2
SU, LPSM · Sorbonne Université and Université Paris Cité, CNRS, LPSM, F-75005 Paris, France · LPSM +2
Training-free conditional diffusion provides a flexible alternative to task-specific conditional model training, but existing samplers often allocate computation inefficiently: independent guided trajectories can vary widely in quality, and additional function evaluations along a single trajectory may not recover from poor early decisions. We propose Tempered Guided Diffusion (TGD), an annealed sequential Monte Carlo framework for training-free conditional sampling with diffusion priors. TGD targets tempered posterior distributions over the clean signal, using noisy diffusion states only as auxiliary variables for proposing reconstructions and propagating particles. Particles are reweighted by incremental likelihood ratios, resampled, and propagated across noise levels, concentrating computation on trajectories plausible under both the prior and observation. Under idealized exact-reconstruction assumptions, full TGD yields a consistent particle approximation to the posterior as the number of particles grows. For expensive reconstruction tasks, Accelerated TGD (A-TGD) retains early particle exploration but prunes to a single high-likelihood trajectory partway through sampling. Experiments on a controlled two-dimensional inverse problem and image inverse problems show improved posterior approximation and favorable wall-clock speed-quality tradeoffs over independent multi-trajectory baselines.
Andreas Makris, Paul Fearnhead, Chris Nemeth
Department of Mathematics and Statistics Lancaster University, UK · Department of Mathematics and Statistics2026 Lancaster University, UK