CLIMB-flow: Coupled Linear Inverse posterior sampling via Multiscale-Based flow
Authors: Zeqiu Yu, Ruizhi Yuan, Mathews Jacob
Organizations: Department of Electrical and Computer Engineering, University of Virginia, Charlottesville, VA, USA · Department of Biostatistics and Health Data Science, University of Pittsburgh, Pittsburgh, PA, USA
Diffusion models are now widely used in Bayesian inverse problems in imaging as priors, where latent diffusion models are often used for larger scale problems to keep the computational complexity and model-size manageable. Unfortunately, the auto-encoder based compression results in loss of spatial detail. In addition, the optimization is converted to a non-linear problem. In this paper, we introduce a posterior sampling algorithm customized for the pyramidal/cascaded architecture, which relies on a coarse to fine hierarchical strategy to generate images in the pixel domain. We present CLIMB-Flow which alternates between three steps: an end-point estimation from the current coarse and noisy image, data-consistent update of the clean image, and re-noising it back to the level the network expects. Together these steps sample the posterior at that scale using an approximate Gibbs sampling from two conditional distributions. Experiments on ImageNet, CelebA, AFHQ and fastMRI span inpainting, deblurring, super-resolution and accelerated MRI, with PSNR gains of 1.37-7.66 dB over the strongest competing method on CelebA and pixel-domain reconstruction up to 512x512.
Figures & tables
Figure 1: Preview of CLIMB-Flow reconstructions. Inference proceeds from coarse to fine, with upsampling (U) linking successive scales. At scale k , a coupled target connects the noisy interpolant xτk and clean image x1k : the velocity network supplies prior information through the former, while measurements constrain the latter. Each panel shows one step of the approximate Gibbs sweep, repeated Sit times before advancing to the next scale.
Type
Model
Inpaint (box)
Inpaint (random)
Super resolution 4 ×
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
P;F
Ours (Mode)
22.21
0.834
0.188
86.80
28.96
0.856
0.110
35.70
25.97
0.712
0.351
97.00
P;F
Ours
22.17
0.779
0.179
77.80
28.18
0.795
0.105
32.40
25.31
0.681
0.266
79.49
P;D
DAPS [ 5 ]
21.43
0.725
0.214
109.85
28.44
0.775
0.135
54.25
25.89
0.694
0.276
83.57
P;D
DPS [ 3 ]
18.94
0.722
0.257
126.52
23.52
0.745
0.297
87.53
21.13
0.489
0.361
106.32
P;D
DDRM [ 2 ]
18.63
0.733
0.254
116.37
–
–
–
–
22.62
0.521
0.324
103.85
Table 1: Quantitative results. Left: ImageNet comparisons with pixel- and latent-domain diffusion baselines. Ours (Mode) achieves the highest SSIM on all three tasks, while Ours achieves the lowest LPIPS/FID on both inpainting tasks and super-resolution. Right: CelebA comparisons with pixel-domain flow solvers and DiffPIR. Ours (Mode) leads PSNR/SSIM on all four tasks, with PSNR gains of 1.37 – 7.66 dB over the best competing baseline per task; Ours improves LPIPS over Mode on three tasks. Bold / underline : best/second-best among all methods in each panel, including both CLIMB-Flow variants; ties are marked equally. P/L: pixel/latent domain; D/F: diffusion/flow.
Figure 2: Qualitative results. (a) 4× super-resolution comparisons with LatentDAPS and ReSample illustrate richer textures and lower LPIPS from sampling versus higher PSNR/SSIM from Mode. (b) Reconstructions on AFHQ 512×512 with box inpainting: the top row shows generation diversity on ImageNet samples and the bottom row shows the convergence trajectory.
Figure 3: fastMRI reconstruction. Comparing with 4 baselines under 8× accelerated MRI reconstruction. The MMSE and reconstruction error maps take 16 reconstructions. CLIMB-Flow algorithms achieves the best PSNR and SSIM, while preserving anatomical structure and fine-grained details.
Pretrained diffusion models represent image distributions through a continuum of progressively smoothed distributions. This multiscale structure organizes generation from global structure to fine detail and supports high-quality, diverse samples. We exploit the same multiscale diffusion prior for linear imaging inverse problems. Rather than using the pretrained model only as a denoiser in an outer iteration, we define a surrogate likelihood whose center is aligned with the clean-image coordinate and whose covariance accounts for residual diffusion uncertainty. This construction defines an explicit surrogate posterior path, from which we derive continuous posterior dynamics. A tunable Langevin component supports target tracking and allows the amount of posterior exploration to be adapted to the application. We prove endpoint consistency and a finite-horizon tracking bound and, in the exact-score setting, first-order weak accuracy. For computation, we derive the Posterior-Dynamics Implicit--Explicit sampler (PD-IMEX), a stable method using one score evaluation per diffusion scale and an implicit data-consistency update. Experiments on deblurring, super-resolution, and inpainting show strong reconstruction quality at 100 score evaluations, coarse-grid stability, and controllable fidelity--diversity behavior.
Zhaoqiang Liu, Tongyao Pang, Ruibing Wang +1
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu 611731, China · Yau Mathematical Sciences Center, Shuangqing Complex, Tsinghua University, Beijing 100084, China
Latent Flow Models have revolutionized compressed-space image synthesis, yet their application to high-fidelity inverse problems remains bottlenecked. In this paper, we trace this dilemma to a fundamental geometric limitation of pre-trained autoencoders, which we term \emph{First-Order Manifold Blindness}. Severe decoder compression (e.g., retaining only ∼2% of the original degrees of freedom) produces a rank-deficient Jacobian, rendering high-frequency measurement residuals in its orthogonal complement invisible to latent gradients even when the decoder can represent the target image. To overcome this bottleneck, we propose Hybrid-Domain Posterior Sampling (HDPS), a decoupled inference framework that disentangles physical measurement consistency from semantic prior modeling. HDPS diverges into the pixel space, leveraging Langevin dynamics to absorb precise orthogonal measurement gradients, and subsequently projects these structural corrections back onto the generative manifold. An optimization-based latent alignment is introduced to filter pixel-space artifacts while avoiding the semantic drift of direct encoding. Extensive experiments on diverse inverse problems demonstrate that HDPS establishes a new state-of-the-art, successfully recovering the high-frequency structural precision that latent-only solvers inherently discard. The code is available at https://github.com/74587887/HDPS.
Hongjie Wu, Yiping Xie, Jiancheng Lv
College of Computer Science, Sichuan University Chengdu, China
Variational inference (VI) is a powerful method for principled posterior inference for scientific inverse imaging. VI learns the posterior distribution, often with a flow-based network, which can cheaply generate posterior samples upon optimization, and can flexibly incorporate score-based or classic priors. However, its application to large-scale image reconstruction is severely hindered by the poor scalability of the flow-based networks. In this work, we introduce ShuffleFlow, a scalable VI framework to address this challenge. Our method breaks down the problem into three parts: a pixel-unshuffling-based image coordinate sampler, a neural field as feature encoder, and a conditional normalizing flow (CNF) as posterior estimator. Specifically, our framework partitions an image into a stack of sub-images with pixel-unshuffling and uses a shared CNF to model the joint distribution of the sub-image stack. We condition the CNF on the output of a neural field, which embeds feature vectors corresponding to pixel-unshuffling sample locations to capture spatial structures, and share the flow's latent variable across the channels to model their correlations. We demonstrate our method's effectiveness and efficiency on both linear and nonlinear imaging inverse problems, and show its ability to more rapidly generate a high-sample-count posterior than diffusion samplers.
Tianao Li, Tjitske Starkenburg, Yu Sun +1
Department of Computer Science, Northwestern University, Evanston, IL 60208