Trajectory Stitching for Solving Inverse Problems with Flow-Based Models
Authors: Alexander Denker, Zeljko Kereta, Carola-Bibiane Schönlieb, Moshe Eliasof
Organizations: Department of Computer Science, University College London · Department of Applied Mathematics and Theoretical Physics, University of Cambridge.
Flow-based generative models have emerged as powerful priors for solving inverse problems. One option is to directly optimize the initial latent code (noise), such that the flow output solves the inverse problem. However, this requires backpropagating through the entire generative trajectory, incurring high memory costs and numerical instability. We propose MS-Flow, which represents the trajectory as a sequence of intermediate latent states rather than a single initial code. By enforcing the flow dynamics locally and coupling segments through trajectory-matching penalties, MS-Flow alternates between updating intermediate latent states and enforcing consistency with observed data. This reduces memory consumption while improving reconstruction quality. We demonstrate the effectiveness of MS-Flow over existing methods on image recovery and inverse problems, including inpainting, super-resolution, and computed tomography.
Figures & tables
Figure 1 : Peak GPU memory for D-Flow vs MS-Flow (Ours) using Euler discretization of the ODE on CelebA. D-Flow scales linearly with the number of timesteps, while MS-Flow is constant.
x∗i+1=argminx∗Φ(x∗)+2α∥x∗−xKi+1∥2.
Algorithm 1 Alternating Minimization with Coordinate Descent Trajectory Optimization of the MS-Flow Objective
Method
Time per iteration
Peak memory
D-Flow (single shooting)
O(nt(Cfwd+Cvjp))
O(ntMnet)
D-Flow (adjoint)
O(nt(Cfwd+Cvjp))
O(Mnet)
MS-Flow (exact gradients)
O(LK(Cfwd+Cvjp))
O(Mnet)
MS-Flow (Jacobian-free)
O(LKCfwd)
O(Mnet)
Table 1 : Time and memory complexity comparison between D-Flow and MS-Flow. Here n is the state dimension, nt the number of ODE discretization steps, K the number of shooting intervals (with K=nt for explicit Euler), L the number of inner iterations, Cfwd and Cvjp the costs of one forward and backward pass through vθ , and Mnet the peak activation memory of the network.
Figure 2 : Evaluation on the OrganCMNIST for sparse-angle CT. Left: The effect of regularization terms. Right: Comparison of MS-Flow with D-Flow across temporal resolutions. Shaded areas represent the standard deviation over the 10 images.
Figure 3 : Convergence of the trajectory loss in Equation 11 for coordinate descent (with and without the Jacobian-free approximation) and full gradient descent.
Method
nt=3
nt=6
nt=12
D-Flow
7.43
17.06
175.57
MS-Flow w/ JFB ( L=1 )
2.44
2.63
4.44
MS-Flow w/ JFB ( L=10 )
18.65
22.00
40.10
MS-Flow w/o JFB ( L=1 )
9.30
19.61
41.05
MS-Flow w/o JFB ( L=10 )
86.65
190.70
405.47
MS-Flow w/ GD ( L=1 )
4.35
5.36
9.42
Table 2 : Average computation time (seconds per image) over 100 iterations for different temporal discretizations ( nt ), averaged over 5 images and 10 runs per image.
Figure 4 : Comparison of the reconstructions for the three image recovery tasks on CelebA. First row: Gaussian deblurring. Second row: Inpainting. Third row: Super-resolution.
Gaussian Deblurring
Box-Inpainting
2x Super-Resolution
4x Super-Resolution
Time/img
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
[s]
Flow-Prior
36.28
0.965
0.009
35.48
0.981
0.011
34.34
0.948
0.015
27.27
0.796
0.094
24.30
OT-ODE
37.72
0.971
0.015
33.44
0.972
0.013
35.14
0.960
0.008
28.18
0.834
0.052
5.19
PnP-Flow
36.24
0.965
0.049
35.87
0.983
0.011
34.64
0.950
0.059
28.28
0.840
0.168
5.97
D-Flow
36.07
0.957
0.029
34.20
0.967
0.012
35.49
0.951
0.019
29.93
0.873
0.062
26.31
MS-Flow (Ours)
39.04
0.968
0.009
34.77
0.974
0.015
36.41
0.957
0.023
28.52
0.838
0.091
14.11
Table 3 : Quantitative comparison of different reconstruction methods across three image recovery tasks. Best results per task and metric (PSNR, SSIM, LPIPS) are highlighted in bold , and second-best are underlined .
Figure 5 : Reconstruction for Gaussian deblurring on FFHQ with σnoise=0.01 .
σnoise=0.01
σnoise=0.1
Method
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
Degraded Image
26.32
0.763
0.388
23.20
0.335
0.684
FlowDPS
27.78
0.751
0.349
26.87
0.697
0.366
FLAIR
29.18
0.783
0.199
28.93
0.775
0.200
MS-Flow (Ours)
30.65
0.837
0.268
28.62
0.772
0.403
Table 4 : Comparison for Gaussian deblurring on FFHQ for different noise levels. Best results in bold, and second-best are underlined.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Example reconstructions on OrganCMnist. First row: filtered back-projection (baseline reconstruction). Second row: MS-Flow (our) reconstruction. Third row: Ground truth.
Image restoration tasks
Method
CT
Super-res
Gaussian Debl.
Box-inpaint
Flow-Priors
–
(0.01, 1e4)
(0.01, 1e4)
(0.01, 1e4)
OT-ODE
–
(0.2, constant)
(0.4, t )
(0.1, t )
PnP-Flow
–
(0.01)
(0.01)
(0.01)
D-Flow
(3, 0.05)
(3, 0.05)
(3, 1)
(3, 1)
MS-Flow
(6, 1, 0.01, 0.1, 5, 0.0001)
(6, 5, 0.01, 0.1, 5, 0.0001)
(6, 1, 0.01, 0.1, 5, 0.0001)
(12, 1, 0.01, 1, 5, 0.000)
Appendix
Table 5 : Hyperparameter settings across image restoration tasks. For Flow-Prior we give (η,λ) , for OT-ODE (t0,γ) , for PnP-Flow (α) , for D-Flow (nt,λ) and for MS-Flow (nt,L,γ,α,η,λ)
PSNR ↑
SSIM ↑
Task
Method
Converged
Best
Converged
Best
Box Inpainting
D-Flow
33.96
35.84
0.9792
0.9736
MS-Flow
36.30
36.46
0.9799
0.9800
Super-resolution
D-Flow
34.67
34.69
0.9314
0.9322
MS-Flow
36.46
36.46
0.9570
0.9570
Gaussian Deblurring
D-Flow
33.48
34.66
0.8696
0.9102
Appendix
Table 6 : Reconstruction quality comparison for CelebA: D-Flow vs. MS-Flow. We show the best PSNR over iterations and the PSNR of the converged image.
Figure 7 : Top: initial shooting-point initialization. Bottom: converged shooting points after optimization. Results shown for Gaussian deblurring, alongside the measurements and ground-truth image.
Figure 8 : Top: initial shooting-point initialization. Bottom: converged shooting points after optimization. Results shown for super-resolution, alongside the measurements and ground-truth image.
Task
Method
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
Low noise ( p=0.05 )
D-Flow
33.26
0.927
0.042
MS-Flow
37.33
0.957
0.013
High noise ( p=0.1 )
D-Flow
28.84
0.846
0.074
MS-Flow
30.23
0.857
0.094
Appendix
Table 7 : Reconstruction quality for salt-and-pepper noise, with Gaussian deblurring on CelebA comparing D-Flow and MS-Flow.
Method
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
D-Flow
34.37
0.968
0.015
MS-Flow ( L=1 iteration)
34.68
0.974
0.015
MS-Flow ( L=5 iterations)
35.62
0.976
0.013
Appendix
Table 8 : Reconstruction quality for box inpainting on CelebA comparing D-Flow and MS-Flow, with respect to the number of inner iterations
Generative models based on flow matching have emerged as a powerful paradigm for inverse problems, offering straighter trajectories and faster sampling compared to diffusion models. However, existing approaches often necessitate differentiating through unrolled paths, leading to numerical instability and prohibitive computational overhead. To address this, we propose P-Flow, a framework that stabilizes the reconstruction process by leveraging a proxy gradient to update the source point. This approach effectively circumvents the numerical instability and memory overhead of long-chain differentiation. To ensure consistency with the prior distribution, we employ a Gaussian spherical projection motivated by the concentration of measure phenomenon in high-dimensional spaces. We further provide a theoretical analysis for P-Flow based on Bayesian theory and Lipschitz continuity. Experiments across diverse restoration tasks demonstrate that P-Flow delivers competitive performance, especially under extreme degradations such as severely ill-posed conditions and high measurement noise.
Flow-based generative models have emerged as powerful image priors for training-free inverse problem solving, capturing coherent semantics and fine-grained structure. Despite these strengths, existing flow-based inverse solvers primarily focus on the design of individual updates, largely overlooking spatio-temporal information allocation under a fixed number of function evaluations (NFEs). Temporally, insufficient early exploration can trap the flow trajectory in an incorrect semantic basin, whereas excessive allocation of NFEs to early stages leaves little budget for late-stage refinement. Spatially, data consistency provides direct constraints only within observed regions, whereas the recovery of missing regions relies mainly on the generative prior. To address these two issues, we introduce two complementary and training-free components, i.e., Spectrum-Adaptive Scheduling (SAS) and Measurement-Prioritized Attention (MPA). For temporal allocation, SAS distributes the available NFEs over flow time according to the degradation spectrum and logSNR geometry, thus better balancing semantic exploration and detail refinement. For spatial propagation, MPA exploits data-prior conflicts to guide information toward weakly constrained regions, thereby enhancing semantic and structural fidelity. Extensive experiments on standard image inverse problems, e.g., super-resolution, motion deblurring, and inpainting, demonstrate that the proposed components can be integrated into existing flow-based inverse solvers in a plug-and-play manner without retraining or additional flow-model evaluations, and can also significantly improve the restoration quality of existing solvers.
We propose NullFlow, a principled framework for one-step generative image reconstruction. Our key idea is to confine the generative flow to a measurement-consistent subspace. Because the flow never leaves this subspace, NullFlow needs no separate data-fidelity corrections, unlike existing solvers. NullFlow samples in a single network evaluation by learning the flow's average velocity, avoiding the step-by-step integration of traditional flow matching methods. We prove that the average velocity of this constrained flow yields a training objective whose global minimizer is a one-step posterior sampler. We show on image inpainting that NullFlow matches state-of-the-art diffusion solvers while cutting inference from hundreds of network evaluations to one.
Xiao Shi, Edward P. Chandler, Chicago Y. Park +2
Computational Imaging Group (CIG) · Washington University in St. Louis