Generative models have found great success as data-driven methods of solving inverse problems. Two popular approaches work either by combining a pretrained generative prior with a known degradation model, or by training a conditional generative model directly from paired data. We target a setting that spans both regimes: unknown degradations can be learned from few paired examples, while known degradations can be learned from self-generated samples. We introduce Likelihood Score Approximation (LSA), a generative framework that keeps a pretrained unconditional model fixed and learns an observation-conditioned model that approximates the likelihood score from paired samples. Within a conditional stochastic-interpolant framework, LSA can be trained in either score or velocity coordinates, independently of the unconditional model's native parameterization, and supports both deterministic and stochastic sampling. We further show empirically that the prior model can be swapped post-training while keeping the same LSA model. Across speech and image inverse problems, LSA operates effectively even at roughly 0.01% of the full training dataset. On the ImageNet-256 benchmark it achieves competitive or better restoration quality than strong posterior-sampling baselines while requiring up to several orders of magnitude fewer network evaluations.
Figures & tables
Figure 1: Prior and posterior vector fields along the Gaussian probability path. (a) Velocity vectors at the same intermediate state xt , differing by κtℓt(x,y) . (b) Score vectors for pt(x) and pt(x∣y) , differing by ℓt(x,y) .
Figure 2: Overview of LSA . Training uses either self-generated pairs with a known degradation operator or limited paired-data with an unknown operator. Prior and LSA parameterizations can be chosen independently.
Figure 3: ImageNet- 5128× super-resolution across paired-data budgets, following the CSE benchmark ( Xu et al., 2025 ) . (a) FID for models trained with 25–1k real pairs, together with the self-generated setting. (b) Qualitative results for LSA .
Data
Model
PESQ ↑
SI-SDR ↑
Avg. nMOS ↑
FAD ↓
EARS-WHAM-v2
1 min
Direct Posterior
1.44
8.70
2.60
1.95
ControlNet
1.48
9.66
3.33
0.74
LSA (ours)
1.63
9.80
3.60
0.58
10 min
Direct Posterior
1.70
11.50
3.44
0.25
ControlNet
1.47
10.18
3.43
0.42
Table 1: Speech restoration results. (E), (U), and (VB) denote EARS, Urgent, and VoiceBank priors. ∗ AnyEnhance is trained on 44.1-kHz audio. Bold : best within each data regime.
Task
Method
NFE
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Task
Method
NFE
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Super- resolution ( 4× )
DMAP
∼ 300
25.39
.661
.229
74.65
Gaussian deblurring
DAPS
1k
26.15
.684
.253
75.68
DAPS
1k
25.89
.694
.276
83.57
ControlNet
26
25.08
.668
.242
99.82
ControlNet
26
25.06
.675
.237
92.15
LSA (ours)
21
25.39
.690
.212
69.34
LSA (ours)
23
25.57
.700
.215
70.38
Motion deblurring
RePS
1k
28.95
.801
.169
53.15
Inpainting (box)
MAS
–
21.15
.817
.168
95.96
DAPS †
1k
–
.769
.175
47.09
DAPS
1k
21.43
.725
.214
109.85
LSA (ours)
42
28.18
.802
.130
32.45
Table 2: ImageNet- 256 restoration on 100 images with additive Gaussian noise ( σy=0.05 ), following the DAPS benchmark ( Zhang et al., 2025a ) . Bold : best. † DAPS rerun from DAPS++ ( Chen et al., 2026 ) .
Figure 4: ImageNet- 256 restoration with 128 real pairs. All paired methods are trained on the same fixed 128 pairs, with DEFT ( Denker et al., 2024 ) included where available. For reference, we also show LSA trained on self-generated prior samples paired with observations from the known degradation operator.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Model
Width
Depth config.
Architecture details
Frozen Parameters
GFLOPs/ NFE
ImageNet- 256
PixelDiT-XL ( 320 ep.)
1152
26
4 pixel blocks; patch size 16
797.38M
311.19
ImageNet- 512
PixelDiT-XL
1152
26
4 pixel blocks; patch size 16
797.38M
1352.24
Speech
NCSN++
128
(1,1,2,2,2,2,2)
2 residual blocks
65.56M
132.70
ADM
128
(1,1,2,2,2,2,2)
2 residual blocks
72.25M
131.42
Appendix
Table 3: Prior model configurations used throughout the experiments. All prior parameters remain fixed during downstream training. Speech GFLOPs counts are measured for one second of audio.
Setting
Model
Width
Depth config.
Architecture details
Trainable params
GFLOPs/ NFE
IN256, 0 real pairs
LSA (NCSN++)
40
(1,1,1,2)
2 residual blocks
2.16M
44.30
IN256, 128 pairs
LSA (NCSN++)
32
(1,1,1,2)
1 residual block
1.01M
19.84
IN256, 128 pairs
ControlNet
692
13
13 control blocks
124.92M
46.80
IN256, 128 pairs
LoRA
–
30
rank=8
2.15M
1.10
IN256 ablation
LSA (PixelDiT)
96
4
2 pixel blocks; patch size 4
0.98M
44.71
IN512, 0 real pairs
LSA (NCSN++)
40
(1,1,1,2)
2 residual blocks
2.16M
176.04
Appendix
Table 4: Model configurations used across all ImageNet experiments. Parameters and GFLOPs/NFE refer to the trainable LSA model or adapter.
Setting
Model
Width
Depth config.
Architecture details
Trainable params
GFLOPs/ NFE
1 min
Conditional / LSA
32
(1,2,2)
2 residual blocks
1.71M
27.60
10 min
Conditional / LSA
64
(1,1,2,2)
1 residual block
5.47M
46.51
Full
Conditional / LSA
128
(1,1,2,2,2,2,2)
2 residual blocks
65.59M
266.03
0 real pairs Phase Retrieval
LSA
64
(1,1,2,2)
2 residual blocks
7.60M
66.22
1 / 10 min
ControlNet
32
7/7
–
1.81M
5.55
Full
ControlNet
64
7/7
–
6.66M
19.64
Appendix
Table 5: Speech model configurations. Parameters and GFLOPs/NFE refer to the trainable model shown in each row. Speech compute is measured for one second of audio.
Task
Method
NFE
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Super-resolution ( 4× )
DEFT
100
24.87
.651
.244
97.15
ControlNet
43
24.78
.665
.247
89.14
LSA (ours)
21
24.93
.683
.229
73.01
Inpainting (box)
DEFT
100
19.20
.759
.194
141.19
ControlNet
10
20.24
.685
.275
152.02
LSA (ours)
22
18.94
.755
.193
134.00
Appendix
Table 6: Detailed ImageNet- 256 restoration results using 128 real pairs. We report NFE, PSNR, SSIM, LPIPS, and FID across all evaluated tasks. Bold denotes the best result for each task.
Figure 6: ImageNet-256 super-resolution, inpainting, and HDR restoration results.
Source
Files
NFE ↓
PSNR ↑
LPIPS ↓
FID ↓
DPS+CSE
25
500
22.26
.405
75.64
LSA (ours)
25
12
23.07
.347
40.76
DPS+CSE
50
500
22.84
.353
55.78
LSA (ours)
50
9
23.20
.335
39.34
CSE
100
500
16.54
.484
85.58
DPS+CSE
100
500
22.98
.323
44.59
Appendix
Table 7: ImageNet- 5128× super-resolution across paired-data budgets. LSA is trained with 25, 50, 100, or 1k real pairs; “self-gen” denotes training on self-generated pairs obtained using the known degradation operator. Baseline results are from Xu et al. (2025) . Bold indicates the best available value within each data budget.
Parameterization
Initialization
tstart
NFE
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Score
Observation marginal
.878
41
24.91
.663
.217
72.86
Velocity
Gaussian
.907
49
24.68
.657
.216
70.23
Velocity
Gaussian
1.000
36
24.86
.661
.221
73.15
Data prediction
Observation marginal
.819
31
25.10
.676
.229
75.72
Data prediction
Gaussian
1.000
29
24.49
.655
.233
74.19
Appendix
Table 8: LSA parameterization ablation for Gaussian deblurring on ImageNet- 256 with 128 pairs and an NCSN++ backbone. For each parameterization, we report the validation-selected sampling start time tstart and initialization. For velocity and data prediction, we additionally report default sampling from the Gaussian endpoint at tstart=1 ; this setting is not defined for the score parameterization due to the endpoint singularity.
FM backbone
Initialization
tstart
NFE
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
NCSN++
Gaussian
.907
49
24.68
.657
.216
70.23
NCSN++
Gaussian
1.000
36
24.86
.661
.221
73.15
PixelDiT
Gaussian
.840
22
25.25
.682
.213
68.31
PixelDiT
Gaussian
1.000
19
24.97
.673
.217
71.52
Appendix
Table 9: LSA backbone ablation for Gaussian deblurring on ImageNet- 256 with 128 pairs. We compare NCSN++ and PixelDiT using Gaussian initialization, reporting both the validation-selected sampling start time tstart and sampling from the Gaussian endpoint at tstart=1 . (Appendix B.4 )
Figure 7: Gaussian-deblurring learning curves for LSA on ImageNet- 256 with a PixelDiT prior, trained using either 128 real pairs or self-generated data. Training time is measured on a single A100 GPU.
Figure 8: Qualitative BWE comparison on three EARS test utterances. The shared top rows show the clean reference and low-pass-filtered observation; each subsequent block compares ControlNet, direct posterior training, and LSA (ours) using 1 minute, 10 minutes, or the full paired training set. Spectrogram magnitudes are measured relative to the clean-reference peak of each utterance, with frequency shown in kHz.
Figure 9: Inference-time prior swapping without retraining LSA . Top: ImageNet- 256 Gaussian deblurring with 128 pairs, where the PixelDiT prior used during training (epoch 320) is replaced by earlier checkpoints. Bottom: 1-min speech-enhancement LSA , where the NCSN++ training prior is replaced by an ADM prior trained on the same EARS data. Values show metric changes relative to the training prior; despite the architecture change, performance remains broadly comparable, with a substantial improvement in LSD.
Figure 10: ODE versus SDE sampling for LSA bandwidth extension with 1 min, 10 min, and full paired data. Changes are relative to the ODE baseline ( ρ=0 ); SDE uses ρ=0.18 , 0.09 , and 0.13 , respectively. Mild stochasticity generally improves performance, providing complementary gains.
Restart
ζ1→ζ0
PESQ ↑
SI-SDR ↑
Avg. nMOS ↑
FAD ↓
Yes
3→0
1.51
7.82
3.81
0.35
No
3→0
1.52
7.67
3.47
1.29
Yes
1
1.47
7.43
3.14
0.70
Yes
2
1.51
10.56
2.93
1.82
No
1
1.36
5.96
2.76
1.11
Appendix
Table 10: Inference-refinement ablation for the 1-min LSA speech-enhancement model. We vary the use of restart refinement and the LSA guidance-weight schedule ζ1→ζ0 , while keeping the trained model fixed (see Appendix B.4 ).
Diffusion and flow-based models learn powerful data priors by training a denoiser to reverse Gaussian corruption. To use this prior to solve a linear inverse problem, one needs to sample from the posterior, but the score that the prior provides is the unconditional score, not the posterior score. Existing methods either steer a fixed pretrained denoiser with approximate measurement-matching corrections, or train a conditional restoration model that abandons the denoising structure of the prior. We derive the exact posterior score in closed form for linear Gaussian inverse problems under general Gaussian interpolants, and show that posterior sampling reduces to a denoising problem at an operator-dependent shifted pivot under an anisotropic noise covariance. We turn this identity into Exact Posterior Score (EPS), a denoising training objective that preserves the input/output structure of standard pretraining and can therefore be trained from scratch or fine-tuned from a pretrained denoiser. At inference, EPS uses the same sampler as the underlying backbone, with no likelihood gradients or projections. We evaluate EPS on five linear inverse problems across FFHQ and ImageNet, where it outperforms training-free and training-based baselines on fidelity, perceptual, and distributional metrics, while using roughly an order of magnitude fewer denoiser evaluations than gradient-based posterior samplers.
Score-based diffusion models demonstrate superior performance in generative tasks but encounter fundamental bottlenecks in inverse problems due to the analytical intractability of the time-dependent likelihood score. To bridge this gap, we propose a novel proximal-based generative modeling (PGM) framework that rigorously circumvents explicit likelihood evaluation. Our framework is built upon a theoretical equivalence between Gaussian convolution in diffusion processes and Moreau-Yosida regularization in nonsmooth optimization. This enables a new sampling mechanism driven by the proposed Moreau score, which admits a closed-form expression via proximal operators. Moreover, we introduce Moreau score matching to learn the proximal operators that rely solely on samples drawn from the prior distribution. Theoretically, PGM eliminates the early-stopping bias inherent in the score-based diffusion model and achieves non-asymptotic convergence. Experiments demonstrate that PGM significantly surpasses state-of-the-art methods in reconstruction quality and sampling time.
Boyang Zhang, Zhiguo Wang, Ya-Feng Liu
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, China · Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing, China · School of Mathematics, Sichuan University, Chengdu, China +2
Vision-Language Latent Diffusion Models (LDMs) (Rombach et al., 2022) provide powerful generative priors for inverse problems. However, existing LDM-based inverse solvers typically require a large number of neural function evaluations (NFEs) and backpropagation through large pretrained components, leading to substantial computational costs and, in some cases, degraded reconstruction quality. We propose a unified Euclidean-Wasserstein-2 gradient-flow framework that jointly performs posterior sampling and prompt optimization in the latent space through a single flow that aligns the prior and posterior with the observed data. Combined with few-step latent text-to-image models, this formulation enables low-NFE inference without backpropagation through autoencoders. Experiments across several canonical imaging inverse problems show that our method achieves state-of-the-art performance with significantly reduced computational cost.
Alessio Spagnoletti, Tim Y. J. Wang, Marcelo Pereyra +1
Laboratoire MAP5, UMR 8145, Université Paris Cité, CNRS · Department of Mathematics, Imperial College London · Heriot-Watt University, MACS & Maxwell Institute