Organizations: Laboratoire MAP5, UMR 8145, Université Paris Cité, CNRS · Heriot-Watt University, School of Mathematical and Computer Sciences & Maxwell Institute for Mathematical Sciences
Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. Current practice falls into two camps. Large foundation image models are deployed as plug-and-play priors with zero-shot approximate likelihood guidance, which introduces significant bias and computational cost. Physics-aware generative models avoid this bias, but each is tied to a specific dataset, task and instrument. We introduce BAM (Bayesian Anything Model), a lightweight foundation model for few-step, physics-aware posterior sampling that generalises robustly to unseen data and tasks, zero-shot or with minimal finetuning. BAM upgrades the operator-conditioned Reconstruct Anything Model (RAM) backbone (Terris et al.) into a conditional flow map, so instrument physics is specified at inference time rather than fixed during training. BAM has just 36M parameters and is pre-trained jointly on large image corpora and libraries of forward operators. A single network then draws posterior samples in a few steps, with no likelihood approximation and no guidance weights to tune. Across linear inverse problems on FFHQ, AFHQ, LSUN, DIV2K and the Kohler camera-shake benchmark, BAM outperforms in just 3 steps both specialised models and leading zero-shot methods in sample quality, at a fraction of their computational cost. BAM gives the community an accessible entry point to generative computational imaging, lowers the economic and environmental cost of training imaging models, and opens a new path for research on physics-aware Bayesian computational imaging. Official page: https://bayesian-anything-model.github.io/
Figures & tables
Figure 1: BAM! One small network, many imaging problems, few steps. Each tile pairs an observation (left) with a BAM! posterior sample (right). A single 36M-parameter network, with the forward operator supplied at inference time, covers ×4 super-resolution (DIV2K, LSUN), Gaussian deblurring (AFHQ), inpainting (FFHQ, LSUN), non-linear JPEG restoration at Q=10 (FFHQ) and blind motion deblurring (Köhler), all in 3 steps. Settings are given in Section 4 .
Figure 2: Pareto frontier drawn by SOTA methods compared to BAM.
Method
NFEs
DIV2K
FFHQ
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
Gaussian deblurring
BAM
3
24.48
0.35
0.06
34.4
29.63
0.25
0.10
45.0
BAM ⋆
3
24.44
0.35
0.04
35.6
30.16
0.25
0.08
51.6
RAM
1
25.63
0.43
0.44
52.6
30.54
0.35
1.15
94.7
RAM ⋆
1
26.00
0.42
0.34
47.0
31.53
0.34
1.07
92.4
Table 1: DIV2K and FFHQ restoration at σy=0.05 . PSNR (dB; ↑ ), LPIPS ( ↓ ), CMMD ( ↓ ), and FID ( ↓ ). BAM variants use thre e steps; BAM ⋆ and RAM ⋆ denote finetuning. Bold (underline) mark the best (second-best) reported values for each dataset and problem; – denotes an unavailable result.
Figure 3: CT reconstruction and 4×4 -block uncertainty and residual maps. Left to right: ground truth, reconstruction A†y , one BAM sample, posterior mean, standard deviation over 64 draws, and residual ∣GT−mean∣ . Yellow boxes and 4× strips show the same lung region in every panel.
Figure 4: DIV2K restorations at σy=0.05 , one example per problem. BAM and BAM ⋆ use three steps. Strips show 4× zoom. More problems and comparisons in Figure 16 .
Figure 5: FFHQ restorations. Deblurring and super-resolution at σy=0.05 ; JPEG at Q=10 with pre-compression noise σJPEG=0.01 . BAM and BAM ⋆ use three steps. Strips show 4× zoom; – marks unavailable results. More problems and comparisons in Figure 17 .
Figure 6: AFHQ restorations at σy=0.05 : inpainting, demosaicing, compressed sensing. BAM and BAM ⋆ use three steps. Strips show 4× zoom. For compressed sensing, the observation column shows A†y ; – marks unavailable results. More examples in Figure 15 .
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: y vs. yσ ablation. Per-sample PSNR for conditioning on yσ and y on the Gaussian deblurring task. Dashed lines indicate the mean PSNR for each variant, while the shaded region shows the per-sample improvement.
Figure 8: Qualitative y vs. yσ ablation on LSDIR Gaussian deblurring. The observation, reconstruction obtained by conditioning BAM on the original measurement y , reconstruction obtained with noise-aware conditioning yσ , and ground truth, for Gaussian deblurring with noise level σy=0.05 . Strips below show 4× linear magnification.
Figure 9: Architecture ablation. Smoothed training LPIPS for the BAM architecture and the base RAM design with stacked operatorsWe show the first 60 k training steps. Lower is better.
Figure 10: Qualitative architecture ablation on DIV2K inpainting. The observation, reconstruction obtained with the original stacked-operator RAM design, reconstruction obtained with the BAM architecture, and ground truth. Strips below show 4× linear magnification.
Figure 11: Contrast-loss ablation on FFHQ. Removing the contrast regularizer can produce visibly over-saturated reconstructions. The auxiliary term stabilizes output contrast during heterogeneous multi-dataset, multi-operator training. Strips below show 4× linear magnification.
Figure 12: Adversarial-loss ablation. Smoothed training LPIPS during FFHQ super-resolution fine-tuning with and without a discriminator loss. Lower is better.
Method
NFEs
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
Gaussian deblurring
BAM 1-step
1
26.56
0.323
0.334
54.38
BAM 2-step
2
29.68
0.249
0.149
42.07
BAM 3-step
3
30.63
0.232
0.089
40.63
SR ×4
BAM 1-step
1
26.52
0.344
0.386
58.85
Appendix
Table 2: BAM sampling-step ablation on FFHQ at σy=0.025 . Results over 64 images per problem. PSNR ( ↑ ), LPIPS ( ↓ ), CMMD ( ↓ ), and FID ( ↓ ). Each step requires one network evaluation. Bold marks the best available value per problem and metric; – denotes an unavailable value.
Method
NFEs
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
Gaussian deblurring
BAM 1-step
1
26.53
0.324
0.329
54.55
BAM 2-step
2
28.97
0.251
0.158
43.36
BAM 3-step
3
29.63
0.250
0.099
44.98
SR ×4
BAM 1-step
1
26.54
0.346
0.396
58.14
Appendix
Table 3: BAM sampling-step ablation on FFHQ at σy=0.05 . Results over 64 images per problem. PSNR ( ↑ ), LPIPS ( ↓ ), CMMD ( ↓ ), and FID ( ↓ ). Each step requires one network evaluation. Bold marks the best available value per problem and metric; – denotes an unavailable value.
Figure 13: BAM sampling-step comparison on FFHQ at σy=0.025 . For compressed sensing, the observation column shows A†y . Yellow boxes select matching regions, with 4× linear magnification below each image.
Figure 14: BAM sampling-step comparison on FFHQ at σy=0.05 . For compressed sensing, the observation column shows A†y . Yellow boxes select matching regions, with 4× linear magnification below each image.
Method
NFEs
σy=0.025
σy=0.05
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
Deblurring
BAM
3
27.41
0.336
0.23
33.52
26.89
0.354
0.26
35.62
BAM ⋆
3
26.85
0.294
0.18
23.89
26.41
0.311
0.19
24.21
RAM
1
27.85
0.417
1.17
52.11
27.46
0.444
1.41
60.40
LATINO
8
23.37
0.440
0.48
64.81
21.48
0.486
0.50
71.87
Appendix
Table 4: AFHQ. PSNR (dB; ↑ ), LPIPS ( ↓ ), CMMD ( ↓ ), and FID ( ↓ ), with 64 images. BAM and BAM ⋆ use three steps. NFEs denote the number of neural function evaluations. Bold marks the best values, and underlining marks the second-best distinct value for each problem, noise level, and metric.
Method
NFEs
σy=0.025
σy=0.05
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
Deblurring
BAM
3
25.01
0.334
0.04
33.51
24.48
0.349
0.06
34.43
BAM ⋆
3
24.91
0.339
0.03
33.80
24.44
0.353
0.04
35.59
RAM
1
26.08
0.401
0.37
41.35
25.63
0.427
0.44
52.55
LATINO
8
23.29
0.473
0.28
74.62
22.00
0.514
0.39
100.27
Appendix
Table 5: DIV2K. PSNR (dB; ↑ ), LPIPS ( ↓ ), CMMD ( ↓ ), and FID ( ↓ ), with 64 images. BAM and BAM ⋆ use three steps. NFEs denote the number of neural function evaluations. Bold marks the best values, and underlining marks the second-best distinct value for each problem, noise level, and metric.
Method
NFEs
σy=0.025
σy=0.05
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
Deblurring
BAM
3
30.63
0.232
0.09
40.63
29.63
0.250
0.10
44.98
BAM ⋆
3
31.17
0.225
0.08
47.33
30.16
0.245
0.08
51.61
RAM
1
31.07
0.327
0.94
83.37
30.54
0.351
1.15
94.74
RAM ⋆
1
32.31
0.316
0.85
82.58
31.53
0.341
1.07
92.38
Appendix
Table 6: FFHQ. PSNR (dB; ↑ ), LPIPS ( ↓ ), CMMD ( ↓ ), and FID ( ↓ ), with 64 images. BAM and BAM ⋆ use three steps. NFEs denote the number of neural function evaluations. Bold marks the best values, and underlining marks the second-best distinct value for each problem, noise level, and metric.
Method
NFEs
σy=0.025
σy=0.05
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
Deblurring
BAM
3
25.78
0.154
1.86
89.35
25.38
0.160
2.15
90.52
BAM ⋆
3
26.16
0.155
1.02
49.41
25.79
0.160
1.26
48.28
RAM
1
25.97
0.213
1.92
77.53
25.58
0.239
2.20
87.09
RAM ⋆
1
27.25
0.304
1.87
76.90
26.59
0.333
2.11
86.54
Appendix
Table 7: LSUN. PSNR (dB; ↑ ), LPIPS ( ↓ ), FID ( ↓ ), and CMMD ( ↓ ) with 300 test images. BAM and BAM ⋆ use three steps. NFEs denote the number of neural function evaluations. Bold marks the best values, and underlining marks the second-best distinct value for each problem, noise level, and metric.
Figure 15: AFHQ restorations. One example per problem, repeated at both noise levels. BAM and BAM ⋆ use three steps. For compressed sensing, the observation column shows A†y . Matched yellow boxes show regions magnified 4× .
Figure 16: DIV2K restorations. One example per problem, repeated at both noise levels. BAM and BAM ⋆ use three steps. For compressed sensing, the observation column shows A†y . Matched yellow boxes show regions magnified 4× .
Figure 17: FFHQ restorations. One example per problem, repeated at both noise levels. BAM and BAM ⋆ use three steps. – indicates an unavailable result. For compressed sensing, the observation column shows A†y . Matched yellow boxes show regions magnified 4× .
Figure 18: LSUN restorations. One example per problem, at σy=0.05 . BAM and BAM ⋆ use three steps; RAM ⋆ denotes dataset-finetuned RAM. – indicates an unavailable result. For compressed sensing, the observation column shows A†y . Matched yellow boxes show regions magnified 4× .
Figure 19: Gaussian sparse-view CT on LIDC-IDRI. The upper panels show the sinogram, the pixelwise standard deviation of 64 BAM draws, and the absolute difference between ground truth and the empirical mean. The lower panels compare filtered backprojection, one BAM draw, the BAM mean, RAM, and the reference slice. Yellow rectangles select the same lung region; the strips immediately below show that region at 4× linear magnification.
Figure 20: Multiscale CT variability and absolute mean error. Each row compares the population standard deviation (left) with ∣GT−mean∣ (right), recomputed after block averaging each sample and the ground truth. All maps are in [0,1] image units. Black denotes zero and brighter values indicate larger variability or absolute error. These display ranges enhance contrast at each scale.
Known kernel
Spatially varying
Method
PSNR
LPIPS
CMMD
FID
PSNR
LPIPS
CMMD
FID
BAM
25.81
0.329
0.05
40.56
23.37
0.340
0.20
40.32
RAM
27.49
0.355
0.25
45.84
22.85
0.388
0.50
50.80
Appendix
Table 8: DIV2K motion deblurring. PSNR (dB; ↑ ), LPIPS ( ↓ ), CMMD ( ↓ ), and FID ( ↓ ), over 64 images. BAM uses three reconstruction steps.
Figure 21: Motion deblurring in three settings. Matched yellow boxes show regions magnified 4× .
Figure 22: Köhler clock detail. Top row, left to right: BAM, ground truth, and Carbajal et al. with joint training. Bottom row: Carbajal et al. without joint training, RAM, and the observation. Joint training refers to the kernel-estimation/deconvolution pipeline of Carbajal et al. (2023) .
Method
PSNR ↑
LPIPS ↓
CMMD ↓
FID ↓
BAM
16.14
0.746
2.11
312.88
BAM ⋆
30.95
0.236
0.12
43.10
RAM
8.96
0.826
2.41
377.62
RAM ⋆
30.44
0.424
0.63
103.79
SILO
23.54
0.446
0.83
110.87
UD2M
30.95
0.283
0.42
55.71
Appendix
Table 9: FFHQ JPEG restoration. Results over 64 images at JPEG quality factor Q=10 . FT denotes finetuning.
Figure 23: FFHQ JPEG restoration. Quality factor 10, with Gaussian noise of standard deviation 0.01 in [0,1] added before compression. All columns refer to the same clean image. Matched yellow boxes show regions magnified 4× .
Problem
BAM PSNR(mean)
BAM mean PSNR
RAM ⋆
Deblurring
32.49
30.61
32.31
Super-resolution
31.57
29.97
32.36
Inpainting
34.20
32.90
33.36
Demosaicing
38.10
36.44
37.36
Compressed sensing
36.18
34.58
33.74
Appendix
Table 10: Single-problem FFHQ results. PSNR (dB; ↑ ) over 64 images and 64 draws per image at σy=0.025 . PSNR(mean) evaluates the mean of the 64 draws; mean PSNR averages the PSNR of those draws. Bold values mark the best in each row.
Figure 24: Posterior draws for a single problem. One FFHQ image per problem at σy=0.025 . The three displayed draws illustrate diversity rather than typical random draws. The mean uses all 64 draws. Matched yellow boxes show regions magnified 4× .
Model
Latency (ms)
Throughput (qps)
PSNR ↑
LPIPS ↓
Baseline
280.56
3.56
22.49
0.400
INT8-Optimized
79.91
12.51
21.20
0.427
Appendix
Table 11: Effect of INT8 quantization on inference efficiency and reconstruction quality. We report latency, throughput, PSNR, and LPIPS for the baseline and INT8-optimized models. Bold values mark the best in each column.
Figure 25: Qualitative comparison of the baseline and INT8-optimized models on Gaussian deblurring. Reconstruction results are shown for the original baseline (left) and the INT8-optimized model (right) on test images from the LSDIR dataset. The close visual agreement between both outputs indicates that the proposed INT8 optimization substantially improves inference efficiency while preserving the reconstruction quality of the baseline model.
Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates three complementary settings: Expert, fixed expert-guided inverse reconstruction; Planner, planner-guided inverse reconstruction; and Forward, forward-system simulation for consistency checking. We benchmark leading proprietary and open-source image-centric multimodal systems, including Gemini, GPT, and Qwen, and compare them with representative task-specific non-agentic baselines. Across tasks, agentic models remain consistently weaker than specialized methods, especially on computational sensing problems such as lensless imaging, event-based reconstruction, time-of-flight imaging, and holography. Planner guidance provides only modest and inconsistent gains over the fixed-prompt Expert baseline. Although the models often generate visually plausible outputs, their reference-based fidelity remains poor, revealing a substantial gap between semantic visual competence and physically grounded imaging performance. ImagingBench provides a unified testbed for measuring this gap and tracking progress in agentic AI for computational imaging.
Generative diffusion models can provide powerful prior probability models for inverse problems in imaging, but existing implementations suffer from two key limitations: (i) the prior density is represented implicitly, and (ii) they rely on likelihood approximations that introduce sampling biases. We address these challenges by introducing a new energy-based model trained for denoising with a covariance-based regularization term that enforces consistency across different measurement conditions. The trained model can compute normalized posterior densities for diverse linear inverse problems, without additional retraining or fine tuning. In addition to preserving the sampling capabilities of diffusion models, this enables previously unavailable capabilities: energy-guided adaptive sampling that adjusts schedules on-the-fly, unbiased Metropolis-Hastings correction steps, and blind estimation of the degradation operator via Bayes rule. We validate the method on multiple datasets (ImageNet, CelebA, AFHQ) and tasks (inpainting, deblurring), demonstrating competitive or superior performance to established baselines.
Nicolas Zilberstein, Santiago Segarra, Eero Simoncelli +1
Rice University, Houston, TX, USA · Flatiron Institute, New York, NY, USA · New York University, New York, NY, USA
The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. Leading open-source alternatives, including Qwen-Image, Hunyuan-Image-3.0 and FLUX.2, are characterized by massive parameter counts (20B to 80B), making them impractical for inference, and fine-tuning on consumer-grade hardware. To address this gap, we propose Z-Image, an efficient 6B-parameter foundation generative model built upon a Scalable Single-Stream Diffusion Transformer (S3-DiT) architecture that challenges the "scale-at-all-costs" paradigm. By systematically optimizing the entire model lifecycle -- from a curated data infrastructure to a streamlined training curriculum -- we complete the full training workflow in just 314K H800 GPU hours (approx. $630K). Our few-step distillation scheme with reward post-training further yields Z-Image-Turbo, offering both sub-second inference latency on an enterprise-grade H800 GPU and compatibility with consumer-grade hardware (<16GB VRAM). Additionally, our omni-pre-training paradigm also enables efficient training of Z-Image-Edit, an editing model with impressive instruction-following capabilities. Both qualitative and quantitative experiments demonstrate that our model achieves performance comparable to or surpassing that of leading competitors across various dimensions. Most notably, Z-Image exhibits exceptional capabilities in photorealistic image generation and bilingual text rendering, delivering results that rival top-tier commercial models, thereby demonstrating that state-of-the-art results are achievable with significantly reduced computational overhead. We publicly release our code, weights, and online demo to foster the development of accessible, budget-friendly, yet state-of-the-art generative models.