Computed tomography (CT) throughput is limited by scan time, which grows with both the number of projections acquired and the detector integration time for each projection. Reconstructing high-quality volumes from sparse-view or low-dose measurements therefore depends on using an informative prior, typically a neural network trained for one specific scan setting and retrained whenever the modality, geometry, or material changes. We investigate whether a single diffusion model trained across several imaging domains can instead serve as a reusable prior for heterogeneous CT reconstruction problems. We evaluate the proposed method using the same diffusion visual transformer model and normalized denoising strength on three datasets that differ in modality, beam geometry, material, and degradation type, spanning additively manufactured metal parts and concrete microstructure imaged with cone-bean X-ray CT and parallel-beam neutron CT respectively. The proposed method improves upon analytic reconstructions in all three cases, demonstrating transferability across the evaluated problems and providing a step toward a reusable foundation prior for heterogeneous CT reconstruction.
Figures & tables
Fig. 1 : Diffusion model training data: cell microscopy, powder micro-XCT, metal AM XCT, and concrete XCT. Top row: simulation masks. Middle row: simulated response. Bottom row: unpaired experimental examples.
Nickel AM
Steel AM
Concrete Microstructure
Modality
X-ray
X-ray
Neutron
Geometry
Cone-beam
Cone-beam
Parallel-beam
Data source
Simulated
Measured
Measured
Angle range [Input, Ref.]
[ 197∘ , 360∘ ]
[ 197∘ , 197∘ ]
[ 180∘ , 180∘ ]
# of views [Input, Ref.]
[145, 2132]
[1000, 1000]
[34, 546]
Integration time [Input, Ref.]
–
[0.6 s, 3.6 s]
[30 s, 30 s]
TABLE I : Overview of three testing datasets, which differ in imaging modality, beam geometry, and material. Bracketed entries give the input and reference acquisitions respectively. Both X-ray datasets are beam hardening and scatter corrected before reconstruction. The cone-beam short scans span 180∘ plus the 17∘ fan angle, while the parallel-beam neutron scan requires only 180∘ . Note that the nickel data is simulated, so it has no detector integration time; the equivalent photon budget is set by the noise added to its projections instead.
Fig. 2 : Qualitative comparison for each dataset. Columns from left to right show the reference reconstruction, the analytic input, the PnP-BM3D reconstruction, and the proposed PnP-Diffusion-ViT reconstruction. The reference reconstructions are generated using full-view FDK for nickel AM, long-integration-time MBIR for steel AM, and full-view FBP for concrete microstructure. Both PnP reconstructions use the same algorithmic parameters ( σ=0.25 , K=5 ). The same frozen prior removes the dominant degradation of all three problems (noise, ring artifacts, and sparse-view streaking) while retaining fine structure that BM3D smooths away.
Dataset
Method
PSNR
SSIM
HFEN
Time (s / slice)
Nickel AM
FDK
9.446
0.026
3.710
0.002
PnP-BM3D
25.769
0.823
0.744
40.8
PnP-Diff-ViT
25.292
0.737
0.733
3.67
Steel AM
FDK
13.650
0.060
3.331
0.008
PnP-BM3D
24.580
0.601
0.872
47.7
PnP-Diff-ViT
26.680
0.700
0.571
9.59
TABLE II : Quantitative comparison against the reference reconstruction. Best metric per dataset is bolded; higher is better for PSNR and SSIM, and lower is better for HFEN. Time is reported in seconds per slice, excluding data loading and output writing.
Generative models, particularly Diffusion Models (DM), have shown strong potential for Computed Tomography (CT) reconstruction serving as expressive priors for solving ill-posed inverse problems. However, diffusion-based reconstruction relies on Stochastic Differential Equations (SDEs) for forward diffusion and reverse denoising, where such stochasticity can interfere with repeated data consistency corrections in CT reconstruction. Since CT reconstruction is often time-critical in clinical and interventional scenarios, improving reconstruction efficiency is essential. In contrast, Flow Matching (FM) models sampling as a deterministic Ordinary Differential Equation (ODE), yielding smooth trajectories without stochastic noise injection. This deterministic formulation is naturally compatible with repeated data consistency operations. Furthermore, we observe that FM-predicted velocity fields exhibit strong correlations across adjacent steps. Motivated by this, we propose an FM-based CT reconstruction framework (FMCT) and an efficient variant (EFMCT) that reuses previously predicted velocity fields over consecutive steps to substantially reduce the number of Neural network Function Evaluations (NFEs), thereby improving inference efficiency. We provide theoretical analysis showing that the error introduced by velocity reuse is bounded when combined with data consistency operations. Extensive experiments demonstrate that FMCT/EFMCT achieve competitive reconstruction quality while significantly improving computational efficiency compared with diffusion-based methods. The codebase is open-sourced at https://github.com/EFMCT/EFMCT.
Jiayang Shi, Lincen Yang, Zhong Li +3
Centrum Wiskunde en Informatica · LIACS, Leiden University · Great Bay University +1
Computed Tomography (CT) is a widely used imaging modality in medical and industrial applications. To limit radiation exposure and measurement time, there is a growing interest in sparse-view CT, where the number of projection views is significantly reduced. Deep neural networks have shown great promise in improving reconstruction quality in sparse-view CT, especially generative diffusion models. However, these methods struggle to scale to large 3D volumes due to several reasons: (i) the high memory and computational requirements of 3D models, (ii) the lack of large 3D training datasets, and (iii) the inconsistencies across slices when using 2D models independently on each slice. We overcome these limitations and scale diffusion-based sparse-view CT reconstruction to large 3D volumes by combining conditional diffusion with explicit data consistency. We propose Conditional Diffusion Posterior Alignment (CDPA) to enable scalable 3D sparse-view CT reconstruction. A 2D U-Net diffusion model is conditioned on an initial 3D reconstruction to improve inter-slice consistency, combined with data-consistency alignment to match measured projections. Experiments on synthetic and real Cone Beam CT (CBCT) data show state-of-the-art performance, with ablations that confirm the synergistic effects of the proposed pipeline. Finally, we show that the same principles also strengthen fast denoising U-Nets, yielding near-diffusion quality at a fraction of the computational cost.
Luis Barba, Johannes Kirschner, Benjamin Bejar
Swiss Data Science Center (SDSC) in Paul Scherrer Institute (PSI), Villigen, Switzerland. · Swiss Data Science Center (SDSC) and ETH Zurich, Switzerland. · Swiss Data Science Center (SDSC) and Paul Scherrer Institute (PSI), Villigen, Switzerland.
Neural representations (NRs), such as neural fields and 3D Gaussians, effectively model volumetric data in computed tomography (CT) but suffer from severe artifacts under sparse-view settings. To address this, we propose DiffNR, a novel framework that enhances NR optimization with diffusion priors. At its core is SliceFixer, a single-step diffusion model designed to correct artifacts in degraded slices. We integrate specialized conditioning layers into the network and develop tailored data curation strategies to support model finetuning. During reconstruction, SliceFixer periodically generates pseudo-reference volumes, providing auxiliary 3D perceptual supervision to fix underconstrained regions. Compared to prior methods that embed CT solvers into time-consuming iterative denoising, our repair-and-augment strategy avoids frequent diffusion model queries, leading to better runtime performance. Extensive experiments show that DiffNR improves PSNR by 3.99 dB on average, generalizes well across domains, and maintains efficient optimization.
Shiyan Su, Ruyi Zha, Danli Shi +2
1Monash University · 2The Australian National University · 3Hong Kong Polytechnic University