Reconstructing 3D Computed Tomography (CT) images from a few X-ray projections is a highly ill-posed inverse problem due to the loss of volumetric information. We propose PhyDiCT, a training-free framework that integrates a differentiable Physics-based forward model, grounded in the Beer-Lambert law, with a text-conditioned Diffusion as a strong prior to reconstruct 3D lung CT images. We refer to our approach as training-free since the prior model is used without fine-tuning, and our goal is to steer the denoising procedure to generate samples consistent with X-ray observations. We guide the diffusion generation using Split Gibbs sampling to jointly optimize for projection fidelity (reward) and consistency with prior knowledge. Also, we introduce a test-time refinement step that enhances image realism and anatomical coherence. We extensively evaluate our method on publicly available 3D CT datasets using both perceptual and semantic metrics, demonstrating that it surpasses existing plug-and-play diffusion and fully trained reconstruction approaches. Our findings highlight that combining a strong generative prior with the underlying physics of image formation substantially improves reconstruction quality, e.g., 7.5% improvement on SSIM compared to full training methods. Code will be released at https://github.com/batmanlab/PhyDiCT.
Figures & tables
Figure 1: (a): Denoiser stage 1. The text prompt c (findings and impression tokens with <CLS> and <SEP> tokens) conditions the iterative generation of image (x0)0 , which is later used for stage 2. (b): Differentiable X-ray rendering. Real-world setup where multiple X-rays (each ω starting at the same s , ending at multiple p on detector) are used to acquire one entire y on the receiver plane. (c): Overview of reward-guided reconstruction. The shadow illustrates distribution envelope of (x0)t , which shrinks as it approaches p(x∣y) . A denoiser Ft is first applied to move the noisy sample xt (lying outside the manifold) toward the clean data distribution p(x∣c) . Subsequent likelihood step shifts the predicted (x0)t to a new variable zt (maybe lying outside) guided by the reward function to drive it closer to the posterior space p(x∣y) .
Figure 2: Comparison of reconstructed CT based on DiffVox. The first row shows raw slices, and following rows are zoomed-in regions. Arrows on first row show minor artifacts, further suppressed with EC. Second row ( bounding box ) highlights that our method accurately reconstructs subtle pathology, while avoiding overemphasis of disease compared to Prior. Third row ( arrows ) highlights the anatomical reconstruction of fissure lines.
Swiss Data Science Center (SDSC) in Paul Scherrer Institute (PSI), Villigen, Switzerland. · Swiss Data Science Center (SDSC) and ETH Zurich, Switzerland. · Swiss Data Science Center (SDSC) and Paul Scherrer Institute (PSI), Villigen, Switzerland.