Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emph{back-projection space} that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.
Figures & tables
Figure 1 : TomoTransformer operates on the disentangled filtered back-projection space by treating each local filtered projection at angle θ as a token. It can process a varying number of projections in the input and generate an arbitrary number of local projections.
Figure 2 : Geometric tokenizer
Method
LoDoPaB
COVID-19
KiTS23
MSD-T10
COLONOG
HNSCC
FFHQ
ImageNet
Average
FBP ( Kak and Slaney, 1988 )
29.24
28.48
30.27
28.13
28.59
28.93
27.97
26.46
28.51
U-Net ( Ronneberger et al., 2015 )
36.76
36.97
40.45
38.38
36.88
40.20
33.43
31.78
36.86
DRUNet ( Zhang et al., 2021 )
37.95
39.35
42.95
40.24
38.34
43.58
35.23
33.42
38.88
NAFNet ( Chen et al., 2022 )
38.27
39.85
43.49
40.66
38.66
44.53
35.68
33.79
39.37
Restormer ( Zamir et al., 2022 )
38.32
39.96
43.58
40.73
38.75
44.73
35.78
33.88
39.47
CTformer ( Wang et al., 2023 )
35.79
36.94
38.33
36.39
35.80
37.72
33.97
31.55
35.81
Table 1 : Fixed sparse-view reconstruction ( 64→256 ): PSNR (dB) per dataset (against the dense 256 -view reference) for TomoTransformer and the baselines. Best per column in bold .
Figure 4 : Fixed sparse-view reconstruction ( 64→256 projections) across three domains: chest CT (LoDoPaB-CT), abdominal CT (KiTS23), and a natural image (ImageNet). The box in the bottom-left of each panel reports PSNR (dB) against the dense 256 -view reference.
Figure 6 : Varied sparse-view reconstruction on LoDoPaB-CT. From a common sparse input of 128 (top) and 64 (bottom) measured views, TomoTransformer and ViewTrans reconstruct a dense 256 -view target.
Figure 7 : Zero-shot denoising. A TomoTransformer pretrained on noiseless data cleans projections corrupted by unseen noise, with no retraining. The blind-spot (BS) and complementary (CR) re-masking passes additionally denoise the measured views. Boxes report PSNR against the noise-free 256 -view FBP (right column).
Figure 8 : Zero-shot reconstruction from the real experimental data of ptychographic reconstructed projections of a nanoscale brain pillar ( Bosch et al., 2025 ) ( 768×768 , 448 axial slices), with TomoTransformer and no fine-tuning . The box in the bottom-left reports per-slice PSNR (dB) against the dense 512 -view FBP of the same measurements.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
LoDoPaB
COVID-19
KiTS23
MSD-T10
COLONOG
HNSCC
FFHQ
ImageNet
Average
FBP ( Kak and Slaney, 1988 )
0.635
0.654
0.662
0.554
0.610
0.395
0.631
0.598
0.592
U-Net ( Ronneberger et al., 2015 )
0.895
0.919
0.948
0.919
0.889
0.895
0.857
0.836
0.895
DRUNet ( Zhang et al., 2021 )
0.910
0.937
0.965
0.939
0.907
0.945
0.892
0.871
0.921
NAFNet ( Chen et al., 2022 )
0.914
0.941
0.968
0.943
0.911
0.954
0.900
0.877
0.926
Restormer ( Zamir et al., 2022 )
0.915
0.943
0.969
0.944
0.912
0.955
0.902
0.879
0.927
CTformer ( Wang et al., 2023 )
0.868
0.908
0.908
0.871
0.855
0.815
0.856
0.809
0.861
Appendix
Table 2 : Fixed sparse-view reconstruction ( 64→256 ): SSIM per dataset (against the dense 256 -view reference) for TomoTransformer and the baselines. Best per column in bold .
FBP
ViewTrans
TomoTransformer
Target
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
128→256
36.68
0.873
37.75
0.914
42.40
0.959
64→256
29.24
0.635
33.07
0.832
37.68
0.910
Appendix
Table 3 : Varied sparse-view reconstruction on LoDoPaB-CT ( →256 views), against the dense 256 -view FBP reference ( PSNR in dB / SSIM , averaged over the 32 test slices). The last row is the very-sparse regime, with only 64 measured views—the extreme sparse end of the training distribution. Best per column in bold .
FBP
TomoTransformer
Target
PSNR
SSIM
PSNR
SSIM
128→256
31.8
0.722
35.8
0.837
128→512
31.8
0.722
38.8
0.916
Appendix
Table 4 : Zero-shot reconstruction on the real nanoscale brain pillar of Section 5.5 , against the dense 512 -view FBP of the same measurements ( PSNR in dB / SSIM , averaged over all 448 axial slices). The pre-trained TomoTransformer receives the same 128 measured views in both rows and is applied with no fine-tuning ; only the number of queried views differs.
Figure 9 : Pre-trained single-purpose baselines of Section 5.2 applied to the real experimental brain projections ( 64→256 ). Each baseline was trained on one fixed geometry, reconstructing a 256 -view FBP from 64 uniformly distributed views. However, the 617 real experimental projections are not uniformly spaced. TomoTransformer is trained on varied sparsity patterns and conditions on each token’s actual angle, and can conveniently handle this real scenario. The box in the bottom-left reports PSNR (dB) against the dense 256 -view FBP of the same measurements.
Method
LoDoPaB
COVID-19
KiTS23
MSD-T10
COLONOG
HNSCC
FFHQ
ImageNet
Average
FBP ( Kak and Slaney, 1988 )
29.24
28.48
30.27
28.13
28.59
28.93
27.97
26.46
28.51
TomoTransformer (standard)
37.41
38.92
42.38
39.77
38.08
39.34
36.06
33.37
38.17
TomoTransformer (GAM)
37.47
39.05
42.68
40.02
38.18
39.61
36.11
33.42
38.32
Appendix
Table 5 : Ablation studies on varied sparse-view reconstruction ( 64→256 ): PSNR (dB) per dataset (against the dense 256 -view reference). The two TomoTransformer variants are identical except for the geometric attention bias of Eq. ( 6 ). Best per column in bold .
Method
LoDoPaB
COVID-19
KiTS23
MSD-T10
COLONOG
HNSCC
FFHQ
ImageNet
Average
FBP ( Kak and Slaney, 1988 )
0.635
0.654
0.662
0.554
0.610
0.395
0.631
0.598
0.592
TomoTransformer (standard)
0.906
0.940
0.963
0.935
0.907
0.917
0.907
0.871
0.918
TomoTransformer (GAM)
0.907
0.941
0.964
0.937
0.908
0.920
0.908
0.872
0.920
Appendix
Table 6 : Ablation studies on varied sparse-view reconstruction ( 64→256 ): SSIM per dataset (against the dense 256 -view reference). Best per column in bold .
Computed Tomography (CT) is a widely used imaging modality in medical and industrial applications. To limit radiation exposure and measurement time, there is a growing interest in sparse-view CT, where the number of projection views is significantly reduced. Deep neural networks have shown great promise in improving reconstruction quality in sparse-view CT, especially generative diffusion models. However, these methods struggle to scale to large 3D volumes due to several reasons: (i) the high memory and computational requirements of 3D models, (ii) the lack of large 3D training datasets, and (iii) the inconsistencies across slices when using 2D models independently on each slice. We overcome these limitations and scale diffusion-based sparse-view CT reconstruction to large 3D volumes by combining conditional diffusion with explicit data consistency. We propose Conditional Diffusion Posterior Alignment (CDPA) to enable scalable 3D sparse-view CT reconstruction. A 2D U-Net diffusion model is conditioned on an initial 3D reconstruction to improve inter-slice consistency, combined with data-consistency alignment to match measured projections. Experiments on synthetic and real Cone Beam CT (CBCT) data show state-of-the-art performance, with ablations that confirm the synergistic effects of the proposed pipeline. Finally, we show that the same principles also strengthen fast denoising U-Nets, yielding near-diffusion quality at a fraction of the computational cost.
Luis Barba, Johannes Kirschner, Benjamin Bejar
Swiss Data Science Center (SDSC) in Paul Scherrer Institute (PSI), Villigen, Switzerland. · Swiss Data Science Center (SDSC) and ETH Zurich, Switzerland. · Swiss Data Science Center (SDSC) and Paul Scherrer Institute (PSI), Villigen, Switzerland.
Computed tomography (CT) throughput is limited by scan time, which grows with both the number of projections acquired and the detector integration time for each projection. Reconstructing high-quality volumes from sparse-view or low-dose measurements therefore depends on using an informative prior, typically a neural network trained for one specific scan setting and retrained whenever the modality, geometry, or material changes. We investigate whether a single diffusion model trained across several imaging domains can instead serve as a reusable prior for heterogeneous CT reconstruction problems. We evaluate the proposed method using the same diffusion visual transformer model and normalized denoising strength on three datasets that differ in modality, beam geometry, material, and degradation type, spanning additively manufactured metal parts and concrete microstructure imaged with cone-bean X-ray CT and parallel-beam neutron CT respectively. The proposed method improves upon analytic reconstructions in all three cases, demonstrating transferability across the evaluated problems and providing a step toward a reusable foundation prior for heterogeneous CT reconstruction.
Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emerging evidence suggests that contiguous transformer layers often behave like repeated applications of similar operations, and multi-view reconstruction transformers refine their predictions progressively across decoder depth. We posit that model depth partially buys iteration, paid for inefficiently in unique parameters, and instead make that iteration explicit in architecture. Our model, DéjàView, applies a single looped transformer block recurrently to per-view features for K refinement steps. Trained once, it exposes K as an inference-time compute knob, matching or outperforming substantially larger feed-forward baselines across five reconstruction benchmarks spanning indoor, outdoor, object-centric, and driving scenes, while using a fraction of their parameters and comparable or lower compute. Importantly, the same looped block formulation outperforms an otherwise identical variant with independent per-step parameters under matched training data and compute, suggesting that explicit iteration is not merely a compute-efficient substitute for capacity but a stronger inductive bias for multi-view 3D reconstruction.
Alessandro Burzio, Tobias Fischer, Sven Elflein +9
NVIDIA · University of Modena and Reggio Emilia, AImageLab · ETH Zürich +1