cs.CVSep 29, 2026

TomoTransformer: Towards a Foundation Model for CT Reconstruction

Authors: AmirEhsan Khorashadizadeh, Benjamín Béjar

Organizations: Swiss Data Science Center (SDSC) in Paul Scherrer Institute (PSI), Villigen, Switzerland.

Abstract

Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emph{back-projection space} that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Conditional Diffusion Posterior Alignment for Sparse-View CT Reconstruction

    Apr 23, 2026Luis Barba, Johannes Kirschner, Benjamin BejarCone-Beam Computed TomographySparse-View

  2. Toward a Foundation Plug-and-Play Prior for Computed Tomography Reconstruction via a Multimodal Diffusion Model

    Aug 24, 2026Haley Duba-Sullivan, Patxi Fernandez-Zelaia, Obaidullah Rahman +1Cone-Beam Computed TomographyDiffusion Priors

  3. Déjà View: Looping Transformers for Multi-View 3D Reconstruction

    May 28, 2026Alessandro Burzio, Tobias Fischer, Sven Elflein +9Feed-Forward 3D Reconstruction3D Reconstruction