cs.CVSep 28, 2026

OTT3R: Multi-View 3D Reconstruction and Fast Dataset Generation at 1% Compute

Authors: Brandon Leblanc, Charalambos Poullis

Organizations: Immersive and Creative Technologies Lab, Concordia University, Montreal, Canada

Abstract

Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks a COLMAP-like system for neural 3D pseudo-label generation. We present OTT3R (RGB-Only Tiny Transformer for 3D Reconstruction), a knowledge distillation framework that addresses both problems on a single workstation equipped with 2 GPUs. Distilling π3π^3 (959M parameters) into a 102M-parameter student yields 9.4×\times compression and up to 7×\times faster inference, trained at 1.6% of VGGT's training compute. An integrated pseudo-label pipeline offers a reliable, high-throughput alternative to COLMAP, generating dense per-pixel point maps and SE(3) camera poses for a 667K-image corpus in 3.5 hours on two commodity GPUs and succeeding on every sequence we tested, including those where COLMAP fails. The general student tracks the teacher on in-distribution monocular depth and, zero-shot, outperforms COLMAP on 7-Scenes and on DTU completion, but it does not replace the teacher on out-of-distribution multi-view geometry. The deployable artifact is the domain-specialized student: after specialization at 0.2% compute, it is 4×\times more accurate than COLMAP on 7-Scenes at 980×\times throughput, with near-teacher completion. Code is available at https://github.com/TheFourthKaramazov/OTT3R

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction

    May 12, 2026Haoyu Zhang, Zeyu Zhang, Zedong Zhou +2Feed-Forward 3D Reconstruction3D Generation

  2. Déjà View: Looping Transformers for Multi-View 3D Reconstruction

    May 28, 2026Alessandro Burzio, Tobias Fischer, Sven Elflein +9Feed-Forward 3D Reconstruction3D Reconstruction

  3. Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction

    May 7, 2026Zecheng Tang, Jiaye Fu, Qiankun Gao +5Feed-Forward 3D Reconstruction3D Reconstruction