Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks a COLMAP-like system for neural 3D pseudo-label generation. We present OTT3R (RGB-Only Tiny Transformer for 3D Reconstruction), a knowledge distillation framework that addresses both problems on a single workstation equipped with 2 GPUs. Distilling π3 (959M parameters) into a 102M-parameter student yields 9.4× compression and up to 7× faster inference, trained at 1.6% of VGGT's training compute. An integrated pseudo-label pipeline offers a reliable, high-throughput alternative to COLMAP, generating dense per-pixel point maps and SE(3) camera poses for a 667K-image corpus in 3.5 hours on two commodity GPUs and succeeding on every sequence we tested, including those where COLMAP fails. The general student tracks the teacher on in-distribution monocular depth and, zero-shot, outperforms COLMAP on 7-Scenes and on DTU completion, but it does not replace the teacher on out-of-distribution multi-view geometry. The deployable artifact is the domain-specialized student: after specialization at 0.2% compute, it is 4× more accurate than COLMAP on 7-Scenes at 980× throughput, with near-teacher completion. Code is available at https://github.com/TheFourthKaramazov/OTT3R
Figures & tables
Figure 1 : Sample 3D reconstructions produced by OTT3R on diverse scenes: outdoor object, indoor room, and object-centric capture, all from RGB-only input.
Figure 2 : Pseudo-label generation pipeline. Given RGB images, we generate dense 3D supervision using π3 . Our configurable sampling strategy selects frames with either strided (hard) or temporal (easy) spacing. Teacher outputs are compressed and cached for efficient student training.
7-Scenes
DTU
Method
Params
Acc ↓
Comp ↓
FPS ↑
Acc ↓
Comp ↓
FPS ↑
π3 Teacher [ 72 ]
959M
1.52
2.03
40
0.11
0.13
59
COLMAP + dense MVS ‡ [ 49 ]
–
24.27 ± 1.20
57.56 ± 8.41
0.28 ± 0.02
0.68
6.69
0.32
Distill3R † [ 31 ]
72M
8.76
8.88
78
–
–
–
OTT3R (general, zero-shot)
102M
10.16
34.84
273
0.91
3.62
331
OTT3R (domain)
102M
6.16
2.84
274
n/a
n/a
n/a
Table 1 : Multi-view reconstruction on 7-Scenes [ 51 ] and DTU [ 1 ] (Acc/Comp in cm, π3 protocol, 518 × 224). The general student is evaluated zero-shot on both benchmarks; the domain student is specialized on 7-Scenes at ∼ 0.2% of VGGT’s training compute. COLMAP runs at 640 × 480 and is evaluated only on the scenes it reconstructs.
Table 4Table 5
Figure 3 : Qualitative multi-view reconstruction on 7-Scenes test set. Top row: π3 teacher (959M). Bottom row: OTT3R student (102M).
Table 7
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Training pipeline. The teacher is run once to populate an offline cache of quantized point maps, confidence, and SE(3) poses; the student then trains against this cache without re-invoking the teacher. Gradients update only the student.
Figure 5 : Out-of-distribution failures on 7-Scenes. Left : general student, zero-shot. Right : domain-specialized student on the same sequence.
Figure 6 : Wide-baseline failures on CO3D. Left : temporal (small-baseline) sampling. Right : strided (wide-baseline) sampling of the same object, where relative-rotation errors misalign the reconstruction.
Transformer-based 3D reconstruction has emerged as a powerful paradigm for recovering geometry and appearance from multi-view observations, offering strong performance across challenging visual conditions. As these models scale to larger backbones and higher-resolution inputs, improving their efficiency becomes increasingly important for practical deployment. However, modern 3D transformer pipelines face two coupled challenges: dense multi-view attention creates substantial token-mixing overhead, and low-precision execution can destabilize geometry-sensitive representations and degrade depth, pose, and 3D consistency. To address the first challenge, we propose Lite3R, a model-agnostic teacher-student framework that replaces dense attention with Sparse Linear Attention to preserve important geometric interactions while reducing attention cost. To address the second challenge, we introduce a parameter-efficient FP8-aware quantization-aware training (FP8-aware QAT) strategy with partial attention distillation, which freezes the vast majority of pretrained backbone parameters and trains only lightweight linear-branch projection layers, enabling stable low-precision deployment while retaining pretrained geometric priors. We further evaluate Lite3R on two representative backbones, VGGT and DA3-Large, over BlendedMVS and DTU64, showing that it substantially reduces latency (1.7-2.0x) and memory usage (1.9-2.4x) while preserving competitive reconstruction quality overall. These results demonstrate that Lite3R provides an effective algorithm-system co-design approach for practical transformer-based 3D reconstruction. Code: https://github.com/AIGeeksGroup/Lite3R. Website: https://aigeeksgroup.github.io/Lite3R.
Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emerging evidence suggests that contiguous transformer layers often behave like repeated applications of similar operations, and multi-view reconstruction transformers refine their predictions progressively across decoder depth. We posit that model depth partially buys iteration, paid for inefficiently in unique parameters, and instead make that iteration explicit in architecture. Our model, DéjàView, applies a single looped transformer block recurrently to per-view features for K refinement steps. Trained once, it exposes K as an inference-time compute knob, matching or outperforming substantially larger feed-forward baselines across five reconstruction benchmarks spanning indoor, outdoor, object-centric, and driving scenes, while using a fraction of their parameters and comparable or lower compute. Importantly, the same looped block formulation outperforms an otherwise identical variant with independent per-step parameters under matched training data and compute, suggesting that explicit iteration is not merely a compute-efficient substitute for capacity but a stronger inductive bias for multi-view 3D reconstruction.
Alessandro Burzio, Tobias Fischer, Sven Elflein +9
NVIDIA · University of Modena and Reggio Emilia, AImageLab · ETH Zürich +1
Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video inputs with hundreds or thousands of frames remains challenging due to the quadratic cost of global attention layers. Recent token-merging methods accelerate these models by compressing the token sequence within the global attention layers, but they apply a uniform reduction to query tokens and key-value tokens, ignoring their functionally distinct roles in 3D reconstruction. In this work, we identify a key property of feed-forward 3D reconstruction models: query tokens encode view-specific geometric requests and are sensitive to compression, while key-value tokens represent shared scene context and tolerate aggressive compression. Guided by this insight, we propose Spark3R, a training-free acceleration framework that decouples the compression of query tokens and key-value tokens by assigning distinct reduction factors, with intra-group token merging applied to query tokens and lightweight token pruning to key-value tokens. Additionally, Spark3R adaptively adjusts the key-value reduction factor across layers, further improving the quality-efficiency trade-off. As a plug-and-play framework requiring no retraining, Spark3R integrates directly into multiple pretrained feed-forward 3D reconstruction models, including VGGT, π3, Depth-Anything-3, and VGGT-Ω, and achieves up to 28× speedup on 1,000-frame inputs while maintaining competitive reconstruction quality.
Zecheng Tang, Jiaye Fu, Qiankun Gao +5
School of Electronic and Computer Engineering, Peking University, Shenzhen 518055, China