cs.CVOct 5, 2026

Less Context, Better Geometry: Masked Geometric Encoder for Robust 3D Foundation Models

Authors: Zhimin Shao, Xijun Liu, Zhaoliang Zhang, Yutao Tang, Abhay Yadav, Rama Chellappa, Cheng Peng

Organizations: Department of Electrical and Computer Engineering Johns Hopkins University · Department of Data Science University of Virginia

Abstract

Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions also can propagate unreliable evidence from occluded or visually similar but geometrically distant views. In this paper, We introduce a Masked Geometric Encoder (MGE), which promotes the learning of robust geometric representations under incomplete cross-view context. During training, MGE strategically drops frame tokens from global attention and distills from a pretrained full-context teacher model. This allows the model to learn an intrinsically richer per-frame representation while providing sufficient intermediate supervision to avoid performance degradation. Through extensive experiments, we show that MGE leads to much stronger performance under occlusion and doppelganger views while retaining high performance on standard benchmarks. Such a richer frame representation also leads to more effective token reduction during inference. To this end, we develop a novel Anchor-Guided Adaptive token merging technique that preserves representative anchor frames while jointly merging redundant tokens from the remaining views. Compared to other efficient inference approaches, we can achieve inference speedup while consistently maintaining higher reconstruction quality, particularly in limited-view settings.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Unified Panoramic Geometry Estimation via Multi-View Foundation Models

    May 25, 2026Vukasin Bozic, Isidora Slavkovic, Dominik Narnhofer +4Omnidirectional ImagesRobust Geometric Model Estimation

  2. Glob3R: Global Structure-from-Motion with 3D Foundation Models

    Jul 10, 2026Junyuan Deng, Heng Li, Kejie Qiu +7Feed-Forward 3D ReconstructionStructure-From-Motion

  3. RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

    Jun 16, 2026Jinhao You, Shuo Lyu, Zhuohang Lyu +5Visual Geometry Grounded TransformerSelf-Supervised Vision Transformers