Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions also can propagate unreliable evidence from occluded or visually similar but geometrically distant views. In this paper, We introduce a Masked Geometric Encoder (MGE), which promotes the learning of robust geometric representations under incomplete cross-view context. During training, MGE strategically drops frame tokens from global attention and distills from a pretrained full-context teacher model. This allows the model to learn an intrinsically richer per-frame representation while providing sufficient intermediate supervision to avoid performance degradation. Through extensive experiments, we show that MGE leads to much stronger performance under occlusion and doppelganger views while retaining high performance on standard benchmarks. Such a richer frame representation also leads to more effective token reduction during inference. To this end, we develop a novel Anchor-Guided Adaptive token merging technique that preserves representative anchor frames while jointly merging redundant tokens from the remaining views. Compared to other efficient inference approaches, we can achieve inference speedup while consistently maintaining higher reconstruction quality, particularly in limited-view settings.
Figures & tables
Figure 1: Method overview. During training, MGE masks the global attention randomly to limit the all-to-all cross-view information flow. To promote richer intermediate feature representation, we further introduce constrain MGE to produce similar 3D features compared to the full-attention teacher encoder output.
Figure 2: Qualitative comparison of multi-view 3D reconstruction. Compared with VGGT and π3 , our method produces more consistent reconstructions in standard scenes (Street, Oats) and in regions with repeated or visually similar structures, including building doors, temple signboards, and church benches, where the baselines often produce shifted or duplicated geometry.
Method
All
Reference
Cross
AUC@30
RRA@15
RTA@15
AUC@30
RRA@15
RTA@15
AUC@30
RRA@15
RTA@15
LoFTR
59.60
60.34
64.68
88.22
89.13
91.30
16.86
16.90
24.78
VGGT
55.56
76.40
65.95
65.75
93.47
75.90
42.21
50.95
53.95
π3
51.23
70.73
62.48
59.42
85.40
69.50
39.10
47.41
52.34
MGE
58.68
77.52
70.07
70.51
94.99
81.48
42.72
51.60
54.60
Table 1: Visual ambiguity evaluation. We report relative-pose accuracy over all image pairs, reference-only pairs, and reference–doppelganger cross pairs. Higher is better.
Condition
Clean
Small
Medium
Large
Methods
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
VGGT
0.149
0.124
0.849
0.178
0.133
0.841
0.268
0.168
0.818
0.327
0.224
0.799
π3
0.122
0.092
0.872
0.114
0.092
0.871
0.181
0.144
0.855
0.278
0.195
0.823
MGE
0.116
0.101
0.871
0.121
0.096
0.863
0.142
0.107
0.860
0.171
0.129
0.847
FastVGGT
0.313
0.351
0.800
0.321
0.281
0.791
0.383
0.359
0.783
0.423
0.314
0.767
Fast- π3
0.163
0.141
0.837
0.184
0.156
0.824
0.188
0.155
0.829
0.225
0.172
0.816
Table 2: Occlusion evaluation on ETH3D.
Figure 3: Comparison of 3D Reconstruction on NRGBD. Following FastVGGT ( Shen et al., 2025 ) , we uniformly sample keyframes with strides from 200 to 3.
Method
Sintel
TUM-dynamics
ScanNet (seen)
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
Fast3R ( Yang et al., 2025 )
0.371
0.298
13.75
0.090
0.101
1.425
0.155
0.123
3.491
CUT3R ( Wang et al., 2025b )
0.217
0.070
0.636
0.047
0.015
0.451
0.094
0.022
0.629
FLARE ( Zhang et al., 2025b )
0.207
0.090
3.015
0.026
0.013
0.475
0.064
0.023
0.971
VGGT ( Wang et al., 2025a )
0.167
0.062
0.491
0.012
0.010
0.311
0.035
0.015
0.382
π3 ( Wang et al., 2026e )
0.074
0.040
0.282
0.014
0.009
0.312
0.031
0.013
0.347
Table 3: Camera pose estimation over standard benchmarks.
Table 7
Figure 4: Ablation of various training strategies on NGRBD
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Qualitative comparison of supervision signals on MegaDepth. From left to right: RGB input, scale-aligned π3 pseudo-GT depth, and dataset GT depth. π3 supplies dense, spatially coherent geometry where the original annotations are incomplete; black pixels indicate missing GT.
Figure 6: More qualitative comparison of multi-view 3D reconstruction. The second row highlights a representative occlusion case, where a person moves over the bench and causes severe cross-view occlusions. Our method better recovers the occluded part and preserves its geometry.
Condition
Clean
Small
Medium
Large
Methods
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
VGGT
0.149
0.124
0.849
0.178
0.133
0.841
0.268
0.168
0.818
0.327
0.224
0.799
MGE-VGGT
0.158
0.132
0.840
0.175
0.128
0.836
0.252
0.160
0.808
0.287
0.185
0.792
Methods
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
VGGT
0.716
0.887
2.122
0.738
0.941
2.155
0.758
0.961
2.948
0.898
1.227
4.673
MGE-VGGT
0.723
0.885
1.790
0.733
0.899
1.774
0.757
0.926
1.936
0.805
0.995
2.111
Appendix
Table 6: Occlusion evaluation on ETH3D for VGGT.
Figure 7: Visualizations of camera poses on visual ambiguity evaluation.
Method
stride=200
stride=100
stride=40
Acc ↓
Comp ↓
CD ↓
NC ↑
Acc ↓
Comp ↓
CD ↓
NC ↑
Acc ↓
Comp ↓
CD ↓
NC ↑
TTT3R
0.097
0.064
0.0805
0.812
0.085
0.051
0.0680
0.811
0.079
0.033
0.0560
0.792
StreamVGGT
0.092
0.083
0.0875
0.827
0.086
0.060
0.0730
0.820
0.090
0.052
0.0710
0.799
FastVGGT
0.047
0.072
0.0595
0.869
0.034
0.047
0.0405
0.871
0.016
0.015
0.0155
0.843
VGGT+AGA
0.041
0.054
0.0475
0.876
0.027
0.028
0.0275
0.879
0.016
0.014
0.0150
0.848
Ours
0.021
0.022
0.0215
0.904
0.018
0.018
0.0180
0.892
0.014
0.014
0.0140
0.856
Appendix
Table 7: Dense reconstruction results under different keyframe strides. Lower is better for Acc, Comp, and CD, while higher is better for NC.
Method
View
7-Scenes
NRGBD
Acc. ↓
Comp. ↓
NC ↑
Acc. ↓
Comp. ↓
NC ↑
Mean
Med.
Mean
Med.
Mean
Med.
Mean
Med.
Mean
Med.
Mean
Med.
Fast3R ( Yang et al., 2025 )
sparse
0.095
0.065
0.144
0.089
0.673
0.759
0.135
0.091
0.163
0.104
0.759
0.877
CUT3R ( Wang et al., 2025b )
0.093
0.049
0.102
0.051
0.704
0.805
0.104
0.041
0.079
0.031
0.822
0.968
FLARE ( Zhang et al., 2025b )
0.085
0.057
0.145
0.107
0.696
0.780
0.053
0.024
0.051
0.025
0.877
0.988
VGGT ( Wang et al., 2025a )
0.044
0.025
0.056
0.033
0.733
0.845
0.051
0.029
0.066
0.038
0.890
0.981
Appendix
Table 8: Point map estimation on 7-Scenes and NRGBD.
Method
Sintel
Bonn
KITTI
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
DUSt3R ( Wang et al., 2024 )
0.662
0.434
0.151
0.839
0.143
0.814
MASt3R ( Leroy et al., 2024 )
0.558
0.487
0.188
0.765
0.115
0.848
MonST3R ( Zhang et al., 2025a )
0.399
0.519
0.072
0.957
0.107
0.884
Fast3R ( Yang et al., 2025 )
0.638
0.422
0.194
0.772
0.138
0.834
MVDUSt3R ( Tang et al., 2025 )
0.805
0.283
0.426
0.357
0.456
0.342
Appendix
Table 9: Video depth estimation on Sintel, Bonn and KITTI.
Figure 16
Figure 9: Qualitative comparison on ETH3D. For each scene, our reconstruction is shown on the left and FastVGGT on the right. Our method generally produces more complete and globally coherent geometry, with fewer missing surfaces and disconnected structures.
Geometry estimation from perspective images has greatly advanced, maturing to the point where off-the-shelf foundation models are able to reconstruct 3D scene structure not only from multi-view imagery, but even from a single view. A natural extension is 3D reconstruction from panoramas, with the exciting prospect of recovering a full 360-degree scene from a single panoramic image. In this work, we introduce PaGeR (Panoramic Geometry Reconstruction), a framework to lift powerful 3D foundation models designed for perspective imagery to the panorama domain. Our strategy is to start from a pre-trained transformer for 3D reconstruction and turn it into a unified high-performance model that predicts scale-invariant depth, metric depth, surface normals, and sky masks from both perspective and omnidirectional images, in a single forward pass. By keeping architectural changes to a minimum and mixing perspective and panoramic images during training, PaGeR retains the rich 3D prior of the underlying foundation model while learning to also estimate geometrically consistent 360-degree scenes from single panoramas. We extensively test our method in both indoor and outdoor environments and find that it delivers state-of-the-art performance and excellent zero-shot performance across a wide range of scenes. Code, data and models are available \href.
Recent 3D geometric foundation models, such as VGGT, provide robust feed-forward 3D reconstruction by directly predicting camera poses and 3D scene points from input images. However, their results remain inaccurate, and scaling them to long sequences or large unordered image sets typically requires chunk-wise processing, which can introduce drift and inconsistency. We present Glob3R, a global SfM-style reconstruction built on 3D foundation models. Our key idea is to explicitly optimize feed-forward geometric predictions. To this end, we augment a frozen Pi3X backbone with a lightweight dense matching head that predicts image warps between selected reference frames and neighboring views. These dense warps are converted into sparse but reliable multi-view feature tracks, which provide correspondence constraints for global optimization. We further introduce a keyframe-based sliding-window association strategy that propagates tracks and relative poses across overlapping windows, enabling scalable reconstruction. Finally, we perform global motion averaging and bundle adjustment to refine camera poses, reduce scale inconsistencies, and recover dense scene geometry. Extensive experiments on indoor, outdoor, large-scale driving, and unordered SfM benchmarks demonstrate that Glob3R achieves robust and accurate reconstruction. It consistently improves over feed-forward foundation-model baselines and recent scalable reconstruction methods, while being more robust than classical SfM pipelines. The refined poses also lead to higher-quality neural rendering, validating the benefit of combining foundation-model priors with global geometric optimization. Project page: https://junyuandeng.github.io/Glob3r
Junyuan Deng, Heng Li, Kejie Qiu +7
The Hong Kong University of Science and Technology · Tongyi Lab, Alibaba Group · Nanjing University +1
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.
Jinhao You, Shuo Lyu, Zhuohang Lyu +5
University of Pennsylvania · University of California, Irvine · Nanyang Technological University +1