Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions also can propagate unreliable evidence from occluded or visually similar but geometrically distant views. In this paper, We introduce a Masked Geometric Encoder (MGE), which promotes the learning of robust geometric representations under incomplete cross-view context. During training, MGE strategically drops frame tokens from global attention and distills from a pretrained full-context teacher model. This allows the model to learn an intrinsically richer per-frame representation while providing sufficient intermediate supervision to avoid performance degradation. Through extensive experiments, we show that MGE leads to much stronger performance under occlusion and doppelganger views while retaining high performance on standard benchmarks. Such a richer frame representation also leads to more effective token reduction during inference. To this end, we develop a novel Anchor-Guided Adaptive token merging technique that preserves representative anchor frames while jointly merging redundant tokens from the remaining views. Compared to other efficient inference approaches, we can achieve inference speedup while consistently maintaining higher reconstruction quality, particularly in limited-view settings.
Figures & tables
Figure 1: Method overview. During training, MGE masks the global attention randomly to limit the all-to-all cross-view information flow. To promote richer intermediate feature representation, we further introduce constrain MGE to produce similar 3D features compared to the full-attention teacher encoder output.
Figure 2: Qualitative comparison of multi-view 3D reconstruction. Compared with VGGT and π3 , our method produces more consistent reconstructions in standard scenes (Street, Oats) and in regions with repeated or visually similar structures, including building doors, temple signboards, and church benches, where the baselines often produce shifted or duplicated geometry.
Method
All
Reference
Cross
AUC@30
RRA@15
RTA@15
AUC@30
RRA@15
RTA@15
AUC@30
RRA@15
RTA@15
LoFTR
59.60
60.34
64.68
88.22
89.13
91.30
16.86
16.90
24.78
VGGT
55.56
76.40
65.95
65.75
93.47
75.90
42.21
50.95
53.95
π3
51.23
70.73
62.48
59.42
85.40
69.50
39.10
47.41
52.34
MGE
58.68
77.52
70.07
70.51
94.99
81.48
42.72
51.60
54.60
Table 1: Visual ambiguity evaluation. We report relative-pose accuracy over all image pairs, reference-only pairs, and reference–doppelganger cross pairs. Higher is better.
Condition
Clean
Small
Medium
Large
Methods
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
VGGT
0.149
0.124
0.849
0.178
0.133
0.841
0.268
0.168
0.818
0.327
0.224
0.799
π3
0.122
0.092
0.872
0.114
0.092
0.871
0.181
0.144
0.855
0.278
0.195
0.823
MGE
0.116
0.101
0.871
0.121
0.096
0.863
0.142
0.107
0.860
0.171
0.129
0.847
FastVGGT
0.313
0.351
0.800
0.321
0.281
0.791
0.383
0.359
0.783
0.423
0.314
0.767
Fast- π3
0.163
0.141
0.837
0.184
0.156
0.824
0.188
0.155
0.829
0.225
0.172
0.816
Table 2: Occlusion evaluation on ETH3D.
Figure 3: Comparison of 3D Reconstruction on NRGBD. Following FastVGGT ( Shen et al., 2025 ) , we uniformly sample keyframes with strides from 200 to 3.
Method
Sintel
TUM-dynamics
ScanNet (seen)
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
Fast3R ( Yang et al., 2025 )
0.371
0.298
13.75
0.090
0.101
1.425
0.155
0.123
3.491
CUT3R ( Wang et al., 2025b )
0.217
0.070
0.636
0.047
0.015
0.451
0.094
0.022
0.629
FLARE ( Zhang et al., 2025b )
0.207
0.090
3.015
0.026
0.013
0.475
0.064
0.023
0.971
VGGT ( Wang et al., 2025a )
0.167
0.062
0.491
0.012
0.010
0.311
0.035
0.015
0.382
π3 ( Wang et al., 2026e )
0.074
0.040
0.282
0.014
0.009
0.312
0.031
0.013
0.347
Table 3: Camera pose estimation over standard benchmarks.
Table 7
Figure 4: Ablation of various training strategies on NGRBD
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Qualitative comparison of supervision signals on MegaDepth. From left to right: RGB input, scale-aligned π3 pseudo-GT depth, and dataset GT depth. π3 supplies dense, spatially coherent geometry where the original annotations are incomplete; black pixels indicate missing GT.
Figure 6: More qualitative comparison of multi-view 3D reconstruction. The second row highlights a representative occlusion case, where a person moves over the bench and causes severe cross-view occlusions. Our method better recovers the occluded part and preserves its geometry.
Condition
Clean
Small
Medium
Large
Methods
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
Acc. ↓
Comp. ↓
NC. ↑
VGGT
0.149
0.124
0.849
0.178
0.133
0.841
0.268
0.168
0.818
0.327
0.224
0.799
MGE-VGGT
0.158
0.132
0.840
0.175
0.128
0.836
0.252
0.160
0.808
0.287
0.185
0.792
Methods
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
ATE ↓
RPE-t ↓
RPE-r ↓
VGGT
0.716
0.887
2.122
0.738
0.941
2.155
0.758
0.961
2.948
0.898
1.227
4.673
MGE-VGGT
0.723
0.885
1.790
0.733
0.899
1.774
0.757
0.926
1.936
0.805
0.995
2.111
Appendix
Table 6: Occlusion evaluation on ETH3D for VGGT.
Figure 7: Visualizations of camera poses on visual ambiguity evaluation.
Method
stride=200
stride=100
stride=40
Acc ↓
Comp ↓
CD ↓
NC ↑
Acc ↓
Comp ↓
CD ↓
NC ↑
Acc ↓
Comp ↓
CD ↓
NC ↑
TTT3R
0.097
0.064
0.0805
0.812
0.085
0.051
0.0680
0.811
0.079
0.033
0.0560
0.792
StreamVGGT
0.092
0.083
0.0875
0.827
0.086
0.060
0.0730
0.820
0.090
0.052
0.0710
0.799
FastVGGT
0.047
0.072
0.0595
0.869
0.034
0.047
0.0405
0.871
0.016
0.015
0.0155
0.843
VGGT+AGA
0.041
0.054
0.0475
0.876
0.027
0.028
0.0275
0.879
0.016
0.014
0.0150
0.848
Ours
0.021
0.022
0.0215
0.904
0.018
0.018
0.0180
0.892
0.014
0.014
0.0140
0.856
Appendix
Table 7: Dense reconstruction results under different keyframe strides. Lower is better for Acc, Comp, and CD, while higher is better for NC.
Method
View
7-Scenes
NRGBD
Acc. ↓
Comp. ↓
NC ↑
Acc. ↓
Comp. ↓
NC ↑
Mean
Med.
Mean
Med.
Mean
Med.
Mean
Med.
Mean
Med.
Mean
Med.
Fast3R ( Yang et al., 2025 )
sparse
0.095
0.065
0.144
0.089
0.673
0.759
0.135
0.091
0.163
0.104
0.759
0.877
CUT3R ( Wang et al., 2025b )
0.093
0.049
0.102
0.051
0.704
0.805
0.104
0.041
0.079
0.031
0.822
0.968
FLARE ( Zhang et al., 2025b )
0.085
0.057
0.145
0.107
0.696
0.780
0.053
0.024
0.051
0.025
0.877
0.988
VGGT ( Wang et al., 2025a )
0.044
0.025
0.056
0.033
0.733
0.845
0.051
0.029
0.066
0.038
0.890
0.981
Appendix
Table 8: Point map estimation on 7-Scenes and NRGBD.
Method
Sintel
Bonn
KITTI
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
DUSt3R ( Wang et al., 2024 )
0.662
0.434
0.151
0.839
0.143
0.814
MASt3R ( Leroy et al., 2024 )
0.558
0.487
0.188
0.765
0.115
0.848
MonST3R ( Zhang et al., 2025a )
0.399
0.519
0.072
0.957
0.107
0.884
Fast3R ( Yang et al., 2025 )
0.638
0.422
0.194
0.772
0.138
0.834
MVDUSt3R ( Tang et al., 2025 )
0.805
0.283
0.426
0.357
0.456
0.342
Appendix
Table 9: Video depth estimation on Sintel, Bonn and KITTI.
Figure 16
Figure 9: Qualitative comparison on ETH3D. For each scene, our reconstruction is shown on the left and FastVGGT on the right. Our method generally produces more complete and globally coherent geometry, with fewer missing surfaces and disconnected structures.