Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D reconstruction and visual localization. Prior work mitigates this issue with pairwise classifiers, but this design limits multi-view contextual reasoning and incurs O(n^2) inference complexity for downstream structure-from-motion (SfM). We present MVDG, a scalable multi-view disambiguation framework built on the 3D foundation model VGGT, which jointly reasons over an arbitrary number of multiview images. By incorporating 3D-aware multi-view features, our method reduces dependence on pairwise comparisons by encoding and decoding views in a single pass. We further observe that direct multi-view fine-tuning of VGGT can be unstable under noisy supervision; motivated by label ambiguity in Doppelgangers, we construct a pseudo-pairwise training set from AerialMegaDepth and show that fine-tuning on sampled subsets yields stable optimization and strong generalization to held-out scenes. Finally, because full SfM evaluation (even with faster pipelines such as GLOMAP) remains expensive, we process a pseudo-pairwise dataset for efficient validation; we derive a predictive relationship between regular SfM metrics and the classification accuracy on this pseudo-pairwise test. Experiments show that our method achieves comparable pairwise accuracy while improving both SfM accuracy and inference speed over baselines.
Figures & tables
Figure 1 : Overview of MVDG for multi-view 3D disambiguation. MVDGclassifies image pairs with similar 2D appearance that may correspond to different viewpoints. We train a classification decoder and a confidence module to select effective encoder depths adaptively. By exiting at early layers when confidence is high, our method improves efficiency while maintaining accuracy. ( Left ) Red -boxed examples are false pairs (different scene/view), and green -boxed examples are true pairs (same scene/view). ( Right ) We show a comparison of structure-from-motion result without vs. with MVDG to remove visually ambiguous images.
Figure 2 : Ambiguous inputs can mislead 3D scene understanding. We feed VGGT [ 19 ] with multi-view images from a VisymScenes scene [ 22 ] containing repetitive patterns. Among the eight input images, two viewpoint groups are evident to human observers (all from the same physical scene): six red -marked images and two orange -marked images. ( Left ) When all views, including transition geometry, are provided, the backbone robustly disambiguates views and reconstructs a coherent scene. ( Right ) When only two views from different viewpoints are provided, VGGT fails to disambiguate and reconstructs an incorrect joint scene.
Figure 3 : Architecture of the proposed pipeline. Given a reference image ( Ir ) and database images ( I1,I2 ), our method encodes all inputs and extracts geometry-aware tokens to classify whether each database image corresponds to the same scene and viewpoint as the reference. We freeze the pre-trained VGGT encoder and train a DPT-style decoder to obtain per-patch features, which are then average pooled into global descriptors. Positionally encoded descriptors are fed to an MLP-based doppelgangers classifier. To accelerate inference, a confidence-based early-exit module halts unnecessary deeper-layer computation. LN denotes LayerNorm.
Accuracy ↑
Precision ↑
Recall ↑
F1↑
Inference time (msec) ↓
VisymScenes
DG++
VisymScenes
DG++
VisymScenes
DG++
VisymScenes
DG++
SALAD
0.8104
0.6233
0.7695
0.5894
0.8862
0.8128
0.8237
0.6833
9.87
Doppelgangers++
0.9355
0.937
0.9095
0.970
0.9673
0.901
0.9375
0.934
165
Ours
0.9441
0.889
0.9215
0.877
0.9712
0.903
0.9457
0.890
103.92
Table 1 : Comparison of image disambiguation methods on pairwise test datasets.
Figure 4 : Qualitative comparison on Doppelgangers and VisymScenes [ 1 , 22 ] . The predicted classification probabilities of ours and Doppelgangers++ are thresholded at 0.5 . Because SALAD [ 5 ] outputs static global descriptors, SALAD* denotes cosine similarity thresholded at 0.3171 , which gives the best F1 score (Sec. 5.3 ). The legend shown in the first top-left negative example applies to all examples; each tile follows the same annotation scheme. Red -boxed examples are false pairs (different scene/view), and green -boxed examples are true pairs (same scene/view).
Table 6
Figure 5 : Correlation between pseudo-pairwise accuracy and downstream SfM metrics.
Figure 6 : Qualitative results on SfM using AerialMegaDepth . For visualization, we clip the point cloud to [−150,150]2 on the ground plane to focus on the scene center; this removes some points and cameras.
AUC@3
AUC@5
AUC@10
Inlier ratio
DG++
Ours
DG++
Ours
DG++
Ours
DG++
Ours
Big Ben
0.326
0.339
0.455
0.537
0.511
0.747
252/748
256/730
Louvre Pyramid
0.830
0.852
0.886
0.896
0.928
0.930
207/296
227/297
Phatheon Paris
0.912
0.910
0.929
0.929
0.944
0.942
491/537
493/536
Ponte Di Rialto
0.068
0.442
0.246
0.466
0.438
0.483
161/329
173/369
Reichstag Building
0.720
0.723
0.832
0.832
0.916
0.914
163/222
161/221
Table 4 : Quantitative SfM evaluation on AerialMegaDepth .
Figure 7 : Qualitative results on the Doppelgangers dataset [ 1 ] . Classification scores are thresholded at 0.5 : scores <0.5 are classified as false pairs, and scores ≥0.5 as true pairs.
Figure 8 : More SfM testing on AerialMegaDepth scenes.
Figure 9 : Heatmaps of encoder-layer attention. We average output patch tokens at each encoder layer and upsample them via bilinear interpolation. The heatmaps show similar attention patterns across l0,…,l11 .
Used Layers
[4]
[4,11]
[4,11,17]
[4,11,17,23]
Training Accuracy
0.9547
0.8929
0.9846
0.8904
Test Accuracy
0.5245
0.6570
0.7028
0.7598
Table 5 : Ablation study on encoder layer selection for classification accuracy.
AUC@3
AUC@5
AUC@10
Inlier ratio
Inference speed
DPT
0.640
0.723
0.801
0.7369
3.38 sec/it
Encoder
0.643
0.726
0.802
0.7360
3.10 sec/it
Table 6 : DPT-style decoder vs. transformer encoder (averaged on AerialMegaDepth ).
Figure 10 : Frequency of earliest layers meeting the confidence threshold. Statistics are computed on the VisymScenes test set [ 22 ] .
Figure 11 : Fine-tuning with our pseudo-pairwise dataset vs. skipping, on standard SfM metrics.