Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D reconstruction and visual localization. Prior work mitigates this issue with pairwise classifiers, but this design limits multi-view contextual reasoning and incurs O(n^2) inference complexity for downstream structure-from-motion (SfM). We present MVDG, a scalable multi-view disambiguation framework built on the 3D foundation model VGGT, which jointly reasons over an arbitrary number of multiview images. By incorporating 3D-aware multi-view features, our method reduces dependence on pairwise comparisons by encoding and decoding views in a single pass. We further observe that direct multi-view fine-tuning of VGGT can be unstable under noisy supervision; motivated by label ambiguity in Doppelgangers, we construct a pseudo-pairwise training set from AerialMegaDepth and show that fine-tuning on sampled subsets yields stable optimization and strong generalization to held-out scenes. Finally, because full SfM evaluation (even with faster pipelines such as GLOMAP) remains expensive, we process a pseudo-pairwise dataset for efficient validation; we derive a predictive relationship between regular SfM metrics and the classification accuracy on this pseudo-pairwise test. Experiments show that our method achieves comparable pairwise accuracy while improving both SfM accuracy and inference speed over baselines.
Figures & tables
Figure 1 : Overview of MVDG for multi-view 3D disambiguation. MVDGclassifies image pairs with similar 2D appearance that may correspond to different viewpoints. We train a classification decoder and a confidence module to select effective encoder depths adaptively. By exiting at early layers when confidence is high, our method improves efficiency while maintaining accuracy. ( Left ) Red -boxed examples are false pairs (different scene/view), and green -boxed examples are true pairs (same scene/view). ( Right ) We show a comparison of structure-from-motion result without vs. with MVDG to remove visually ambiguous images.
Figure 2 : Ambiguous inputs can mislead 3D scene understanding. We feed VGGT [ 19 ] with multi-view images from a VisymScenes scene [ 22 ] containing repetitive patterns. Among the eight input images, two viewpoint groups are evident to human observers (all from the same physical scene): six red -marked images and two orange -marked images. ( Left ) When all views, including transition geometry, are provided, the backbone robustly disambiguates views and reconstructs a coherent scene. ( Right ) When only two views from different viewpoints are provided, VGGT fails to disambiguate and reconstructs an incorrect joint scene.
Figure 3 : Architecture of the proposed pipeline. Given a reference image ( Ir ) and database images ( I1,I2 ), our method encodes all inputs and extracts geometry-aware tokens to classify whether each database image corresponds to the same scene and viewpoint as the reference. We freeze the pre-trained VGGT encoder and train a DPT-style decoder to obtain per-patch features, which are then average pooled into global descriptors. Positionally encoded descriptors are fed to an MLP-based doppelgangers classifier. To accelerate inference, a confidence-based early-exit module halts unnecessary deeper-layer computation. LN denotes LayerNorm.
Accuracy ↑
Precision ↑
Recall ↑
F1↑
Inference time (msec) ↓
VisymScenes
DG++
VisymScenes
DG++
VisymScenes
DG++
VisymScenes
DG++
SALAD
0.8104
0.6233
0.7695
0.5894
0.8862
0.8128
0.8237
0.6833
9.87
Doppelgangers++
0.9355
0.937
0.9095
0.970
0.9673
0.901
0.9375
0.934
165
Ours
0.9441
0.889
0.9215
0.877
0.9712
0.903
0.9457
0.890
103.92
Table 1 : Comparison of image disambiguation methods on pairwise test datasets.
Figure 4 : Qualitative comparison on Doppelgangers and VisymScenes [ 1 , 22 ] . The predicted classification probabilities of ours and Doppelgangers++ are thresholded at 0.5 . Because SALAD [ 5 ] outputs static global descriptors, SALAD* denotes cosine similarity thresholded at 0.3171 , which gives the best F1 score (Sec. 5.3 ). The legend shown in the first top-left negative example applies to all examples; each tile follows the same annotation scheme. Red -boxed examples are false pairs (different scene/view), and green -boxed examples are true pairs (same scene/view).
Table 6
Figure 5 : Correlation between pseudo-pairwise accuracy and downstream SfM metrics.
Figure 6 : Qualitative results on SfM using AerialMegaDepth . For visualization, we clip the point cloud to [−150,150]2 on the ground plane to focus on the scene center; this removes some points and cameras.
AUC@3
AUC@5
AUC@10
Inlier ratio
DG++
Ours
DG++
Ours
DG++
Ours
DG++
Ours
Big Ben
0.326
0.339
0.455
0.537
0.511
0.747
252/748
256/730
Louvre Pyramid
0.830
0.852
0.886
0.896
0.928
0.930
207/296
227/297
Phatheon Paris
0.912
0.910
0.929
0.929
0.944
0.942
491/537
493/536
Ponte Di Rialto
0.068
0.442
0.246
0.466
0.438
0.483
161/329
173/369
Reichstag Building
0.720
0.723
0.832
0.832
0.916
0.914
163/222
161/221
Table 4 : Quantitative SfM evaluation on AerialMegaDepth .
Figure 7 : Qualitative results on the Doppelgangers dataset [ 1 ] . Classification scores are thresholded at 0.5 : scores <0.5 are classified as false pairs, and scores ≥0.5 as true pairs.
Figure 8 : More SfM testing on AerialMegaDepth scenes.
Figure 9 : Heatmaps of encoder-layer attention. We average output patch tokens at each encoder layer and upsample them via bilinear interpolation. The heatmaps show similar attention patterns across l0,…,l11 .
Used Layers
[4]
[4,11]
[4,11,17]
[4,11,17,23]
Training Accuracy
0.9547
0.8929
0.9846
0.8904
Test Accuracy
0.5245
0.6570
0.7028
0.7598
Table 5 : Ablation study on encoder layer selection for classification accuracy.
AUC@3
AUC@5
AUC@10
Inlier ratio
Inference speed
DPT
0.640
0.723
0.801
0.7369
3.38 sec/it
Encoder
0.643
0.726
0.802
0.7360
3.10 sec/it
Table 6 : DPT-style decoder vs. transformer encoder (averaged on AerialMegaDepth ).
Figure 10 : Frequency of earliest layers meeting the confidence threshold. Statistics are computed on the VisymScenes test set [ 22 ] .
Figure 11 : Fine-tuning with our pseudo-pairwise dataset vs. skipping, on standard SfM metrics.
Visual aliasing, also known as the doppelganger problem, remains a key challenge for structure-from-motion (SfM): visually similar but physically distinct surfaces can produce incorrect image matches and degrade reconstruction quality. Previous work mitigates this issue with geometry-aware foundation-model features, but places a heavy transformer classifier on top of the backbone, making large-scale disambiguation expensive. We introduce XDG, an efficient visual disambiguation model designed for scalable SfM. Our key observation is that a 3D foundation model already performs the cross-view geometric reasoning necessary for visual disambiguation, so doppelganger classification should adapt the backbone representation directly rather than relearn pair reasoning in a separate heavy decoder. XDG fine-tunes Depth Anything 3 with lightweight LoRA adapters and repurposes its camera tokens as compact pair-level classification tokens. A compact MLP head predicts whether a candidate image pair observes the same 3D surface. Extensive experiments show that XDG provides a favorable accuracy-efficiency tradeoff: it remains competitive with the state-of-the-art disambiguation method across pairwise and reconstruction benchmarks and delivers more than a 3x inference speedup. On individual LaMAR scenes containing thousands of images, XDG saves more than 10 hours of visual disambiguation processing. Code is available at https://github.com/xtcpete/xdg.
Gonglin Chen, Ben Southall, Hanyuan Xiao +8
1USC Institute for Creative Technologies · University of Southern California · 3SRI International
Multi-view 3D reconstruction, namely, structure-from-motion followed by multi-view stereo, is a fundamental component of 3D computer vision. In general, multi-view 3D reconstruction suffers from an unknown scale ambiguity unless a reference object of known size is present in the scene. In this article, we show that multi-view images captured using a dual-pixel (DP) sensor can automatically resolve the scale ambiguity, without requiring a reference object or prior calibration. Specifically, the defocus blur observed in DP images provides sufficient information to determine the absolute scale when paired with depth maps (up to scale) recovered from multi-view 3D reconstruction. Based on this observation, we develop a simple yet effective linear method to estimate the absolute scale, followed by the intensity-based optimization stage that aligns the left and right DP images by shifting them back toward each other using cross-view blur kernels. Experiments demonstrate the effectiveness of the proposed approach across diverse scenes captured with different cameras and lenses. Code and data are available at https://github.com/lilika-makabe/dp-sfm-tpami.git
Lilika Makabe, Kohei Ashida, Hiroaki Santo +2
Graduate School of Information Science and Technology, The University of Osaka, Japan
Global Structure-from-Motion (SfM) is an efficient paradigm for recovering camera poses and sparse 3D structure from unordered images. However, its reliance on scale-ambiguous epipolar geometry makes global positioning sensitive to noisy baseline estimates and weak view-graph constraints, while false edges from visually ambiguous pairs can further degrade reconstruction. We propose DGSfM, a depth-aware global SfM pipeline that uses monocular depth maps as a scalable prior while preserving explicit multi-view optimization. For each image pair, we use a depth-aware relative pose solver to convert scale-ambiguous epipolar constraints into scale-aware relative pose constraints. We further improve robustness through view-graph filtering and depth-consistency-based correspondence pruning, which suppress false edges and matches that remain plausible under epipolar geometry alone. Finally, global scale averaging and depth-guided pose-point initialization align monocular depth maps into a common reconstruction scale and provide stable initialization for global positioning and bundle adjustment. Experiments on ETH3D and IMC2021 show that DGSfM consistently improves over strong global SfM baselines across sparse and dense matching front-ends, achieving substantial gains in pose accuracy. Code is available at https://github.com/sithu31296/DGSfM.
Sithu Aung, Viktor Kocur, Yaqing Ding +2
VRG, FEE, Czech Technical University in Prague · FMPH, Comenius University in Bratislava · Southeast University +1