Emergent Multi-View Geometry Through Self-Distillation
Authors: David Nordström, Thibaut Loiseau, Vincent Lepetit, Michael Felsberg, Guillaume Bourmaud, Fredrik Kahl
Organizations: Chalmers University of Technology, Sweden · LIGM, Ecole des Ponts, Univ. Gustave Eiffel, CNRS, France · Linköping University, Sweden · Univ. Bordeaux, CNRS, Bordeaux INP, IMS, UMR 5218, France
Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM, and Muskie on correspondence estimation, camera pose estimation, and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.
Figures & tables
Figure 1: Poincar3 main results. Poincar3 outperforms state-of-the-art SSL baselines across multi-view geometric tasks, including pose estimation, point cloud estimation, and image matching.
Figure 2: Emerging matching capabilities without supervision. Given query keypoints, we visualize the tracks formed by selecting the patch with the highest attention activation. Despite receiving neither correspondence labels nor explicit attention supervision, the model learns to identify patch correspondences across images, suggesting the emergence of 3D-aware representations from SSL.
Figure 3: Poincar3 learns visual features without labels via multi-view self-distillation. The teacher, an EMA of the student, processes a full image sequence, while the student processes a masked and photometrically augmented subset. After the multi-view transformer, student and teacher embeddings are aligned using an image-level loss on CLS tokens and a patch loss on masked tokens.
Figure 4
Figure 5: Qualitative feature comparison. Given a query patch (green star), we visualize its feature correlations across frames and mark the maximum response (green circle). RGB reconstruction produces diffuse responses, whereas our method yields localized responses at corresponding patches.
Figure 6
Multi-View Relative Pose
Point Cloud Estimation
Method
RE10K
ScanNet++
MegaDepth
ETH3D
DTU
Metric →
@3 ∘
@30 ∘
@3 ∘
@30 ∘
@3 ∘
@30 ∘
Acc.
NC
Acc.
NC
Train only heads on top of the frozen backbone
DINOv3
0.0
18.9
0.0
9.4
1.5
55.5
1.18
0.59
10.76
0.54
Muskie
1.1
36.3
0.0
19.0
0.1
46.9
1.12
0.60
11.43
0.55
MuM
0.0
29.8
0.0
18.7
0.1
48.5
1.12
0.60
12.60
0.54
Table 1: Feed-forward reconstruction. Reporting relative pose accuracy by AUC over 10 random frames and point cloud accuracy by median accuracy (acc.) in mm and normal consistency (NC) by the cosine of the angle between the normals. Training with 4 H200 GPUs for 3 days each.
Method
ScanNet ( Dai et al., 2017 )
NAVI ( Jampani et al., 2023 )
PCK@ →
5px
10px
25px
50px
5px
10px
25px
50px
Feature Nearest-Neighbor Matching
Feed-Forward Reconstruction
VGGT- Ω
24.5
32.8
52.9
69.0
16.8
25.9
51.6
70.9
DA3
21.5
22.9
28.2
35.2
14.6
18.6
30.5
46.9
π3
26.7
38.6
61.3
75.6
18.2
30.3
60.2
78.8
Table 2: Multi-view correspondence estimation. Zero-shot patch tracking across 8 views. While Poincar3 has strong representations at all layers, we compare to the best layer for fair comparison.
Figure 8: Multi-view correspondence tracks. We visualize the predicted tracks (blue), the ground-truth (green), and the error (red). Poincar3 gives accurate zero-shot multi-view geometry.
Figure 10Figure 11
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Poincaré adapter. We visualize the alignment between visual features and camera trajectories with a lightweight Poincaré adapter , following Fig. 1 of Chen et al. (2026) . Applied to Poincar3, it unrolls the feature space toward the ground-truth camera trajectories.
Figure 12: PCA feature visualization. We visualize the 3 principal components as RGB for different sequences from the same scene.
Figure 13: Additional multi-view correspondence visualizations. Including also DINOv3 as a baseline.
Method
Nearest-Neighbor
Linear Probe
PCK →
8px
16px
32px
8px
16px
32px
Feed-forward Reconstruction
VGGT- Ω
2.3
6.8
17.0
26.4
47.9
67.5
DA3
0.8
2.4
6.8
29.2
49.0
66.7
π3
8.0
17.1
30.2
32.9
55.3
73.5
Self-supervised Models
Appendix
Table 5: Lightweight two-view matching. Matching robustness on ScanNet-1500 ( Dai et al., 2017 ; Sarlin et al., 2020 ) .
Context
ImageNet-1K ( Deng et al., 2009 ; Krizhevsky et al., 2012 )
ADE20K ( Zhou et al., 2017 )
DINOv3
85.1
49.8
Muskie
27.5
6.4
MuM
27.5
4.6
Poincar3
34.4
9.6
Appendix
Table 6: Semantic performance. Image classification accuracy on ImageNet-1K and semantic segmentation (mIoU) on ADE20K.
Figure 14: Layer-wise performance. We plot the multi-view correspondence accuracy at different depths. We find that Poincar3 encodes multi-view geometry throughout the network, while the performance of supervised methods collapses at later layers.
Datasets
Type / Source
Weight
№Scenes
SpatialVID ( Wang et al., 2026a )
Outdoor / Video
1
176,749
DL3DV ( Ling et al., 2024 )
Mixed / Video
1
10,000
RealEstate10K ( Zhou et al., 2018 )
Indoor / Video
1
7,850
MegaDepth ( Li and Snavely, 2018 )
Outdoor / MVS
1
169
AerialMD ( Vuong et al., 2025 )
Aerial / MVS
1
124
BlendedMVS ( Yao et al., 2020 )
Aerial / Mesh
1
493
Appendix
Table 7: Dataset mixture for Poincar3. The top part contains large-scale internet video datasets with noisy or no annotations, while the bottom part contains 3D datasets with annotations. The weight is proportional to the probability of sampling from the respective dataset.
School of Computation, Information, and Technology, Technical University of Munich, Germany · Munich Center for Machine Learning, Germany · Department of Computing, Imperial College London, United Kingdom