Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models---their internal knowledge of 3D geometry---into a single-view estimator, we can obtain enhanced 3D consistent foundational features. Our key idea is to construct a multi-view teacher by fusing pretrained 2D foundation features with multi-view geometric features, and refining the fused representation with a discriminative ranking objective. Through our discriminative distillation framework, we enforce the learned features to be both 3D consistent and locally distinctive, while keeping them aligned with the feature space of the original foundation model to preserve the semantic structure of the pretrained representation. Consistency and local discriminability are critical for 3D computer vision problems such as forming semantic and geometric correspondences across images. To demonstrate the effectiveness of our method, we perform comprehensive experiments spanning multiple angles: direct feature analysis, dense prediction transfer, and explicit 3D lifting and rendering. Across these evaluations, our method consistently produces stronger 3D-aware foundation features that improve multi-view consistency and local discriminability while preserving the semantic transferability of the original representation.
Figures & tables
Figure 1: Our 3D-aware distillation turns pretrained 2D foundation features into representations that are more consistent across views while retaining their semantic utility. The resulting features produce cleaner cross-view correspondences and remain coherent after explicit 3D feature splatting (left). Across 3D-aware finetuning and distillation baselines [ 69 , 68 , 50 ] built from the same 2D foundation backbone [ 40 ] , our representation provides a more balanced profile over geometric correspondence, semantic transfer, dense prediction, and 3D lifting evaluations (right). The radar chart axes use per-metric min–max normalization with the minimum set to 50%.
Figure 2: Overview of our method. We fuse pretrained 2D foundation features [ 40 ] with frozen multi-view geometry features [ 30 ] to construct a multi-view teacher. We refine the teacher with geometry-supervised ranking and feature-space anchoring, encouraging valid 3D correspondences to rank above incompatible candidates while preserving the original semantic feature space. We then distill it into a single-view student, which is used as the final feature extractor at inference time.
Separation
Geometric correspondence
Semantic matching
Method
ScanNet [ 10 ]
ScanNet [ 10 ]
Navi-W [ 24 ]
DAVIS [ 42 ]
SPair-71k [ 37 ]
Margin ↑
@10px ↑
@20px ↑
@0.05 ↑
@0.1 ↑
J&F ↑
@0.05 ↑
@0.1 ↑
ViT-B
DINOv2 (baseline)
0.4108
21.72
36.11
57.97
75.94
67.82
52.98
68.56
Fit3D (ECCV’24)
0.3494
22.89
39.50
31.26
53.81
66.70
30.25
44.86
MEF (ICLR’25)
0.4288
39.17
55.06
67.71
80.87
66.82
56.78
72.07
SnD (ICLR’26)
0.4520
28.65
45.67
59.53
76.86
68.05
53.68
69.65
Table 1: Direct feature evaluation compared with recent 3D-aware feature finetuning methods [ 69 , 68 , 50 ] . Benchmarks cover ScanNet [ 10 ] feature separation and geometric correspondence, Navi-Wild [ 24 ] object correspondence, SPair-71k [ 37 ] semantic keypoint matching, and DAVIS [ 42 ] mask propagation. Best and second-best results are highlighted within each backbone group. Our method achieves the best separation, outperforms existing methods by a large margin on ScanNet, and performs strongly in all cases.
Figure 3: Qualitative nearest-neighbor correspondence results on ScanNet [ 10 ] and Navi-Wild [ 24 ] . Green and red lines denote correct and incorrect matches, respectively. Our features produce cleaner correspondences with fewer mismatches in these examples.
Figure 4: Viewpoint-gap analysis for ScanNet [ 10 ] correspondence (ViT-B, PCK@10px and PCK@20px). Our method remains robust as viewpoint changes increase.
Figure 5: Qualitative comparison of semantic segmentation using frozen features with linear probing.
Semantic segmentation
Depth estimation
ScanNet [ 10 ]
ScanNet++ [ 67 ]
ScanNet [ 10 ]
NYUv2 [ 52 ]
Method
mIoU ↑
mAcc ↑
aAcc ↑
mIoU ↑
mAcc ↑
aAcc ↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
ViT-B
DINOv2 [ 40 ]
53.69
66.29
77.63
31.69
41.41
81.47
0.1553
0.3718
78.12
0.1705
0.6566
74.90
Fit3D [ 69 ]
51.35
63.37
76.96
30.58
38.31
84.06
0.2045
0.4599
66.70
0.1975
0.7426
68.03
MEF [ 68 ]
52.49
65.17
76.93
29.40
38.62
81.43
0.1677
0.3937
75.06
0.1661
0.6222
75.98
SnD [ 50 ]
56.36
69.40
79.43
31.95
41.67
82.62
0.1449
0.3446
80.89
0.1537
0.6018
78.56
Table 2: Linear probing comparisons. We evaluate semantic segmentation on ScanNet [ 10 ] and ScanNet++ [ 67 ] , and depth estimation on ScanNet [ 10 ] and NYUv2 [ 52 ] . Our method is best or second-best on every segmentation metric and remains competitive with the SnD variants [ 50 ] on depth, while outperforming both variants on ScanNet geometric correspondence in Table 1 .
Figure 6: Qualitative comparison of depth estimation using frozen features with linear probing. Our method provides coherent depth estimates.
Figure 7: Qualitative examples of feature uplifting and rendering on pretrained 3D Gaussians [ 4 ] . PCA-projected 2D features are compared with rendered features using the same uplifting method [ 36 ] . Our features remain more coherent across views and after rendering.
Method
mIoU ↑
mAcc ↑
aAcc ↑
Per-view mIoU ↑
Multi-view agree. ↑
DINOv2 [ 40 ]
62.30
74.30
82.69
54.49
77.34
Fit3D [ 69 ]
59.54
71.04
81.45
54.21
80.73
MEF [ 68 ]
61.54
73.73
82.24
54.07
77.71
SnD [ 50 ]
62.79
74.57
82.99
56.51
80.28
SnD [ 50 ] (+blending)
63.26
75.18
83.25
57.14
80.80
Ours
63.32
74.85
83.36
57.79
82.49
Table 3: 3D point segmentation on ScanNet [ 10 ] with a linear probe on features aggregated over views. Per-view mIoU applies the same probe to the feature of each view without averaging, and multi-view agreement is the rate at which two views of the same point predict the same class. Ours matches the strongest baseline after aggregation and is best per view, with the most consistent predictions across views.
Figure 8: Feature consistency between image features and rendered features, measured as nearest-neighbor matching accuracy (PCK@10px) when matching across them. Our method retains the highest accuracy, demonstrating consistency.
Figure 9: Distribution of per-query feature separation margins on the ScanNet [ 10 ] test pairs. Vertical lines mark the mean of each method.
Method
Mean
Std.
DINOv2 [ 40 ]
0.4108
0.1935
Fit3D [ 69 ]
0.3494
0.1276
SnD [ 50 ]
0.4520
0.1388
SnD [ 50 ] (+blending)
0.3526
0.1233
MEF [ 68 ]
0.4288
0.1709
Ours
0.4632
0.1226
Appendix
Table 5: Mean and standard deviation of the per-query feature separation margins on the ScanNet test pairs.
Block
Stage
Sep. ↑
Corr. ↑
13
Early
0.3302
27.94
19
Middle
0.3682
37.03
39
Final
0.0860
13.54
Appendix
Table 6: DA3 feature layer analysis. Frozen DA3-Giant features from early, middle, and final encoder blocks, evaluated directly on ScanNet feature separation and correspondence (PCK@10px).
Model
Scale
Params
Dim.
Mode
Sep. ↑
Corr. ↑
Sem. ↑
Seg. ↑
VGGT [ 60 ]
ViT-L + Agg.
1.26B
2048
Multi
0.3096
36.98
65.89
53.26
DA3 [ 30 ]
ViT-G
1.14B
1536
Multi
0.3682
37.03
48.74
39.90
DUNE [ 46 ]
ViT-B
86M
768
Single
0.1846
33.10
59.52
47.76
Ours
ViT-B
92M
768
Single
0.4632
66.01
75.46
57.00
Appendix
Table 7: Evaluation of off-the-shelf 3D-aware foundation features. All features are evaluated frozen; Mode indicates multi-view (Multi) or single-view (Single) inference.
Feature
ScanNet
7Scenes
ETH3D
DINOv2 [ 40 ]
15.1 / 26.9
76.1 / 90.3
16.7 / 36.5
Fit3D [ 69 ]
19.0 / 38.0
77.3 / 92.9
16.0 / 38.8
MEF [ 68 ]
30.7 / 48.1
88.3 / 98.4
26.2 / 54.0
SnD [ 50 ]
20.6 / 36.9
81.1 / 92.3
15.2 / 36.9
SnD [ 50 ] (+blending)
18.8 / 34.1
77.0 / 93.0
15.2 / 38.8
Ours
65.9 / 85.4
95.4 / 100.0
49.0 / 80.6
Appendix
Table 8: Downstream camera pose estimation with PnP-RANSAC. Each entry reports pose recall R@10 / R@25 (%), within 10 cm/ 10∘ and 25 cm/ 10∘ .
Backbone
Model
Sep. ↑
Corr. ↑
Sem. ↑
Seg. ↑
ViT-L
DINOv2-L-reg
0.4315
23.68
67.83
55.84
+ Ours
0.5413
61.39
77.75
57.91
ViT-B
DINOv2-B-reg
0.4167
26.61
68.21
54.67
+ Ours
0.5127
60.74
75.72
55.72
ViT-B
DINOv3-B
0.3251
30.26
69.93
57.10
+ Ours
0.4455
64.76
73.85
56.92
Appendix
Table 9: Generalization across foundation backbones. Each backbone is compared with DDMS trained on top of it (+ Ours), using the metrics of Table 7 .
ADE-20K [ 70 ]
Pascal-VOC [ 16 ]
KITTI [ 18 ]
Model
mIoU ↑
mAcc ↑
aAcc. ↑
mIoU ↑
mAcc ↑
aAcc. ↑
RMSE ↓
AbsRel ↓
δ1↑
DINOv2
44.29
55.45
79.67
81.30
87.52
95.85
4.0027
0.1228
85.90
Ours
47.14
57.81
82.15
84.97
90.16
96.82
3.7735
0.1207
86.17
Appendix
Table 10: Out-of-distribution dense prediction transfer. We evaluate frozen-feature linear probes on ADE-20K [ 70 ] and Pascal-VOC [ 16 ] for semantic segmentation, and on KITTI [ 18 ] for depth estimation.
Figure 10: Qualitative out-of-distribution transfer with frozen-feature linear probes: semantic segmentation on ADE20K [ 70 ] and Pascal VOC [ 16 ] , and depth estimation on KITTI [ 18 ] .
Figure 11: Additional nearest-neighbor correspondence results on ScanNet [ 10 ] , compared with all baselines. Green and red lines indicate correct and incorrect matches.
Figure 12: Additional correspondence results on Navi-Wild [ 24 ] , where the same object appears under different backgrounds and viewpoints. DDMS is compared with DINOv2.
Figure 13: Additional semantic keypoint matching results on SPair-71k [ 37 ] , where keypoints are matched across different instances of the same category.
Figure 14: Additional semantic segmentation results from frozen-feature linear probing. All methods use the same linear probe.
Figure 15: Additional depth estimation results from frozen-feature linear probing, comparing DDMS with DINOv2.
Figure 16: Additional 3D point segmentation results on ScanNet [ 10 ] . Features are aggregated from posed RGB-D views onto the point cloud and classified with a linear probe.
Figure 17: Additional feature splatting results on pretrained 3D Gaussian scenes. For each method, the original 2D features are shown next to the features rendered from held-out views after uplifting.
Shanghai Key Laboratory of Intelligent Sensing and Recognition, Shanghai Jiao Tong University · Dim12 AI Inc · Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University