We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors. In contrast to previous approaches that normally evaluate depth and camera intrinsics separately or evaluate point-clouds with geometric similarity metrics, which cannot directly reflect models' point-to-point distance estimation capability, Pumpire directly assesses point-to-point distances from the reconstructed geometry. To this end, we collect a large-scale and diverse dataset (pumpire-6k) comprising 100 real-world scenes, each annotated with physically measured point-pair distances and containing 64 frames, for a total of 6,400 frames. Building on this dataset, we establish a holistic evaluation protocol that covers both image- and video-level 3D foundation models and enables direct assessment of point-pair distance errors and cross-setting comparison. We conduct extensive experiments across 29 baseline configurations of representative 3D foundation models and provide a comprehensive analysis of the results. By offering this benchmark, we target the more fundamental ability to perceive and estimate physical scale in the reconstructed 3D space, which prior evaluation protocols have largely overlooked. The project page can be found at https://pumpire.github.io/
Figures & tables
Figure 1: Overview of the Pumpire evaluation protocol. Pumpire provides a unified evaluation of image- and video-level 3D foundation models under geometry estimation and completion settings. Estimation models recover geometry from RGB inputs, whereas completion models additionally condition on depth priors. Pumpire back-projects depth maps predicted by 3D foundation models with their associated camera intrinsics and quantifies the results with the proposed metrics. Their predicted distance is compared with the physically measured ground truth using the proposed metrics as introduced in Section 4 .
Figure 2: Limitations of existing depth benchmarks.
Optical properties
N
Texture patterns
N
Geometric structures
N
Common
84
Common
83
Common
63
Transparent surfaces
12
Repetitive textures
13
Slender structures
25
Reflective surfaces
10
Textureless regions
5
Near sharp edges
11
Table 1: Surface coverage of pumpire-6k. “N” denotes the number of scenes containing each property. Categories are not mutually exclusive, as a scene may exhibit multiple properties.
N
AE
RE
LE
RSD
δ1.05
δ1.10
δ1.25
Normal
76
0.026
0.038
0.036
0.062
79.32
92.56
98.66
Challenging
24
4.047
8.390
1.228
0.887
21.03
27.60
36.72
Overall
100
0.991
2.043
0.322
0.260
65.33
76.97
83.80
Table 2: Performance of raw sensor depth on normal and challenging scenes. N denotes the number of scenes. AE is reported in meters, and δ metrics are reported as percentages.
Figure 3: Representative challenging scenes with unreliable sensor depth. These cases correspond to the challenging subset identified via the RSD threshold ( >0.1 ).
Normal
Challenging
Model
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
Metric3D
1.180
1.695
0.828
0.347
8.02
14.68
27.82
8.00
5.004
9.273
1.299
0.672
7.68
14.13
23.11
8.00
MoGe-3
0.208
0.292
0.236
0.084
18.89
35.49
59.89
6.00
0.158
0.275
0.228
0.112
12.63
26.82
58.20
3.00
Depth Pro
0.135
0.196
0.234
0.154
17.74
32.98
65.30
5.57
0.171
0.426
0.250
0.220
15.82
31.58
68.10
2.86
Unidepth
0.159
0.220
0.207
0.134
20.00
38.38
68.44
4.86
0.943
1.941
0.676
0.337
15.10
26.24
42.90
6.29
MoGe-2
0.176
0.242
0.204
0.086
23.40
40.93
63.08
4.71
0.202
0.355
0.258
0.206
15.95
30.21
57.03
3.43
Table 3: Quantitative results for image-level geometry estimation. The top-3 results are highlighted as first , second , and third , the same highlighting applies below.
Normal
Challenging
Model
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
MapAnything
0.295
0.404
0.312
0.061
9.14
18.79
37.66
5.57
0.560
1.750
0.440
0.190
4.90
8.40
43.40
4.43
MASt3R
0.126
0.188
0.207
0.088
16.26
31.90
62.48
4.57
1.036
2.206
0.677
0.515
11.47
18.59
38.26
5.57
pi3x
0.129
0.184
0.206
0.037
18.35
34.08
60.80
4.00
0.185
0.402
0.254
0.136
11.48
25.38
63.09
2.14
MUSt3R
0.120
0.174
0.193
0.095
22.71
38.96
67.50
3.29
0.996
2.116
0.671
0.497
10.69
19.36
41.99
4.86
CUT3R
0.110
0.155
0.156
0.090
22.65
42.39
76.12
2.57
0.330
0.646
0.342
0.310
17.49
33.85
63.27
2.71
Table 4: Quantitative results for video-level geometry estimation. pi3x and MapAnything also support depth-prior conditioning; their results are reported in Table 6 .
Normal
Challenging
Model
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
PromptDA
0.081
0.111
0.108
0.145
54.26
72.88
87.75
9.71
1.412
2.610
0.871
0.727
9.90
17.84
28.91
5.86
Marigold-dc
0.077
0.112
0.079
0.153
63.38
79.38
92.33
9.29
3.813
8.175
1.296
0.783
12.04
18.29
29.23
8.71
CDM-L515
0.046
0.057
0.056
0.050
64.14
86.99
97.10
7.00
2.569
7.541
0.759
0.672
26.43
38.93
50.59
3.00
CDM-D435
0.043
0.053
0.051
0.052
71.57
87.69
96.69
6.43
2.721
5.079
0.907
0.462
26.69
37.24
48.37
3.14
DepthLab
0.031
0.046
0.044
0.064
69.65
89.14
98.70
5.57
4.010
8.394
1.237
0.837
17.06
24.02
34.38
8.71
Table 5: Quantitative results for image-level geometry completion with sensor depth prior. We input the sensor depth as the prior without any preprocessing.
Normal
Challenging
Model
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
pi3x
0.097
0.131
0.139
0.062
26.43
44.76
80.46
5.00
2.811
5.781
1.079
0.911
8.42
15.91
37.83
4.43
MapAnything
0.028
0.040
0.038
0.053
77.44
93.67
98.76
4.00
3.334
7.172
1.117
0.839
20.96
28.61
39.64
4.14
CAPA-MoGe2
0.019
0.027
0.027
0.027
85.42
96.95
99.76
2.29
2.351
4.692
1.037
0.524
21.43
26.26
37.07
2.43
CAPA-Unidepthv2
0.018
0.027
0.027
0.027
86.15
97.25
99.76
1.86
2.522
4.99
1.041
0.607
25.69
31.94
41.97
2.57
CAPA-VGGT
0.016
0.022
0.022
0.028
91.44
98.71
99.89
1.29
2.352
4.811
0.974
0.604
31.48
38.14
44.87
1.43
Table 6: Quantitative results for video-level geometry completion with sensor depth prior. We input the sensor depth for every view as the prior without any preprocessing. “CAPA_” means CAPA framework supported by “” pretrained model.
Benchmark
Depth Pro
MoGe-2
MetricAnything
Unidepth
Unidepthv2
Depth benchmark
4.25
2.83
2.75
2.08
2.83
Intrinsics benchmark
2.89
1.00
1.50
3.46
3.83
Overall rank
3.57
1.92
2.13
2.77
3.33
Table 7: Average rank (Rank ↓ ) of each model on the depth benchmark (12 metrics) and the intrinsics benchmark (24 metrics). Bold : best; underlined : second best.
Figure 4: (a) Image-level geometry estimation (orange) and completion (blue) versus raw RealSense D435 depth: δ1.05 on normal scenes and RE on challenging scenes. Red dashed lines mark the sensor baseline; green regions indicate better performance. (b) Robustness of geometry estimation (purple) and completion (orange): challenging-to-normal RE ratios for image-level models (left), and δ1.05 on normal (x-axis) versus challenging (y-axis) scenes for multi-view models (right). Lower RE ratios indicate less degradation; points closer to the dashed y=x line indicate more consistent accuracy across scene difficulty. The suffix “-c” in MapAnything-c and pi3x-c denotes their geometry completion settings with depth priors.
Figure 5: (a) Image-level completion accuracy ( δ1.05 ) with 100, 1,000, or 10,000 randomly sampled depth points, or full sensor depth (Full). (b) Video-level completion RE as the fraction of views with depth priors increases from 0.1 to 1.0; shading indicates the gap between models. Enlarged markers in (a,b) highlight each model’s best setting. (c) RE versus RSD of predicted distances for representative image- and video-level estimation and completion models. Points closer to the lower-left corner indicate greater accuracy and cross-frame consistency.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Left : Overview of our data acquisition pipeline, including scene preparation, multi-view RGB-D capture, and the generation of raw data for subsequent annotation. Mid & Right : Overview of the coarse-to-fine data annotation pipeline. Valid frames are first sub-sampled from the raw image sequence, followed by anchor-frame annotation, automatic keypoint tracking, deviation-based re-annotation, and frame-by-frame verification to produce the final annotated sequence.
Input
Task
Models
Image-level
Estimation
Metric3D Yin et al. (2023) , Unidepth Piccinelli et al. (2024) , MoGe-2 Wang et al. () , MoGe-3 Kong et al. (2026) , Depth Pro Bochkovskiy et al. (2025) , MetricAnything Ma et al. (2026) , Metric3D v2 Hu et al. (2024) , Unidepthv2 Piccinelli et al. (2025)
Completion
CDM Liu et al. (2025) (CDM-L515, CDM-D435), DepthLab Liu et al. (2024) , PriorDA Wang et al. (2025c) , Lingbot-Depth Tan et al. (2026) , InfiniDepth Yu et al. (2026a) , Omni-dc Zuo et al. (2025) , LDCM Yu et al. (2026b) , Marigold-dc Viola et al. (2025) , PromptDA Lin et al. (2025)
Video-level
Estimation
MapAnything Keetha et al. (2026) , MASt3R Leroy et al. (2024) , MUSt3R Cabon et al. (2025) , Depth Anything 3 Lin et al. () , pi3x Wang et al. (2026a) , CUT3R Wang et al. (2025b)
Completion
MapAnything Keetha et al. (2026) , pi3x Wang et al. (2026a) , CAPA Ke et al. (2026) (CAPA-VGGT Wang et al. (2025a) , CAPA-Unidepthv2 Piccinelli et al. (2025) , CAPA-MoGe2 Wang et al. () )
Appendix
Table 8: Overview of the evaluated baselines. The evaluated CDM variants and CAPA backbones are listed in parentheses.
Model
AE ↓
RE ↓
LE ↓
RSD ↓
δ1.05↑
δ1.10↑
δ1.25↑
Rank ↓
Metric3D
1.319 ( +0.139 )
1.895 ( +0.200 )
0.893 ( +0.065 )
0.340 ( -0.007 )
5.16 ( -2.86 )
9.89 ( -4.79 )
23.42 ( -4.40 )
8.00
Depth Pro
0.144 ( +0.009 )
0.210 ( +0.014 )
0.240 ( +0.006 )
0.175 ( +0.021 )
16.53 ( -1.21 )
31.68 ( -1.30 )
62.77 ( -2.53 )
5.14
Unidepth
0.205 ( +0.046 )
0.287 ( +0.067 )
0.246 ( +0.039 )
0.121 ( -0.013 )
14.68 ( -5.32 )
27.22 ( -11.16 )
52.55 ( -15.89 )
6.86
MoGe-2
0.172 ( -0.004 )
0.229 ( -0.013 )
0.208 ( +0.004 )
0.095 ( +0.009 )
17.00 ( -6.40 )
31.50 ( -9.43 )
62.29 ( -0.79 )
4.86
MoGe-3
0.185 ( -0.023 )
0.253 ( -0.039 )
0.220 ( -0.016 )
0.093 ( +0.009 )
17.19 ( -1.70 )
33.68 ( -1.81 )
63.34 ( +3.45 )
4.43
MetricAnything
0.150 ( +0.005 )
0.195 ( +0.001 )
0.183 ( +0.007 )
0.098 ( +0.009 )
20.19 ( -4.81 )
37.34 ( -3.94 )
66.86 ( -0.10 )
3.29
Appendix
Table 9: Image-level geometry estimation using ground-truth intrinsics. For the seven evaluation metrics, values in parentheses denote changes relative to the corresponding normal-subset results in Table 3 ; green and red indicate improvements and degradations, respectively.
Model
NYU-D
KITTI
DIODE
ETH3D
AbsRel ↓
L1 (m) ↓
δ1.25↑
AbsRel ↓
L1 (m) ↓
δ1.25↑
AbsRel ↓
L1 (m) ↓
δ1.25↑
AbsRel ↓
L1 (m) ↓
δ1.25↑
Depth Pro
0.09
0.25
0.93
0.14
2.34
0.83
0.40
4.46
0.41
0.38
3.25
0.33
MoGe-2
0.08
0.21
0.96
0.21
3.59
0.45
0.33
2.62
0.54
0.10
0.62
0.88
MetricAnything
0.10
0.27
0.94
0.09
1.60
0.94
0.34
2.49
0.65
0.11
0.67
0.90
Unidepth
0.06
0.14
0.98
0.05
1.04
0.98
0.27
2.64
0.67
0.58
3.20
0.14
Unidepthv2
0.07
0.18
0.96
0.09
1.58
0.95
0.78
7.07
0.54
0.21
1.23
0.68
Appendix
Table 10: Depth benchmark results for image-level metric geometry estimation models. The best values are highlighted in green , and the second-best ones in yellow .
Table 11: Intrinsics benchmark results for image-level metric geometry estimation models. @1/5/10∘ refer to AUC@1/5/10∘ . vFoV, hFoV refer to vertical FoV and horizontal FoV. The best values are highlighted in green , and the second-best ones in yellow . Methods trained on evaluated datasets are in gray and excluded from the ranking to ensure a fair comparison
Figure 7: Effect of sampling stride on video-level geometry estimation models (normal subset only). Each curve corresponds to one model, and the highlighted marker on each curve indicates that model’s best-performing stride.
Figure 8: Qualitative comparison for image-level geometry estimation models between back-projecting with ground-truth intrinsics and models’ predicted intrinsics. Ground-truth annotations are shown in red , (ground-truth intrinsics + depth maps) are shown in blue , (predicted intrinsics + depth maps) are shown in yellow .
Figure 9: Qualitative comparison for image-level geometry completion models on different prior depth patterns with varying sparsity levels. The results showcase that Lingbot-Depth and PriorDA perform robustly under different patterns, while CDM-L515 exhibits a substantial performance drop due to the large pattern domain gap between its training data and the evaluated prior patterns.
Figure 10: Qualitative results for MapAnything Keetha et al. (2026) with varying fractions of views provided with depth priors. The distance of “Sensor Captured” is obtained by averaging the sensor measured distances of the ten views sampled. The results demonstrate that, as the ratio of depth priors increased, MapAnything gains continuous performance improvement. “Ratio#x” indicates the ratio of prior input to the model.
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data. We introduce Metric Anything, a simple and scalable pretraining framework that learns metric depth from noisy, diverse 3D sources without manually engineered prompts, camera-specific modeling, or task-specific architectures. Central to our approach is the Sparse Metric Prompt, created by randomly masking depth maps, which serves as a universal interface that decouples spatial reasoning from sensor and camera biases. Using about 20M image-depth pairs spanning reconstructed, captured, and rendered 3D data across 10000 camera models, we demonstrate-for the first time-a clear scaling trend in the metric depth track. The pretrained model excels at prompt-driven tasks such as depth completion, super-resolution and Radar-camera fusion, while its distilled prompt-free student achieves state-of-the-art results on monocular depth estimation, camera intrinsics recovery, single/multi-view metric 3D reconstruction, and VLA planning. We also show that using pretrained ViT of Metric Anything as a visual encoder significantly boosts Multimodal Large Language Model capabilities in spatial intelligence. These results show that metric depth estimation can benefit from the same scaling laws that drive modern foundation models, establishing a new path toward scalable and efficient real-world metric perception. We open-source MetricAnything at http://metric-anything.github.io/metric-anything-io/ to support community research.
3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.
Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-completion networks recover metric depth at the cost of geometric fidelity, cross-domain robustness, or speed. We present FounRef, a training-free method that aligns a frozen monocular foundation prior with sparse metric anchors to produce dense metric depth. FounRef is modular by design: its depth prior, anchor source, and refinement solver can each be replaced independently. We instantiate FounRef with MoGe-2 and LiDAR anchors. FounRef validates each anchor against the prior's dense depth prediction, rejecting inconsistencies caused by cross-sensor misalignment that geometry-only filters cannot detect. It then applies global and local metric corrections through a structure-preserving solver, retaining the prior's fine-grained geometry. FounRef requires no task-specific training and operates out of the box across unfamiliar cameras and scenes. On out-of-domain data, it delivers up to 24% lower depth error, 92% lower surface-normal noise, and almost 15x faster inference than DMD3C, a state-of-the-art depth-completion network. By decoupling metric alignment from geometry prediction, FounRef provides an accurate, geometrically faithful, and efficient approach to dense metric depth that can directly benefit from future advances in foundation models and metric sensors.