Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-completion networks recover metric depth at the cost of geometric fidelity, cross-domain robustness, or speed. We present FounRef, a training-free method that aligns a frozen monocular foundation prior with sparse metric anchors to produce dense metric depth. FounRef is modular by design: its depth prior, anchor source, and refinement solver can each be replaced independently. We instantiate FounRef with MoGe-2 and LiDAR anchors. FounRef validates each anchor against the prior's dense depth prediction, rejecting inconsistencies caused by cross-sensor misalignment that geometry-only filters cannot detect. It then applies global and local metric corrections through a structure-preserving solver, retaining the prior's fine-grained geometry. FounRef requires no task-specific training and operates out of the box across unfamiliar cameras and scenes. On out-of-domain data, it delivers up to 24% lower depth error, 92% lower surface-normal noise, and almost 15x faster inference than DMD3C, a state-of-the-art depth-completion network. By decoupling metric alignment from geometry prediction, FounRef provides an accurate, geometrically faithful, and efficient approach to dense metric depth that can directly benefit from future advances in foundation models and metric sensors.
Figures & tables
Figure 1: Accuracy is not the only axis. Metric error against surface dispersion on KITTI ( n=500 frames, identical anchors for every method); marker area grows with runtime. A single configuration change moves our point to 0.663 m at 11.2∘ without affecting runtime
Figure 2: Global calibration of four priors on the same frame. Top: raw predictions use incompatible depth scales; DA-V2 is relative. Bottom: after fitting the same sparse anchors, all priors share a metric scale while preserving their edges and relative layout. All depth panels use the 0 – 80m range. Pixels a prior marks invalid (sky) are drawn at the far end (MoGe-2 provides an invalid map, while depth pro predicts them at 10,000m ).
Figure 3: Anchor filtering on a cluttered KITTI crop. RePLAy retains returns through windshields, mirrors, and moving-vehicle silhouettes, whereas FounRef rejects them because they disagree with the locally refined smooth surface D1 .
Figure 4: Solver comparison on a sparse-anchor nuScenes crop. All methods start from the same calibrated prior and anchors. Pixel-wise Poisson and TGV produce anchor-centred surface bumps, whereas FounRef shares corrections on a coarse bilateral grid. Depth panels share the [1,50]m range.
Figure 5: One solver, four operating points on nuScenes ( 0 – 50m ). σs is the bilateral grid’s spatial bandwidth in pixels; reducing it localizes corrections, improving RMSE but increasing normal dispersion.
KITTI
nuScenes
Waymo
GOOSE
NYU
RMSE ↓
far ↓
surf. ↓
RMSE ↓
far ↓
surf. ↓
RMSE ↓
far ↓
surf. ↓
RMSE ↓
far ↓
surf. ↓
RMSE ↓
far ↓
surf. ↓
ms ↓
Ours
0.726
2.08
10.8
0.687
2.00
3.5
0.745
1.43
1.7
1.687
4.42
8.8
0.174
0.40
6.6
78
Ours ( σs 4)
0.663
1.83
11.2
0.489
1.27
3.6
0.587
1.10
2.2
1.073
3.24
8.9
0.164
0.37
6.7
78
OMNI-DC
1.137
3.61
45.8
0.353
0.87
34.6
0.590
1.12
41.7
0.702
2.01
33.1
0.099
0.20
10.2
1011
Prior-DA
1.583
4.83
41.4
0.661
2.19
16.6
0.833
1.69
13.2
0.742
2.03
15.8
0.135
0.25
13.0
1621
PromptDA
2.237
7.32
11.9
2.058
7.41
8.8
2.770
6.48
5.8
1.545
4.91
11.0
0.218
0.40
11.1
580
Table 1: Depth completion across five validation sets ( n=500 ) using identical anchors for every method and no dataset-specific recalibration. KITTI and NYU use dense ground truth; the other datasets use a held-out 20% of the cloud. RMSE and far (RMSE over target depths 30<z≤50m ) are in metres, surf. is normal dispersion ( ∘ ), and ms is mean latency on the Waymo validation set, our worst-case setting, on an A100 (most models cannot run at fp16 precision).
solver
RMSE ↓
far ↓
surf. ↓
imp. →1
ms ↓
Ours
1.050
2.583
1.77 ∘
0.38
14
Screened Poisson
0.880
1.904
4.75 ∘
5.25
42
TGV 2 -L2
0.934
1.949
6.32 ∘
—
300
Table 2: Solver comparison on KITTI val against accumulated dense depth. All solvers receive identical prior, calibration, anchors, and weights; only the regularizer differs. RMSE : error over all valid pixels; far : RMSE in the 30 – 50 m band. surf. : median angle between the output normals and calibrated prior normals in smooth regions. imp. : correction at anchor pixels relative to a few pixels away, so 1.0 leaves no trace of the anchor pattern.
Figure 6: Qualitative bake-off across five datasets. Metric accuracy doesn’t stand on its own, without appropriate runtime and coherent structure.
accuracy
structure
prior
RMSE ( ↓ )
AbsRel ( ↓ )
RMSE >50m ( ↓ )
incoh ( ↓ )
detail ( ↑ )
ms ( ↓ )
MoGe-2
2.83
0.029
14.57
0.007
0.025
49
Metric3D-V2
2.76
0.028
14.41
0.008
0.004
77
DA-V3
2.98
0.034
15.00
0.020
0.020
107
DA-V2
3.05
0.037
15.44
0.050
0.112
55
UniDepth-V2
3.59
0.047
17.28
0.020
0.007
49
Table 3: Prior selection (KITTI val. , n=100 , A100, fp16). All rows: FounRef refined. Scored to 80 m so each prior’s bias beyond the 50 m trust cap is visible; anchors stop at 50 m, so the > 50 m column measures unsupported extrapolation. Structure is measured in log-depth on LiDAR-derived masks: incoh = high-frequency spectral fraction on flat surfaces, detail = the same on RGB-textured regions. Read them together: low+low = washed, high detail + low incoh = clean and sharp.
set
filter
drop
vs. GT
vs. D1
ms
(%)
err. (D/K)
×
err. (D/K)
×
single
Ours
2.3
1.08 / 0.04
27
1.10 / 0.03
37
13
RePLAy
4.5
0.49 / 0.04
12
0.43 / 0.04
11
452
Oracle
2.5
0.95 / 0.03
32
0.64 / 0.04
16
—
temporal
Ours
2.2
0.80 / 0.04
20
0.88 / 0.03
29
13
RePLAy
2.2
0.26 / 0.04
7
0.27 / 0.05
6
868
Table 4: Anchor-level filter comparison on KITTI validation ( n=100 ). Errors are relative errors at dropped/kept anchors against dense GT or the independent reference D1 . Higher drop/keep ratios indicate better separation. Runtime is the filter’s own cost on one A100 (fp16). Temporal is over 10 sweeps.The oracle measures the difference between the raw and filtered point clouds, as processed by the Authors of KITTI.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: The pipeline on one nuScenes frame, by row: left the state after a step, right what that step changed (blue nearer, red farther, grey unchanged, each panel on its own scale). Top: anchors in, and what survives the anchor filter ( 4.7% removed), dilated for legibility.
stage
kitti
waymo
goose
RMSE (m)
calibrated prior
1.407
2.522
3.268
+D1 (filter reference)
0.698
0.959
1.101
+D2 (full method)
0.567
0.826
1.131
RMSE 30–50 m
calibrated prior
4.220
4.423
8.713
Appendix
Table 5: Pipeline ablation on the accumulated, unfiltered cloud. Anchors and held-out target are screened by the method’s own MoGe reference; no competitor appears here, so these are inference numbers rather than the cross-method protocol of the main tables.
Figure 8: Anchor-density sweep.
method
100%
50%
25%
10%
5%
1%
kitti
ours
0.832
0.865
0.906
0.969
1.024
1.150
OMNI-DC
1.095
0.887
0.663
0.708
0.796
0.991
Prior-Depth-Anything
1.667
0.760
0.753
0.841
0.921
1.167
DMD3C
0.548
0.536
0.533
0.563
0.661
1.596
nuscenes
Appendix
Table 6: Anchor-density sweep: RMSE (m) as a fraction of the returns is withheld. Every method receives the identical thinned cloud on each frame, drawn once per frame and shared, and is scored against the same target. 100% is the temporally accumulated cloud: 135351 anchors per frame on KITTI ( 31.6% of all pixels), 35815 on nuScenes ( 2.5% ), 109119 on Waymo ( 4.4% ) and 102247 on GOOSE ( 4.2% ).
σs
KITTI
nuScenes
Waymo
GOOSE
NYU
RMSE
disp.
RMSE
disp.
RMSE
disp.
RMSE
disp.
RMSE
disp.
16 (default)
0.832
12.01
0.968
6.94
0.917
4.85
2.254
9.35
0.269
7.53
12
0.794
11.88
0.908
6.84
0.859
4.80
2.043
9.29
0.264
7.53
8
0.736
11.71
0.826
6.72
0.791
4.77
1.798
9.18
0.258
7.53
6
0.700
11.60
0.776
6.65
0.748
4.78
1.651
9.18
0.255
7.53
4
0.656
11.49
0.713
6.61
0.704
4.85
1.513
9.36
0.253
7.53
Appendix
Table 7: Solver operating points: the bilateral cell σs swept at constant λ , over 50 frames per dataset. RMSE in metres and normal dispersion in degrees side by side, so the accuracy–structure trade of each setting reads across a row.
variant
KITTI
nuScenes
Waymo
GOOSE
NYU
ms
RMSE
disp.
RMSE
disp.
RMSE
disp.
RMSE
disp.
RMSE
disp.
default configuration
0.832
12.01
0.968
6.94
0.917
4.85
2.254
9.35
0.269
7.53
35
full-resolution solve
0.793
11.52
1.021
6.58
0.931
4.62
2.622
9.35
0.282
7.50
40
CG 2 iterations
0.868
11.25
1.195
6.54
1.074
4.46
3.744
9.69
0.293
7.48
27
CG 16 iterations
0.831
12.16
0.961
7.27
0.916
5.02
2.327
9.44
0.270
7.52
42
bistochastic 8 steps
0.832
11.97
0.967
6.92
0.917
4.84
2.253
9.20
0.269
7.53
35
Appendix
Table 8: Solver components, each toggled alone from the default configuration, over 50 frames per dataset; RMSE in metres and normal dispersion in degrees side by side, ‘ms’ the mean solve time. Full-resolution solve drops the half-resolution grid. CG iterations is the conjugate-gradient budget of the final solve. Bistochastic normalisation is the grid balancing that makes the bilateral operator well conditioned. A component is worth keeping only if switching it off costs more than the runtime it saves.
perturbation
kitti
nuscenes
waymo
goose
nyu
projection shift
0 px
0.832
0.968
0.917
2.254
0.269
1 px
0.911
0.969
0.932
2.251
0.269
2 px
1.042
1.003
0.972
2.254
0.270
3 px
1.219
1.043
1.033
2.264
0.272
5 px
1.624
1.151
1.178
2.293
0.274
Appendix
Table 9: Anchor perturbation, ours. Only the anchors are corrupted; the evaluation target is the unperturbed held-out cloud (dense GT on KITTI/NYU), so the numbers measure damage to the prediction rather than a moved target.
Figure 9: Solver sensitivity, two frames per dataset: the anchors given to the solver, depth at the default σs=16 , at σs=4 , and their signed difference (blue nearer, red farther, grey unchanged, each panel on its own scale). At σs=4 the correction concentrates into blobs at individual anchors instead of travelling along surfaces — the 21 – 33% RMSE it buys is paid for in surface damage. NYU is omitted: its two settings differ by under a centimetre.
Figure 10: Predicted depth for every evaluated method on one frame of each dataset, alongside the input image. Frames are drawn at random from each dataset, except GOOSE, where a forest track was chosen over the open field the draw returned. Our sky and other pixels where the monocular prior is undefined are drawn at the depth cap, since we emit no prediction there. The supervised networks hold up on the driving scenes that resemble their training data and come apart away from them, most visibly indoors. On the driving rows every method returns a plausible map — and on those same frames their normal dispersion spans 1.7 to 56 degrees (Figure 11 ). What separates these methods is not visible in this representation.
Figure 11: Surface normals of the same predictions, with each panel’s RMSE (m) and normal dispersion (deg) beneath it. The completion networks imprint the sensor into the surface: the ruling visible across the road is the LiDAR sweep pattern of the anchor column reproduced as geometry. The point metric does not register it — DMD3C and BP-Net are both more accurate than we are on the KITTI frame while carrying seven to eight times the normal error, because reproducing a measurement exactly is rewarded by RMSE whether or not the surface between measurements survives. We estimate a smooth metric correction over a prior we do not modify, so the anchor pattern cannot enter the geometry in the first place.
Figure 12: The same depth maps, turned back into 3D. For each scene the upper row is the input image and each method’s depth map; the lower row is the anchors the solver was given, drawn over the image, and each depth map re-projected into a coloured point cloud, all seen from a direction the camera never occupied ( 35∘ azimuth, 18∘ elevation). From the original viewpoint all four look alike by construction, since each is consistent with the same image. The depth maps are indeed hard to separate — but only ours re-projects to continuous surfaces. On Waymo the baselines smear the buildings and trees behind the road into sheets of stray points that hang over the scene; on NYU, where OMNI-DC and Prior-DA are metrically ahead of us, the room still resolves for all three and only DMD3C disintegrates. This is what the normal-dispersion column measures, and what a downstream consumer of the depth map receives.
Dense and accurate depth estimation is essential for robotic manipulation, grasping, and navigation, yet currently available depth sensors are prone to errors on transparent, specular, and general non-Lambertian surfaces. To mitigate these errors, large-scale monocular depth estimation approaches provide strong structural priors, but their predictions can be potentially skewed or mis-scaled in metric units, limiting their direct use in robotics. Thus, in this work, we propose a training-free depth grounding framework that anchors monocular depth estimation priors from a depth foundation model in raw sensor depth through factor graph optimization. Our method performs a patch-wise affine alignment, locally grounding monocular predictions in metric real-world depth while preserving fine-grained geometric structure and discontinuities. To facilitate evaluation in challenging real-world conditions, we introduce a benchmark dataset with dense scene-wide ground truth depth in the presence of non-Lambertian objects. Ground truth is obtained via matte reflection spray and multi-camera fusion, overcoming the reliance on object-only CAD-based annotations used in prior datasets. Extensive evaluations across diverse sensors and domains demonstrate consistent improvements in depth performance without any (re-)training. We make our implementation publicly available at https://anchord.cs.uni-freiburg.de.
Simon Dorer, Martin Büchner, Nick Heppert +1
Department of Computer Science, University of Freiburg, Germany. · 2Zuse School ELIZA
Monocular depth estimation (MDE) typically produces depth estimations that are defined up to an unknown scale or shift. When only sparse metric anchors are available, recovering accurate metric depth becomes challenging yet necessary for practical applications. We address this problem by formulating metric depth recovery as image-adaptive scale field modeling. Instead of directly correcting the depth, we reformulate the correction as a low-dimensional linear combination of image-adaptive basis maps. These maps are derived from semantic and geometric cues encoded in the MDE estimations and intermediate representations. The weights of basis maps are efficiently determined from sparse metric anchors via a least-squares problem. This formulation yields improved metric depth accuracy, strong robustness under extreme anchor sparsity, and an interpretable decomposition of spatial scale variations. Extensive experiments across multiple datasets and representative MDE models demonstrate the effectiveness and general applicability of our approach.
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data. We introduce Metric Anything, a simple and scalable pretraining framework that learns metric depth from noisy, diverse 3D sources without manually engineered prompts, camera-specific modeling, or task-specific architectures. Central to our approach is the Sparse Metric Prompt, created by randomly masking depth maps, which serves as a universal interface that decouples spatial reasoning from sensor and camera biases. Using about 20M image-depth pairs spanning reconstructed, captured, and rendered 3D data across 10000 camera models, we demonstrate-for the first time-a clear scaling trend in the metric depth track. The pretrained model excels at prompt-driven tasks such as depth completion, super-resolution and Radar-camera fusion, while its distilled prompt-free student achieves state-of-the-art results on monocular depth estimation, camera intrinsics recovery, single/multi-view metric 3D reconstruction, and VLA planning. We also show that using pretrained ViT of Metric Anything as a visual encoder significantly boosts Multimodal Large Language Model capabilities in spatial intelligence. These results show that metric depth estimation can benefit from the same scaling laws that drive modern foundation models, establishing a new path toward scalable and efficient real-world metric perception. We open-source MetricAnything at http://metric-anything.github.io/metric-anything-io/ to support community research.