Panoramic depth estimation captures the complete 360∘ scene geometry, being essential for robotics and AR/VR applications. While perspective depth models have achieved remarkable zero-shot generalization via large-scale training, panoramic methods lag behind, especially for open-world scenes, due to data scarcity. To bridge this gap, we introduce DA360, a panoramic-adapted version of Depth Anything V2. Our key insight is that the base DAV2 model, trained on perspective images to predict affine-invariant disparity, already exhibits good zero-shot performance on panoramas. Building on this, we design a lightweight adaptation framework that (i) learns a per-image shift from the ViT class token with scale-invariant supervision, transforming affine-invariant disparity into scale-invariant disparity that directly yields well-formed 3D point clouds, and (ii) integrates circular padding into the DPT decoder to eliminate seam artifacts, ensuring spatial coherence. Fine-tuned on a combination of synthetic indoor and outdoor panoramic data, DA360 is evaluated on standard real-world indoor benchmarks and our newly curated outdoor dataset, Metropolis. Results show that DA360 not only outperforms the original DAV2 by over 50% and 12% relative error reduction indoors and outdoors, but also surpasses prior specialized methods like PanDA by about 25--35% across all tests, establishing state-of-the-art zero-shot panoramic depth estimation.
Figures & tables
Figure 3: Comparisons with SOTA zero-shot monocular depth models on outdoor panoramic images of SUN360.
Figure 4: Our framework is a generalizable panoramic depth estimation model, which fine-tunes a zero-shot pinhole depth model with synthetic panorama depth datasets to produce scale-invariant and boundary-consistent panoramic depth maps.
Models
Backbone
Matterport3D
Metropolis
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
Marigold-v1.1 ( Ke et al. 2025 )
SD 2.0
0.2097
65.11
0.7357
26.87
DAV2 ( Yang et al. 2024b )
ViT-S
0.2032
64.56
0.2169
66.34
DAV2-ind ( Yang et al. 2024b )
0.2170
63.98
0.5667
24.89
DAV2-out ( Yang et al. 2024b )
0.2695
54.18
0.4274
34.71
DAV2 ( Yang et al. 2024b )
ViT-B
0.1966
66.03
0.2061
67.64
Table 1: Evaluating pinhole models on panorama datasets.
Figure 5: Circular padding for equirectangular representation, divided into two phases of vertical and horizontal directions.
Models
Matterport3D
Stanford2D3D
Metropolis
AbsRel ↓
MAE ↓
RMSE ↓
RMSE log ↓
δ1↑
AbsRel ↓
MAE ↓
RMSE ↓
RMSE log ↓
δ1↑
AbsRel ↓
MAE ↓
RMSE ↓
RMSE log ↓
δ1↑
DreamCube ( Huang et al. 2025 )
0.2736
0.5201
0.7258
0.1364
56.60
0.2592
0.4130
0.5773
0.1282
61.24
0.4928
15.872
20.682
0.3036
36.16
DAC (L) ( Guo et al. 2025 )
0.1442
0.3185
0.5013
0.0806
82.52
0.1261
0.2166
0.3495
0.0718
86.23
0.5524
17.068
21.748
0.2337
34.09
UniK3D (L) ( Piccinelli et al. 2025 )
0.1046
0.2133
0.3906
0.0645
92.17
0.1011
0.1584
0.2706
0.0607
91.88
0.4842
15.687
20.847
0.2170
33.82
DA 2 (L) ( Li et al. 2025 )
0.1032
0.2256
0.4207
0.0658
89.49
0.0688
0.1165
0.2508
0.0485
95.55
0.3895
11.658
16.131
0.1833
45.82
DAV2 (S) ( Yang et al. 2024b )
0.2032
0.4553
0.7566
0.1128
64.56
0.2561
0.4907
0.7906
0.1416
55.74
0.2169
9.4807
16.057
0.1426
66.34
Table 2: Quantitative comparison. DA360 and DA 2 are evaluated with scale alignment, while all other methods use affine alignment. The postfixes (S), (B), and (L) denote the backbone scales, i.e., ViT-Small, ViT-Base, and ViT-Large, respectively.
Figure 6: Analysis of the curated Metropolis test set. (a) Camera positions colored by elevation, showing 2.2 km × 1.9 km spatial coverage in Detroit. (b) Histogram of semantic category counts per image; the average is 24 and the range is [14,36] , indicating varying scene complexity.
Figure 8: Zero-shot qualitative comparison on Matterport3D ( Chang et al. 2017 ) , Stanford2D3D ( Armeni et al. 2017 ) , Metropolis ( Mapillary 2024 ) , and SUN360 ( Xiao et al. 2012 ) .
Figure 10: Effectiveness of shift learning and loss configuration.
Figure 11: Comparison of w/o and w/ circular padding.
Models
Structured3D
Deep360
Matterport3D
Metropolis
w/o circular padding
0.0226
0.0958
0.1153
0.2661
w/ circular padding
0.0210
0.0871
0.1132
0.2273
Table 4: AbsRel on a 5-pixel ERP boundary margin.
Figure S3: Typical problematic samples in Metropolis: a) depth value overflow due to encoding limitations; b) overly sparse valid depth; c) apparent depth errors. Val: ratio of valid depth points; Min: minimum depth value; Max: maximum depth value.
Figure S5: Six examples of the curated Metropolis test set.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S6: Comparison of disparity distributions between TartanAir V2 ( Patel et al. 2025 ) and Structured3D ( Zheng et al. 2020 ) : (a) histogram of disparity map median values; (b) histogram of disparity map median absolute deviation (MAD) values; (c) and (d) distributions of median and MAD computed from 5,000 randomly selected disparity maps from TartanAir V2 and Structured3D, respectively.
Figure S8: Visual comparisons of different model parameter initialization strategies: ( Random ), DINOv2 , and DAV2 (corresponding to the model SI in the main paper).
Model
Pretrained Models
Structured3D
Deep360
Matterport3D
Metropolis
Names
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
DAV2
Depth Anything V2 ( Yang et al. 2024b )
0.0263
0.0988
99.07
0.0879
6.2732
89.34
0.1004
0.5047
91.02
0.1966
14.587
72.78
DINOv2
DINOv2 ( Oquab et al. 2023 )
0.0348
0.1218
98.36
0.0967
6.9283
87.60
0.1158
0.5611
88.09
0.2421
17.059
62.97
Random
None/Random
0.1444
0.3278
82.01
0.5039
29.6632
37.06
0.3000
1.0870
50.83
0.5266
17.675
32.57
Appendix
Table S1: Ablation study on different model parameter initialization strategies.
Models
Pre-trained
Finetuning
Evaluation
Matterport3D
Stanford2D3D
Metropolis
Models
Methods
Alignment
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
PanDA(S)*
DAV2-ind
PanDA
scale+shift
0.1346
0.5090
83.59
0.0957
0.3261
90.80
0.2898
15.037
57.52
PanDA(S) †
DAV2-out
PanDA
scale+shift
0.1414
0.5328
83.26
0.0926
0.3400
92.87
0.2994
15.362
56.90
PanDA(S) ∘
DAV2
PanDA
scale+shift
0.1319
0.5028
85.20
0.0930
0.3293
91.75
0.2517
14.197
64.15
DA360(S)*
DAV2-ind
DA360
scale
0.1076
0.5780
89.48
0.0754
0.3569
93.95
0.1990
14.673
72.77
DA360(S) †
DAV2-out
DA360
scale
0.1059
0.5358
90.33
0.0739
0.3210
94.73
0.2465
15.065
67.02
Appendix
Table S2: Different pre-trained models and finetuning methods.
Models
Matterport3D
Metropolis
1 ×
2 ×
1 ×→ 2 ×
1 ×
2 ×
1 ×→ 2 ×
DA360 (S)
0.0963
0.1259
0.0948
0.1922
0.2332
0.1919
DA360 (B)
0.0877
0.1209
0.0861
0.1798
0.2269
0.1783
DA360 (L)
0.0793
0.1075
0.0783
0.1630
0.2054
0.1622
Appendix
Table S3: AbsRel at different resolution settings.
Models
Inference
Run
Evaluation
Matterport3D
Stanford2D3D
Metropolis
Style
Time
Alignment
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
MoGe-2(L)
end2end
0.101s
scale+shift
0.2920
0.7827
54.65
0.3731
0.9519
45.65
0.6764
23.152
30.84
MoGe-2(L)
crop&fuse
196s
scale+shift
0.0835
0.3519
93.89
0.1134
0.2984
86.15
0.2141
12.644
72.04
DAV2(L)
crop&fuse
152s
scale+shift
0.1195
0.5704
86.34
0.1491
0.5633
78.90
0.2230
16.201
62.66
DA360(L)
end2end
0.026s
scale
0.0793
0.4407
94.07
0.0710
0.2930
93.53
0.1630
13.015
80.47
Appendix
Table S4: Comparison with crop&fuse baselines.
Figure S9: Additional visualizations on SUN360: six examples with depth maps and point clouds from our DA360.
Figure S11: Visual comparison on example I of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S13: Visual comparison on example II of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S15: Visual comparison on example III of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S17: Visual comparison on example IV of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S19: Visual comparison on example V of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S21: Visual comparison on example VI of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S23: Visual comparison on example VII of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S25: Visual comparison on example VIII of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S27: Visual comparison on example IX of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S29: Visual comparison on example X of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Recent depth foundation models trained on perspective imagery achieve strong performance, yet generalize poorly to 360∘ images due to the substantial geometric discrepancy between perspective and panoramic domains. Moreover, fully fine-tuning these models typically requires large amounts of panoramic data. To address this issue, we propose RePer-360, a distortion-aware self-modulation framework for monocular panoramic depth estimation that adapts depth foundation models while preserving powerful pretrained perspective priors. Specifically, we design a lightweight geometry-aligned guidance module to derive a modulation signal from two complementary projections (i.e., ERP and CP) and use it to guide the model toward the panoramic domain without overwriting its pretrained perspective knowledge. We further introduce a Self-Conditioned AdaLN-Zero mechanism that produces pixel-wise scaling factors to reduce the feature distribution gap between the perspective and panoramic domains. In addition, a cubemap-domain consistency loss further improves training stability and cross-projection alignment. By shifting the focus from complementary-projection fusion to panoramic domain adaptation under preserved pretrained perspective priors, RePer-360 surpasses standard fine-tuning methods while using only 1% of the training data. Under the same in-domain training setting, it further achieves an approximately 20% improvement in RMSE. The code is available at https://github.com/munimo/RePer360.
Cheng Guan, Chunyu Lin, Zhijie Shen +2
Institute of Information Science, Beijing Jiaotong University, Beijing 100044, China. · Visual Intelligence +X International Cooperation Joint Laboratory of MOE, Beijing 100044, China.
While monocular depth estimation has achieved significant progress, achieving generalized metric depth estimation for both narrow field-of-view (FoV) perspectives and 360∘ panoramas remains an unsolved challenge. Existing methods are often tailored to specific camera types and struggle to produce accurate metric depth that generalizes across diverse settings. This limitation stems from two key challenges: the inherent geometric discrepancy between perspective and panoramic cameras, and the scarcity of panoramic training data with metric annotations. In this work, we introduce DepthMaster, a unified metric depth estimation framework. Rather than employing specialized networks to learn spherical distortions, we reformulate the problem by decomposing panoramic images into overlapping perspective patches. Crucially, distinct from prior projection-based methods that rely on ad-hoc architectural modifications to handle boundaries, we introduce a novel Correspondence Consistency Loss (CCL) and inject virtual projection cameras as geometric priors, allowing us to seamlessly stitch the patches while avoiding specialized operators and keeping the backbone largely compatible with standard Transformer designs. This strategy also resolves the geometric differences by unifying all inputs into a canonical perspective representation, and effectively circumvents data scarcity by directly unlocking powerful metric priors from vast perspective datasets. Trained on a mixed dataset that contains only one panorama dataset, DepthMaster achieves state-of-the-art zero-shot performance on 13 diverse datasets, outperforming not only universal methods but also leading specialist models in both perspective and panoramic domains.
Pengfei Wang, Shihao Wang, Liyi Chen +3
Visual Computing Lab The Hong Kong Polytechnic University
Geometry estimation from perspective images has greatly advanced, maturing to the point where off-the-shelf foundation models are able to reconstruct 3D scene structure not only from multi-view imagery, but even from a single view. A natural extension is 3D reconstruction from panoramas, with the exciting prospect of recovering a full 360-degree scene from a single panoramic image. In this work, we introduce PaGeR (Panoramic Geometry Reconstruction), a framework to lift powerful 3D foundation models designed for perspective imagery to the panorama domain. Our strategy is to start from a pre-trained transformer for 3D reconstruction and turn it into a unified high-performance model that predicts scale-invariant depth, metric depth, surface normals, and sky masks from both perspective and omnidirectional images, in a single forward pass. By keeping architectural changes to a minimum and mixing perspective and panoramic images during training, PaGeR retains the rich 3D prior of the underlying foundation model while learning to also estimate geometrically consistent 360-degree scenes from single panoramas. We extensively test our method in both indoor and outdoor environments and find that it delivers state-of-the-art performance and excellent zero-shot performance across a wide range of scenes. Code, data and models are available \href.