Panoramic depth estimation captures the complete 360∘ scene geometry, being essential for robotics and AR/VR applications. While perspective depth models have achieved remarkable zero-shot generalization via large-scale training, panoramic methods lag behind, especially for open-world scenes, due to data scarcity. To bridge this gap, we introduce DA360, a panoramic-adapted version of Depth Anything V2. Our key insight is that the base DAV2 model, trained on perspective images to predict affine-invariant disparity, already exhibits good zero-shot performance on panoramas. Building on this, we design a lightweight adaptation framework that (i) learns a per-image shift from the ViT class token with scale-invariant supervision, transforming affine-invariant disparity into scale-invariant disparity that directly yields well-formed 3D point clouds, and (ii) integrates circular padding into the DPT decoder to eliminate seam artifacts, ensuring spatial coherence. Fine-tuned on a combination of synthetic indoor and outdoor panoramic data, DA360 is evaluated on standard real-world indoor benchmarks and our newly curated outdoor dataset, Metropolis. Results show that DA360 not only outperforms the original DAV2 by over 50% and 12% relative error reduction indoors and outdoors, but also surpasses prior specialized methods like PanDA by about 25--35% across all tests, establishing state-of-the-art zero-shot panoramic depth estimation.
Figures & tables
Figure 3: Comparisons with SOTA zero-shot monocular depth models on outdoor panoramic images of SUN360.
Figure 4: Our framework is a generalizable panoramic depth estimation model, which fine-tunes a zero-shot pinhole depth model with synthetic panorama depth datasets to produce scale-invariant and boundary-consistent panoramic depth maps.
Models
Backbone
Matterport3D
Metropolis
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
Marigold-v1.1 ( Ke et al. 2025 )
SD 2.0
0.2097
65.11
0.7357
26.87
DAV2 ( Yang et al. 2024b )
ViT-S
0.2032
64.56
0.2169
66.34
DAV2-ind ( Yang et al. 2024b )
0.2170
63.98
0.5667
24.89
DAV2-out ( Yang et al. 2024b )
0.2695
54.18
0.4274
34.71
DAV2 ( Yang et al. 2024b )
ViT-B
0.1966
66.03
0.2061
67.64
Table 1: Evaluating pinhole models on panorama datasets.
Figure 5: Circular padding for equirectangular representation, divided into two phases of vertical and horizontal directions.
Models
Matterport3D
Stanford2D3D
Metropolis
AbsRel ↓
MAE ↓
RMSE ↓
RMSE log ↓
δ1↑
AbsRel ↓
MAE ↓
RMSE ↓
RMSE log ↓
δ1↑
AbsRel ↓
MAE ↓
RMSE ↓
RMSE log ↓
δ1↑
DreamCube ( Huang et al. 2025 )
0.2736
0.5201
0.7258
0.1364
56.60
0.2592
0.4130
0.5773
0.1282
61.24
0.4928
15.872
20.682
0.3036
36.16
DAC (L) ( Guo et al. 2025 )
0.1442
0.3185
0.5013
0.0806
82.52
0.1261
0.2166
0.3495
0.0718
86.23
0.5524
17.068
21.748
0.2337
34.09
UniK3D (L) ( Piccinelli et al. 2025 )
0.1046
0.2133
0.3906
0.0645
92.17
0.1011
0.1584
0.2706
0.0607
91.88
0.4842
15.687
20.847
0.2170
33.82
DA 2 (L) ( Li et al. 2025 )
0.1032
0.2256
0.4207
0.0658
89.49
0.0688
0.1165
0.2508
0.0485
95.55
0.3895
11.658
16.131
0.1833
45.82
DAV2 (S) ( Yang et al. 2024b )
0.2032
0.4553
0.7566
0.1128
64.56
0.2561
0.4907
0.7906
0.1416
55.74
0.2169
9.4807
16.057
0.1426
66.34
Table 2: Quantitative comparison. DA360 and DA 2 are evaluated with scale alignment, while all other methods use affine alignment. The postfixes (S), (B), and (L) denote the backbone scales, i.e., ViT-Small, ViT-Base, and ViT-Large, respectively.
Figure 6: Analysis of the curated Metropolis test set. (a) Camera positions colored by elevation, showing 2.2 km × 1.9 km spatial coverage in Detroit. (b) Histogram of semantic category counts per image; the average is 24 and the range is [14,36] , indicating varying scene complexity.
Figure 8: Zero-shot qualitative comparison on Matterport3D ( Chang et al. 2017 ) , Stanford2D3D ( Armeni et al. 2017 ) , Metropolis ( Mapillary 2024 ) , and SUN360 ( Xiao et al. 2012 ) .
Figure 10: Effectiveness of shift learning and loss configuration.
Figure 11: Comparison of w/o and w/ circular padding.
Models
Structured3D
Deep360
Matterport3D
Metropolis
w/o circular padding
0.0226
0.0958
0.1153
0.2661
w/ circular padding
0.0210
0.0871
0.1132
0.2273
Table 4: AbsRel on a 5-pixel ERP boundary margin.
Figure S3: Typical problematic samples in Metropolis: a) depth value overflow due to encoding limitations; b) overly sparse valid depth; c) apparent depth errors. Val: ratio of valid depth points; Min: minimum depth value; Max: maximum depth value.
Figure S5: Six examples of the curated Metropolis test set.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S6: Comparison of disparity distributions between TartanAir V2 ( Patel et al. 2025 ) and Structured3D ( Zheng et al. 2020 ) : (a) histogram of disparity map median values; (b) histogram of disparity map median absolute deviation (MAD) values; (c) and (d) distributions of median and MAD computed from 5,000 randomly selected disparity maps from TartanAir V2 and Structured3D, respectively.
Figure S8: Visual comparisons of different model parameter initialization strategies: ( Random ), DINOv2 , and DAV2 (corresponding to the model SI in the main paper).
Model
Pretrained Models
Structured3D
Deep360
Matterport3D
Metropolis
Names
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
DAV2
Depth Anything V2 ( Yang et al. 2024b )
0.0263
0.0988
99.07
0.0879
6.2732
89.34
0.1004
0.5047
91.02
0.1966
14.587
72.78
DINOv2
DINOv2 ( Oquab et al. 2023 )
0.0348
0.1218
98.36
0.0967
6.9283
87.60
0.1158
0.5611
88.09
0.2421
17.059
62.97
Random
None/Random
0.1444
0.3278
82.01
0.5039
29.6632
37.06
0.3000
1.0870
50.83
0.5266
17.675
32.57
Appendix
Table S1: Ablation study on different model parameter initialization strategies.
Models
Pre-trained
Finetuning
Evaluation
Matterport3D
Stanford2D3D
Metropolis
Models
Methods
Alignment
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
PanDA(S)*
DAV2-ind
PanDA
scale+shift
0.1346
0.5090
83.59
0.0957
0.3261
90.80
0.2898
15.037
57.52
PanDA(S) †
DAV2-out
PanDA
scale+shift
0.1414
0.5328
83.26
0.0926
0.3400
92.87
0.2994
15.362
56.90
PanDA(S) ∘
DAV2
PanDA
scale+shift
0.1319
0.5028
85.20
0.0930
0.3293
91.75
0.2517
14.197
64.15
DA360(S)*
DAV2-ind
DA360
scale
0.1076
0.5780
89.48
0.0754
0.3569
93.95
0.1990
14.673
72.77
DA360(S) †
DAV2-out
DA360
scale
0.1059
0.5358
90.33
0.0739
0.3210
94.73
0.2465
15.065
67.02
Appendix
Table S2: Different pre-trained models and finetuning methods.
Models
Matterport3D
Metropolis
1 ×
2 ×
1 ×→ 2 ×
1 ×
2 ×
1 ×→ 2 ×
DA360 (S)
0.0963
0.1259
0.0948
0.1922
0.2332
0.1919
DA360 (B)
0.0877
0.1209
0.0861
0.1798
0.2269
0.1783
DA360 (L)
0.0793
0.1075
0.0783
0.1630
0.2054
0.1622
Appendix
Table S3: AbsRel at different resolution settings.
Models
Inference
Run
Evaluation
Matterport3D
Stanford2D3D
Metropolis
Style
Time
Alignment
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
AbsRel ↓
RMSE ↓
δ1↑
MoGe-2(L)
end2end
0.101s
scale+shift
0.2920
0.7827
54.65
0.3731
0.9519
45.65
0.6764
23.152
30.84
MoGe-2(L)
crop&fuse
196s
scale+shift
0.0835
0.3519
93.89
0.1134
0.2984
86.15
0.2141
12.644
72.04
DAV2(L)
crop&fuse
152s
scale+shift
0.1195
0.5704
86.34
0.1491
0.5633
78.90
0.2230
16.201
62.66
DA360(L)
end2end
0.026s
scale
0.0793
0.4407
94.07
0.0710
0.2930
93.53
0.1630
13.015
80.47
Appendix
Table S4: Comparison with crop&fuse baselines.
Figure S9: Additional visualizations on SUN360: six examples with depth maps and point clouds from our DA360.
Figure S11: Visual comparison on example I of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S13: Visual comparison on example II of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S15: Visual comparison on example III of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S17: Visual comparison on example IV of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S19: Visual comparison on example V of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S21: Visual comparison on example VI of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S23: Visual comparison on example VII of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S25: Visual comparison on example VIII of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S27: Visual comparison on example IX of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Figure S29: Visual comparison on example X of SUN360 ( Xiao et al. 2012 ) with SOTA depth estimation models.
Institute of Information Science, Beijing Jiaotong University, Beijing 100044, China. · Visual Intelligence +X International Cooperation Joint Laboratory of MOE, Beijing 100044, China.