Equirectangular projection (ERP) is the standard representation for 360∘ imagery, and robust dense feature matching on ERP underpins panoramic stereo, view synthesis, and omnidirectional SLAM. Dense matchers trained on flat images degrade systematically on ERP because the chart introduces three coupled distortions -- topological, metric, and area -- that standard coarse matching and visibility estimation do not explicitly model. We show that correcting the three distortions at the coarse-stage interfaces where they arise -- pairwise distortions in attention, per-pixel distortion in covisibility gating -- improves PCK@1∘ from 0.229 to 0.275 on Matterport3D under a fixed coarse scaffold, with the refiner architecture unchanged -- our central result. Concretely, SCCM (Spherically Consistent Coarse Matching) augments a chart-naive cross-attention/dual-softmax coarse matcher with two sphere-derived priors: Spherical Positional Attention (SPA) pairs a yaw-periodic RoPE (topology) with a tangent-plane bias (metric), and Area-Aware Covisibility (AAC) applies a pre-sigmoid log-area correction (area). The chart-naive scaffold serves as a controlled reference, separating the scaffold-replacement effect from the spherical-prior effect. Instantiated in the RoMa V1 framework with the same frozen encoder, refiner architecture, and loss, SCCM also outperforms the ERP-native EDM (0.163) and an ERP-retrained RoMa V1 (0.198) under a unified ERP dense matching protocol, while perspective-trained matchers largely fail on ERP. It further transfers zero-shot to Stanford2D3D and, when trained on outdoor Holo360D, leads there as well.
Figures & tables
Figure 1 : SCCM coarse matcher: pipeline overview. SCCM injects sphere-derived priors into a fixed chart-naïve coarse scaffold (denoted R1 in Sec. 4 ), while keeping the encoder, refiner architecture, and loss unchanged. The inherited scaffold consists of cross-/self-attention, covisibility gating, and gated dual-softmax. The blue modules are the proposed sphere-aware priors: yaw-periodic RoPE (topology) and Tangent-Plane Bias (metric) in SPA (Fig. 2 ), and Log-Area Correction (area) in AAC (Fig. 3 ); orange boxes denote coarse matching signals passed to gated dual-softmax, grey boxes are inherited unchanged, and purple boxes are outputs. The output is a dense warp WA→B with per-pixel certainty c .
Figure 2 : Spherical Positional Attention (SPA). From coarse query/key tokens Q,K with sphere coordinates (φ,λ) , SPA forms two sphere-aware positional terms that merge at the pre-softmax attention logit. (top) Yaw-periodic RoPE makes the relative phase strictly 2π -periodic in longitude, so seam-neighbor tokens remain adjacent across the ERP boundary. (bottom) Tangent-Plane Bias, instantiating the CPB framework [ 18 ] on the sphere, replaces the planar chart offset Δq,kchart with the spherical log-map offset δq,k computed in the query tangent plane at rq , producing a geodesic-aware additive bias bq,k . The resulting logit is Aq,k=Aq,kRoPE+bq,k (Sec. 4.1 ).
Figure 3 : Area-Aware Covisibility (AAC). Given a per-pixel feature xp and latitude φp , AAC adds a latitude-dependent log-area term to the covisibility logit before the sigmoid: ℓpAAC=ℓp+αLAClogmax(cosφp,ϵ) . (left) the bare logit ℓp=MLP(xp) is area-unaware, whereas ERP pixels represent smaller spherical area near the poles; (middle) the log-area term is 0 at the equator and negative toward the poles; (right) after sigmoid, the corrected logit yields an area-aware gate that down-weights pixel-overrepresented polar candidates before dual-softmax matching, with learnable αLAC=softplus(αraw) (one per image side, init =1 ). See Sec. 4.2 .
Method
Train data
PCK@1 ∘↑
PCK@3 ∘↑
PCK@5 ∘↑
MAE ∘↓
Med. ∘↓
(a) Matterport3D test — in-distribution ( 15,682 pairs)
Perspective baselines (trained on perspective, zero-shot on ERP)
RoMa V1 [ 5 ]
ScanNet [ 6 ] (indoor, persp.)
0.023
0.058
0.081
57.46
49.74
RoMa V2 [ 4 ]
Mixed (incl. indoor, persp.)
0.040
0.126
0.201
36.10
18.09
ERP-native baselines (released checkpoints, not retrained)
SphereGlue † [ 11 ]
Custom Synth. (indoor, ERP)
0.028
0.072
0.095
64.50
65.45
Table 1 : Main full-system comparison on (a) Matterport3D, (b) zero-shot Stanford2D3D, and (c) outdoor Holo360D, where all three matchers are trained on Holo360D from their Matterport3D checkpoints under one protocol. All methods are evaluated end-to-end under the same ERP dense matching metrics. Retrained/controlled rows in (a, b) share the protocol of Sec. 5.1 , while external baselines use their released public weights. The mechanism-isolated contribution of the sphere-aware priors, measured as chart-naïve → SCCM under a fixed scaffold, is reported in Tab. 2 . Stanford2D3D uses the overlap range 0.30≤ov≤0.80 ; details in Supp. Sec. B. Best bold , second underlined . † SphereGlue [ 11 ] is sparse and not directly comparable to dense methods.
Figure 4 : Precision analysis on Matterport3D test (full test set, 15,682 pairs). (a) PCK@ 1∘ by absolute latitude ∣φ∣ : SCCM is consistently highest across latitude bands, with a larger margin toward the poles (shaded). (b) PCK@ 1∘ by absolute longitude ∣λ∣ : SCCM shows a smaller drop near the ERP seam ( ∣λ∣=180∘ , shaded) than EDM, while the chart-naïve scaffold lies between EDM and SCCM.
Figure 5 : Qualitative comparison on Matterport3D test. Each method panel shows the predicted warp masked by correctness: white denotes inaccurate or non-covisible pixels. The reported percentage is the fraction of GT-covisible pixels with angular error below 1∘ . SCCM recovers larger accurate regions than the dense baselines (EDM, RoMa V1 (ERP-retrained)) and the chart-naïve scaffold across representative overlap levels, including high-latitude ceiling/floor regions.
PCK ∘ ( ↑ )
Error ∘ ( ↓ )
Row
@1
@3
@5
MAE
Med.
Primary role
(R1) chart-naïve (no PE)
0.229
0.609
0.769
5.68
2.23
planar scaffold
ctrl: EDM abs. PE §
0.224
0.593
0.756
5.97
2.31
(input-PE control)
ctrl: standard RoPE §
0.236
0.613
0.764
6.79
2.17
(rel-PE control)
(R2a) + yaw-periodic RoPE
0.267 ( +3.8 )
0.653
0.796
5.87 ( +0.19 )
1.93
strict precision
(R2b) + TPB (full SPA)
0.266 ( −0.1 )
0.652
0.799
5.58 ( −0.29 )
1.95
error-tail (MAE)
Table 2 : Controlled ablation on the fixed chart-naïve scaffold (Matterport3D test). R2a–R3 cumulatively add yaw-periodic RoPE, TPB, and LAC/AAC to the no-prior scaffold R1; the two § rows are positional-encoding controls on R1. Values in parentheses on PCK@ 1∘ and MAE denote the marginal change from the previous chain row (PCK in pp, MAE in degrees), revealing the modules’ functional roles: yaw-periodic RoPE drives strict precision, TPB reduces the angular-error tail, and LAC adds area-gated precision. Chain marginals on PCK@ 1∘ fold the TPB × LAC interaction into the final row (factorial decomposition in Supp. Sec. M). Protocol as in Sec. 5.1 . Best in bold .
Pose AUC ↑
Acc (m) ↓
Comp (m) ↓
Recon. F ↑
Method
5∘
10∘
20∘
Mean
Med.
Mean
Med.
@5
@10
@20
EDM [ 6 ]
7.97
17.86
29.51
1.452
0.470
0.685
0.336
14.5
27.5
42.7
RoMa V1 (retr.) [ 5 ]
11.12
27.26
47.42
0.884
0.294
0.573
0.250
16.0
30.7
48.2
chart-naïve
13.66
31.50
51.76
0.787
0.267
0.512
0.236
17.4
32.7
50.6
SCCM (ours)
17.41
35.84
55.09
0.772
0.247
0.496
0.218
19.3
35.5
53.2
Table 3 : Downstream geometry on Matterport3D under a unified evaluator. We report relative-pose AUC and certainty-gated reconstruction metrics. Reconstruction triangulates matches under the GT pose with a fixed certainty gate ( τ=0.5 ) for all methods; both tasks are evaluated on the 11,574 test pairs where every method yields ≥200 confident matches; details in Supp. Sec. C. Best bold , second underlined .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Optimizer
AdamW ( β1=0.9 , β2=0.999 , wd 0.01 ); non-zero LR only on trainable modules
LR schedule
base 1×10−4 , constant, ×0.1 at step 112,500 ( 90% of the 125K steps)
spherical warp from per-pixel depth + relative pose (intrinsics-free ERP)
MP3D split
official 61/11/18 scenes =54,015/4,625/15,682 pairs; native 1024×512
Appendix
Table 1 : Shared training protocol for all models retrained on Matterport3D, used for the fixed-scaffold ablations.
Figure 1 : Top-down reconstruction comparison on three representative test pairs (qualitative; aggregate reconstruction metrics are in Tab. 3, main). Each panel projects the triangulated cloud onto the X–Z plane (camera A at origin). On these examples the baselines (EDM, RoMa V1 (ERP-retrained)) and our chart-naïve scaffold produce distorted room contours, whereas SCCM’s cloud (red) aligns more closely with the GT reference (green).
Method
AUC@ 5∘
AUC@ 10∘
AUC@ 20∘
EDM [ 6 ]
11.52
24.55
38.17
RoMa V1 (ERP-retrained) [ 5 ]
14.69
33.60
54.28
chart-naïve
17.56
37.68
57.97
SCCM
22.06
42.36
61.17
Appendix
Table 2 : Pose on the overlap-selected subset ( ov>0.5 , 9,268 pairs) under our unified ray-essential evaluator. We report Pose-AUC using max(rotation,translation) error. This is not an EDM-official reproduction. Best bold , second underlined .
Set
Method
PCK@ 1∘↑
PCK@ 3∘↑
PCK@ 5∘↑
MAE ↓
Med ↓
MP3D
chart-naïve
0.2299
0.6086
0.7674
5.667
2.225
SCCM
0.2748
0.6653
0.8059
5.360
1.875
Δ
+.045
+.057
+.039
−.307
−.350
95% CI
[+.042,+.048]
[+.053,+.061]
[+.035,+.042]
[−.42,−.19]
[−.40,−.30]
Stanford2D3D
chart-naïve
0.1789
0.4746
0.6267
18.86
3.250
SCCM
0.2291
0.5637
0.6974
17.09
2.450
Appendix
Table 3 : Pair-level bootstrap ( 10,000 paired resamples) on Matterport3D test ( 15,682 pairs) and zero-shot Stanford2D3D ( 8,744 pairs). PCK is fraction-of-pixels ( ↑ ); MAE/median in degrees ( ↓ ). Every Δ (SCCM − chart-naïve) 95% CI excludes zero, and SCCM wins all 10,000/10,000 resamples for every metric on both datasets. Best in bold . Values are computed at full precision from per-pair records collected in a separate evaluation pass; entries and Δ s may differ from the main-paper tables in the last reported digits on both datasets (aggregation and evaluation-pass differences).
PCK@ 1∘↑
PCK@ 3∘↑
PCK@ 5∘↑
MAE ↓
Med ↓
Matterport3D test
EDM
pixel-uniform
0.163
0.400
0.509
16.78
4.79
area-weighted
0.186
0.419
0.518
17.16
4.55
RoMa V1 (ERP-retrained)
pixel-uniform
0.198
0.538
0.711
6.60
2.76
area-weighted
0.238
0.565
0.716
6.84
2.44
chart-naïve
pixel-uniform
0.229
0.609
0.769
5.68
2.23
Appendix
Table 4 : Sphere-area-weighted metrics : each pixel weighted by cosφ vs the pixel-uniform grid, on Matterport3D test and zero-shot Stanford2D3D. SCCM leads all baselines (EDM, ERP-retrained RoMa V1, the chart-naïve scaffold) on every metric under both pixel-uniform and sphere-area-weighted aggregation, on both datasets. The pixel-uniform columns reproduce the main-paper tables; the area-weighted columns re-weight the same test pixels by cosφ .
Method
upright (resampled)
pitch 10∘
pitch 20∘
pitch 30∘
EDM [ 6 ]
0.163
0.105
0.044
0.018
RoMa V1 (ERP-retrained) [ 5 ]
0.198
0.096
0.042
0.022
chart-naïve
0.230
0.099
0.043
0.023
SCCM (ours)
0.273
0.142
0.064
0.034
Appendix
Table 5 : Off-gravity tilt robustness on Matterport3D test ( 15,682 pairs). A fixed SO(3) pitch is applied to both views and the relative pose is conjugated accordingly, so the ground truth remains consistent up to resampling artifacts. We report PCK@ 1∘ under increasing pitch. All ERP matchers degrade off-gravity; SCCM is not tilt-equivariant but retains the highest absolute accuracy at every tested angle. Best per column in bold .
ϵ
PCK@ 1∘↑
PCK@ 3∘↑
PCK@ 5∘↑
MAE ∘↓
Med ∘↓
10−4
0.2748
0.6653
0.8059
5.360
1.875
10−3 (used)
0.2748
0.6653
0.8059
5.360
1.875
10−2
0.2748
0.6653
0.8059
5.360
1.875
Appendix
Table 6 : Sensitivity to the numerical safety clamp ϵ used in Eq. (5) of the main paper, on the full Matterport3D test split ( 15,682 pairs) with the full SCCM model. Varying ϵ over [10−4,10−2] leaves all metrics unchanged at reported precision, confirming that the clamp is not driving the result. PCK ↑ ; MAE/median in degrees ↓ .
Model
αLACA/αLACB
PCK@ 1∘↑
PCK@ 3∘↑
PCK@ 5∘↑
MAE ∘↓
SCCM (learned α )
0.978/1.028
0.2748
0.6653
0.8059
5.360
SCCM ( α=1 , infer.)
1.000/1.000
0.2748
0.6654
0.8059
5.360
Appendix
Table 7 : LAC strength sanity check on the full Matterport3D test split. The trained SCCM checkpoint is evaluated with its learned per-view LAC strengths (αLACA,αLACB) and with α fixed to 1 at inference time . The two settings yield metrics identical to within 10−4 , supporting that LAC’s contribution comes from placing the analytic log-area term before the sigmoid rather than from scalar tuning.
Configuration
Trainable params
GFLOPs
(R1) chart-naïve (cross-attn + DS)
137.17 M
751.15
+ SPA (RoPE + TPB)
+≈1.1 k
≈9.0 ( ≈0.1 cached)
+ AAC (LAC)
+2
≈0.0
SCCM total
137.17 M
≈760.2 ( ≈751.3 cached)
Increase vs. baseline
<0.002%
≈1.2% ( ≈0.01% cached)
Appendix
Table 8 : Computational footprint relative to the R1 chart-naïve baseline: trainable-parameter and FLOP additions at the medium ERP resolution ( 448×896 ). Inference latency is reported separately in Tab. 9 .
Method
Latency (ms)
Overhead
chart-naïve
115.76
—
SCCM (TPB uncached)
139.95
+20.90%
SCCM (TPB cached)
121.60
+5.04%
Appendix
Table 9 : TPB bias caching latency. Wall-clock latency on a single RTX 4090 at medium ERP resolution, averaged over 100 timed forwards after warm-up. The cached variant precomputes the content-independent TPB bias for the fixed ERP grid and reuses it at inference time.
Figure 2 : Band-wise marginal decomposition of R1 → R2a → R2b → R3 (Tab. 2, main). Bars show each module’s marginal effect: blue + yaw-periodic RoPE (R2a − R1); orange + TPB (R2b − R2a); green + LAC/AAC (R3 − R2b). Left column: Δ PCK@ 1∘ (pp); right column: Δ MAE (deg). Top row: latitude bands ∣φ∣ ; bottom row: longitude bands ∣λ∣ . Positive Δ PCK and negative Δ MAE indicate improvement.
Row
PCK@ 1∘↑
MAE ↓
P50 ↓
P75 ↓
P90 ↓
P95 ↓
P99 ↓
R2a (RoPE)
0.2667
5.87
1.92
4.18
9.29
19.49
94.23
R2b (full SPA)
0.2662
5.58
1.94
4.13
8.88
17.69
89.27
Δ
−0.05
−0.29
+0.02
−0.05
−0.41
−1.80
−4.96
Appendix
Table 10 : TPB error-tail quantiles on Matterport3D test (full split, ∼3.0 B valid pixels). R2a is yaw-periodic RoPE only; R2b adds TPB (full SPA). Angular-error quantiles P k are in degrees; in the Δ row, PCK@ 1∘ is in percentage points and all other columns in degrees. The Δ PCK@ 1∘ of −0.05 pp is computed from full-precision values before rounding (the body Tab. 2 rounds the two rows to 0.267 and 0.266 , i.e. −0.1 pp). TPB leaves strict precision (PCK@ 1∘ ) and the median (P50) nearly unchanged but lowers the high-error quantiles, confirming its role as tail calibration.
Cell
TPB
LAC
PCK@ 1∘↑
PCK@ 3∘↑
PCK@ 5∘↑
MAE ↓
R2a (RoPE)
✗
✗
0.2679
0.6558
0.7987
5.789
R2b ( + TPB)
✓
✗
0.2662
0.6521
0.7989
5.576
+ LAC (no TPB)
✗
✓
0.2664
0.6583
0.8017
5.527
R3 (SCCM)
✓
✓
0.2748
0.6653
0.8059
5.360
Effect (95% CI)
Δ PCK@ 1∘ (pp)
Δ MAE (deg)
TPB alone (R2b − R2a)
−0.17[−0.37,+0.04]
−0.21[−0.33,−0.10]
Appendix
Table 11 : TPB × LAC factorial on Matterport3D test ( 15,682 pairs): all four combinations over the R2a (yaw-periodic RoPE) base under the shared protocol of Tab. 1 . Effects are paired-bootstrap estimates ( 10,000 resamples). All four cells are evaluated in a single evaluation pass; the body table and Tab. 10 stem from earlier passes, and repeated evaluation of the same checkpoint can shift the last reported digits through floating-point nondeterminism (observed up to ∼10−3 in PCK@ 1∘ and ∼0.1∘ in MAE). The body table’s final-row marginal ( +0.9 pp) corresponds to LAC given TPB here ( +0.86 pp), i.e., the interaction plus LAC’s solo effect. Neither prior improves strict precision alone; their combination does.
Figure 3 : SCCM’s off-gravity failure mode , shown on a single Matterport3D test pair (the same pair throughout). Top : input view A upright vs. under a synthetic 30∘ off-gravity pitch (Sec. 0.F ), which tips the ERP chart off the gravity horizon. Bottom : SCCM’s per-pixel angular error for the warp A→B (black = non-covisible; colorbar 0∘ to ≥30∘ ). Upright, SCCM attains PCK@ 1∘=0.37 on this pair (mostly low error); under tilt the gravity-aligned prior no longer matches the tilted chart and accuracy drops markedly (PCK@ 1∘=0.18 ), with error concentrated where the chart distortion is largest. Holding the pair fixed isolates tilt—not scene content—as the cause; full-set tilt averages are in Tab. 5 . This is a geometry-specific failure outside the intended upright-ERP setting.
Stage
chart-naïve
SCCM
Δ (pp)
coarse anchor ( 1/16 )
0.122
0.163
+4.1
refiner stage ( 1/8 )
0.133
0.187
+5.4
refiner stage ( 1/4 )
0.188
0.234
+4.6
refiner stage ( 1/2 )
0.219
0.263
+4.4
final dense
0.229
0.275
+4.6
Appendix
Table 12 : Refinement-stage PCK@ 1∘ diagnostic on the full Matterport3D test split (no masking/object filtering). PCK@ 1∘ after the coarse stage and after each refinement stage. SCCM’s advantage is present before refinement and preserved through the architecturally unchanged RoMa V1 refiner; the similar coarse-to-final lift for both models indicates that the gain enters primarily through coarse-anchor formation rather than being created by the refiner.
Method
PCK@ 0.35∘↑
PCK@ 0.5∘↑
PCK@ 1∘↑
PCK@ 3∘↑
PCK@ 5∘↑
MAE ∘↓
Med. ∘↓
RoMa V1 (ERP-retrained)
0.123
0.182
0.322
0.592
0.726
5.66
2.12
chart-naïve (R1)
0.119
0.180
0.331
0.626
0.761
5.02
1.93
SCCM
0.137
0.202
0.357
0.645
0.776
4.94
1.76
Appendix
Table 13 : Outdoor training on Holo360D ( 8,000 test pairs). All three matchers start from their Matterport3D checkpoints and are trained under one protocol; SCCM leads on every metric, including the two error statistics. Best in bold .
Method
Matterport3D
Stanford2D3D
DKM [ 3 ] (CVPR’23)
0.035
0.026
MASt3R † [ 7 ] (ECCV’24)
0.057
0.028
LoFTR † [ 12 ] (CVPR’21)
0.105
0.058
VGGT [ 13 ] (CVPR’25)
0.000
0.000
EDM [ 6 ] (reference)
0.163
0.104
Appendix
Table 14 : Additional zero-shot baselines (PCK@ 1∘ ). † Sparse matchers scored on their own matches only. EDM from Tab. 1 for reference.
Configuration
seed 1 (paper)
seed 2
seed 3
s.d. (pp)
chart-naïve (R1)
0.229
0.230
0.223
0.4
SCCM
0.275
0.283
0.272
0.6
margin (pp)
+4.6
+5.3
+4.9
Appendix
Table 15 : Multi-seed retraining of the ablation endpoints (Matterport3D test, PCK@ 1∘ ). Paper seed first.
Method
8-pt + RANSAC
Robust 360-8PA
EDM [ 6 ]
8.28
8.29
RoMa V1 (ERP-retrained)
11.69
11.68
chart-naïve (R1)
14.17
14.16
SCCM
17.96
17.96
Appendix
Table 16 : Solver swap on 15,112 Matterport3D test pairs: Pose AUC@ 5∘ with our 8-point + RANSAC estimator vs. Robust 360-8PA on identical matches.
The equirectangular projection (ERP) is widely used for panoramic depth estimation, but its spatially varying distortion makes geometry-consistent feature modeling challenging. We revisit panoramic depth estimation by decoupling contextual modeling in native spherical space from dense ERP prediction. To this end, we propose a Fibonacci Spherical Graph (FSG) as an intermediate reasoning space to lift ERP features onto quasi-uniform Fibonacci nodes on the sphere and capture local and long-range dependencies through complementary spherical neighborhoods. The resulting spherical discretization distributes graph nodes approximately uniformly over the spherical surface, reducing the over-representation of highly stretched regions during relational modeling. Operating on a compact set of Fibonacci nodes also avoids the computational burden of constructing and processing a graph at full ERP resolution. To bridge spherical reasoning and dense prediction, we propose a Spherical Context Conditioning (SCC) module that adaptively modulates dense ERP features with the enhanced spherical representation, allowing spherical context to guide pixel-aligned depth prediction. Extensive experiments on three benchmarks demonstrate that the proposed method consistently achieves superior depth accuracy over existing approaches.
Zhijie Shen, Chunyu Lin, Shuai Zheng +4
Beijing Jiaotong University · Hefei University of Technology · Shandong University
Dense feature matching aims to estimate all correspondences between two images of a 3D scene and has recently been established as the gold standard due to its high accuracy and robustness. However, existing dense matchers still fail or perform poorly for many hard real-world scenarios, and high-precision models are often slow, limiting their applicability. In this paper, we attack these weaknesses on a wide front through a series of systematic improvements that together yield a significantly better model. In particular, we construct a novel matching architecture and loss, which, combined with a curated diverse training distribution, enables our model to solve many complex matching tasks. We further make training faster through a decoupled two-stage matching-then-refinement pipeline, and at the same time, significantly reduce refinement memory usage through a custom CUDA kernel. Finally, we leverage the recent DINOv3 foundation model along with multiple other insights to make the model more robust and unbiased. In our extensive set of experiments, we show that the resulting novel matcher sets a new state-of-the-art, being significantly more accurate than its predecessors. Code is available at https://github.com/Parskatt/romav2
Johan Edstedt, David Nordström, Yushan Zhang +7
Linköping University · Chalmers University of Technology · University of Amsterdam +1
Omnidirectional stereo images provide full-surround perception but violate the geometric assumptions of classical disparity estimation: in spherical or fisheye views, epipolar correspondences follow curved great-circle paths, producing two-dimensional displacements that cannot be treated as single-axis disparity before geometric rectification. In this work, we adopt a standard spherical-to-equirectangular (ERP) projection as a preprocessing step, which straightens epipolar curves and restores a one-dimensional disparity structure - horizontal for left-right rigs and vertical for top-bottom rigs. Building on our previously introduced RAFT + Epipolar-Aligned Channel Selection (EACS) framework, originally developed for rectilinear and ERP stereo, we examine whether the same modular pipeline remains accurate when the input originates from spherical stereo imagery. After ERP projection, dense optical flow from RAFT is reduced to disparity by retaining only the baseline-aligned flow component. Experiments on synthetic fisheye stereo datasets show that this spherical-to-ERP-to-RAFT+EACS pipeline produces accurate, smooth, and structurally consistent disparity maps at real-time speed. These findings confirm that established ERP preprocessing can be effectively combined with our earlier RAFT+EACS method to enable practical, interpretable, and efficient disparity estimation from spherical stereo, providing a straightforward pathway for extending conventional stereo pipelines to 360 imaging.
Sahereh Obeidavi, Dieter Landes
Faculty of Electrical Engineering and Computer Science, Coburg University of Applied Science, German