cs.CVFeb 26, 2026

SO3UFormer: Learning Intrinsic Spherical Features for Rotation-Robust Panoramic Dense Prediction

Authors: Qinfeng ZhuYunxi JiangLei Fan

Organizations: University of Liverpool, United Kingdom. · CNRS, France. · Xi’an Jiaotong-Liverpool University, China.

Abstract

Panoramic dense-prediction models, spanning semantic segmentation and depth estimation, are typically trained under a strict gravity-aligned assumption. Real-world captures, however, routinely violate it: handheld devices jitter and aerial platforms change attitude, so the camera is rarely upright. Under such 3D reorientation, standard spherical Transformers overfit global latitude cues and collapse. We introduce SO3UFormer, an architecture that learns intrinsic spherical features largely decoupled from the underlying coordinate frame, through three geometric components: (1) removing absolute latitude encoding, which breaks the dependence on the gravity axis; (2) quadrature-consistent spherical attention, which corrects for non-uniform sampling density; and (3) a gauge-aware relative positional bias built from local tangent-plane angles rather than global axes. A logit-space \emph{SO(3)}-consistency regularizer, used only during training, further suppresses residual discretization effects. To benchmark robustness, we introduce Pose35, a variant of Stanford2D3D perturbed by random rotations within ±35\pm 35^\circ, and evaluate under a full, arbitrary \emph{SO(3)} stress test. There, the baseline SphereUFormer collapses from 67.53 \emph{mIoU} on Pose35 to 25.26 under the full \emph{SO(3)} test, whereas SO3UFormer reaches 72.03 on Pose35 and retains 70.67 under the same test. Similarly, on a second real-world dataset (Matterport3D) for segmentation and on panoramic depth estimation, SO3UFormer remains essentially rotation-invariant while the gravity-anchored baseline again loses most of its accuracy. Code and models are available at https://github.com/zhuqinfeng1999/SO3UFormer.

Explore similar work

May 25, 2026cs.CV

Unified Panoramic Geometry Estimation via Multi-View Foundation Models

Geometry estimation from perspective images has greatly advanced, maturing to the point where off-the-shelf foundation models are able to reconstruct 3D scene structure not only from multi-view imagery, but even from a single view. A natural extension is 3D reconstruction from panoramas, with the exciting prospect of recovering a full 360-degree scene from a single panoramic image. In this work, we introduce PaGeR (Panoramic Geometry Reconstruction), a framework to lift powerful 3D foundation models designed for perspective imagery to the panorama domain. Our strategy is to start from a pre-trained transformer for 3D reconstruction and turn it into a unified high-performance model that predicts scale-invariant depth, metric depth, surface normals, and sky masks from both perspective and omnidirectional images, in a single forward pass. By keeping architectural changes to a minimum and mixing perspective and panoramic images during training, PaGeR retains the rich 3D prior of the underlying foundation model while learning to also estimate geometrically consistent 360-degree scenes from single panoramas. We extensively test our method in both indoor and outdoor environments and find that it delivers state-of-the-art performance and excellent zero-shot performance across a wide range of scenes. Code, data and models are available \href\href{https://github.com/prs-eth/PaGeR}{\text{here}}.
Vukasin Bozic, Isidora Slavkovic, Dominik Narnhofer +4
May 13, 2026cs.CV

PanoWorld: Towards Spatial Supersensing in 360^\circ Panorama World

Multimodal large laboratory models (MLLMs) still struggle with spatial understanding under the dominant perspective-image paradigm, which inherits the narrow field of view of human-like perception. For navigation, robotic search, and 3D scene understanding, 360-degree panoramic sensing offers a form of supersensing by capturing the entire surrounding environment at once. However, existing MLLM pipelines typically decompose panoramas into multiple perspective views, leaving the spherical structure of equirectangular projection (ERP) largely implicit. In this paper, we study pano-native understanding, which requires an MLLM to reason over an ERP panorama as a continuous, observer-centered space. To this end, we first define the key abilities for pano-native understanding, including semantic anchoring, spherical localization, reference-frame transformation, and depth-aware 3D spatial reasoning. We then build a large-scale metadata construction pipeline that converts mixed-source ERP panoramas into geometry-aware, language-grounded, and depth-aware supervision, and instantiate these signals as capability-aligned instruction tuning data. On the model side, we introduce PanoWorld with Spherical Spatial Cross-Attention, which injects spherical geometry into the visual stream. We further construct PanoSpace-Bench, a diagnostic benchmark for evaluating ERP-native spatial reasoning. Experiments show that PanoWorld substantially outperforms both proprietary and open-source baselines on PanoSpace-Bench, H* Bench, and R2R-CE Val-Unseen benchmarks. These results demonstrate that robust panoramic reasoning requires dedicated pano-native supervision and geometry-aware model adaptation. All source code and proposed data will be publicly released.
Changpeng Wang, Xin Lin, Junhan Liu +5
Mar 6, 2026cs.CV

RePer-360: Releasing Perspective Priors for 360^\circ Depth Estimation via Self-Modulation

Recent depth foundation models trained on perspective imagery achieve strong performance, yet generalize poorly to 360^\circ images due to the substantial geometric discrepancy between perspective and panoramic domains. Moreover, fully fine-tuning these models typically requires large amounts of panoramic data. To address this issue, we propose RePer-360, a distortion-aware self-modulation framework for monocular panoramic depth estimation that adapts depth foundation models while preserving powerful pretrained perspective priors. Specifically, we design a lightweight geometry-aligned guidance module to derive a modulation signal from two complementary projections (i.e., ERP and CP) and use it to guide the model toward the panoramic domain without overwriting its pretrained perspective knowledge. We further introduce a Self-Conditioned AdaLN-Zero mechanism that produces pixel-wise scaling factors to reduce the feature distribution gap between the perspective and panoramic domains. In addition, a cubemap-domain consistency loss further improves training stability and cross-projection alignment. By shifting the focus from complementary-projection fusion to panoramic domain adaptation under preserved pretrained perspective priors, RePer-360 surpasses standard fine-tuning methods while using only 1% of the training data. Under the same in-domain training setting, it further achieves an approximately 20% improvement in RMSE. The code is available at https://github.com/munimo/RePer360.
Cheng Guan, Chunyu Lin, Zhijie Shen +2