Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery
Authors: Niels Sombekke, Rob G. J. Wijnhoven, Martin R. Oswald
Organizations: University of Amsterdam (UvA), Amsterdam, The Netherlands · Spotr, The Hague, The Netherlands
Abstract
We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared DINOv2 backbone. The design naturally handles a variable number of street-level views per building without padding or fixed-size pooling, and jointly predicts multi-label roof element and roof material classes. We construct a large-scale dataset of 32,135 buildings (61,672 segments) spanning ten countries, pairing satellite images with up to eight street-level views per segment and evaluating four masking strategies for isolating the target building. We propose an RGB-M masking strategy that appends the building footprint mask as a fourth input channel, providing a soft spatial prior that outperforms hard cropping across both modalities. The Perceiver IO fusion model improves over all other fusion strategies and yields substantial per-class gains for attributes visible from street level (e.g., +11.3 AP for slate, +1.3 AP for dormers), though the satellite-only baseline retains a slight advantage in macro-averaged mAP for classes that are predominantly visible from above. These results establish a scalable, flexible architecture for multi-modal building inspection that can accommodate heterogeneous inputs and multiple output tasks.
Current cross-view localization methods predominantly rely on satellite imagery as the aerial modality. Although recent work explores planimetric maps (e.g., OpenStreetMap tiles), these approaches often lag in performance. Yet both modalities are widely available and possess complementary properties. Satellite images are closer to ground-level camera imagery, offering finer detail, whereas planimetric maps contain annotated objects (e.g., streetlamps) and remain informative in areas where the ground is occluded, such as by foliage. Despite this, only one prior work provides an end-to-end method to fuse the two modalities, and it does not demonstrate their potential within state-of-the-art methods. To combine the strengths of both modalities, we propose a new fusion module that augments standard encoders and demonstrates that integrating satellite imagery with planimetric maps improves state-of-the-art single-modality methods. The module comprises (i) cross-modal conditioning, which processes each modality's encoding with awareness of the other, and (ii) a patch-level fusion rule that controls the granularity of information exchange. We achieve state-of-the-art results, reducing the mean localization error by 30.13%. Qualitatively, the fusion adaptively selects the more informative modality, improving overall accuracy.
Monocular building height estimation from optical imagery is important for characterizing urban vertical structure, yet remains challenging due to the heterogeneity of urban building morphology and the indirect relationship between optical image appearance and building height. The recently launched PhiSat-2 satellite provides a promising open-access data source for this task, with 4.75m spatial resolution and seven multispectral bands spanning the visible to near-infrared range. However, its suitability for monocular building height estimation has not been systematically assessed. This study presents an initial open-reference assessment of PhiSat-2 imagery for this task by constructing a PhiSat-2--Height Dataset (PHDataset) and proposing a Two-Stream Ordinal Network (TSONet). PHDataset integrates global PhiSat-2 imagery with open building-height references and contains 9,475 co-registered patch pairs from 26 cities worldwide. TSONet jointly learns dense height estimation and auxiliary footprint prediction, using footprint-aware structural guidance and ordinal height modeling to better exploit PhiSat-2 spatial--spectral information. Specifically, a Cross-Stream Exchange Module (CSEM) enables adaptive interaction between the height and footprint streams, while a Feature-Enhanced Bin Refinement (FEBR) module performs coarse-to-fine ordinal query refinement with multi-level features. Experiments on PHDataset show that TSONet outperforms representative competing methods, reducing MAE and RMSE by over 13.2% and 9.7%, respectively, while improving IoU and F1-score by over 14.0% and 10.1%. Additional analyses further indicate that PhiSat-2 imagery contains useful spatial--spectral cues for monocular building height estimation at an intermediate spatial resolution.
Sentinel-2 imagery offers open access, global coverage, and frequent revisit times, making it attractive for practical building mapping at scale; however, its native 10m resolution makes building vs non-building classification challenging, particularly for small or sub-pixel buildings, and performance can vary with both seasonality and the heterogeneity of built-up environments. This paper introduces a Sentinel-2 building-detection framework designed to systematically quantify these effects and to support more formalised, practice-oriented model selection. We construct a dedicated multi-temporal Sentinel-2 dataset over the Warsaw region and derive binary ground-truth masks by rasterising official Polish topographic database (BDOT10k) building footprints onto the Sentinel-2 pixel grid. Using two established convolutional segmentation backbones (U-Net and DeepLabV3+), we first perform scene-specific fine-tuning to select a robust architecture and identify the best monthly models for L1C and L2A products separately. We then conduct cross-temporal inference by applying each best monthly model to all scenes, enabling an assessment of (i) which months provide favourable training and inference conditions, (ii) how performance transfers between seasons, (iii) the impact of processing level, and (iv) how these effects differ across built-up typologies. Based on these results, we provide practical guidance for routine Sentinel-2 building classification under varying acquisition periods and settlement characteristics.
Michał Romaszewski, Kamil Drejer, Katarzyna Kołodziej +7