cs.CVOct 5, 2026

Vision Transformer Ensembles for Panoramic Street Segmentation

Authors: Yunus Serhat Bıçakçı

Organizations: Department of Artificial Intelligence and Machine Learning, Faculty of Applied Sciences, Marmara University, Istanbul, Türkiye · Geospatial Data Science Group, School of Geographical and Earth Sciences, University of Glasgow, Glasgow, UK

Abstract

Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Nine pretrained segmentation systems are compared using approximately equal computation budgets. The candidates include DeepLabV3+, SegFormer, UPerNet, Mask2Former, DINOv3 with a linear decoder, and an Encoder only Mask Transformer using DINOv3. The two leading candidates are trained independently with three random seeds and longer budgets. Equal averaging of class probabilities from the three Encoder only Mask Transformer models, evaluated at three image scales with horizontal reflection, produces 60.95% mean intersection over union and 71.16% mean F1 on the 84 image public validation split. The submitted predictions receive 57.08% mean intersection over union and 67.96% mean F1 on the hidden test leaderboard. Producing all 249 test masks takes 251.49 seconds including model initialization and provenance checks on one NVIDIA RTX 5090. Peak allocated GPU memory is 2.70 GiB. The study reports all eligible models, all inference variants, class level errors, source conditions, and reproducibility checks, providing a documented challenge workflow with existing architectures.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MambaPanoptic: A Vision Mamba-based Structured State Space Framework for Panoptic Segmentation

    May 12, 2026Qing Cheng, Damiano Bertolini, Wei Zhang +3Panoptic SegmentationFeature Pyramid