Vision Transformer Ensembles for Panoramic Street Segmentation
Authors: Yunus Serhat Bıçakçı
Organizations: Department of Artificial Intelligence and Machine Learning, Faculty of Applied Sciences, Marmara University, Istanbul, Türkiye · Geospatial Data Science Group, School of Geographical and Earth Sciences, University of Glasgow, Glasgow, UK
Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Nine pretrained segmentation systems are compared using approximately equal computation budgets. The candidates include DeepLabV3+, SegFormer, UPerNet, Mask2Former, DINOv3 with a linear decoder, and an Encoder only Mask Transformer using DINOv3. The two leading candidates are trained independently with three random seeds and longer budgets. Equal averaging of class probabilities from the three Encoder only Mask Transformer models, evaluated at three image scales with horizontal reflection, produces 60.95% mean intersection over union and 71.16% mean F1 on the 84 image public validation split. The submitted predictions receive 57.08% mean intersection over union and 67.96% mean F1 on the hidden test leaderboard. Producing all 249 test masks takes 251.49 seconds including model initialization and provenance checks on one NVIDIA RTX 5090. Peak allocated GPU memory is 2.70 GiB. The study reports all eligible models, all inference variants, class level errors, source conditions, and reproducibility checks, providing a documented challenge workflow with existing architectures.
Figures & tables
Split
RGB panoramas
Available masks
Training
497
497
Validation
84
84
Hidden test
249
0
Table 1: Official split used in the experiments. Hidden test annotations were not available to the training or selection procedure.
Figure 1: The selection procedure uses public validation labels throughout. Hidden test labels are never used for training, checkpoint selection, or inference variant selection.
System
Base learning rate
Encoder factor
Decay
Clip
DeepLabV3+ ResNet50
10−4
1
0.01
1
SegFormer B2 and B5
6×10−4
0.1
0.01
1
DINOv3 linear B and L
6×10−4
0.1
0.01
1
EoMT DINOv3 L
6×10−5
1
0.01
1
UPerNet ConvNeXt L
10−4
0.3
0.05
1
UPerNet Swin L
6×10−5
1
0.01
1
Table 2: Eligible optimizer settings. The encoder learning rate equals the base rate multiplied by the encoder factor. Gradient clipping uses the recorded norm threshold.
System
mIoU
mF1
Rare IoU
Seconds
GiB
EoMT DINOv3 L
57.565
68.515
26.843
624.20
10.58
DINOv3 L linear
52.151
63.566
24.841
628.02
5.71
Mask2Former Swin L
51.926
62.885
15.591
609.14
9.36
DINOv3 B linear
49.110
60.300
21.979
603.15
1.99
UPerNet ConvNeXt L
48.784
59.043
13.314
609.52
9.59
UPerNet Swin L
46.876
56.323
5.802
605.99
13.05
Table 3: Eligible pilot results at approximately 600 seconds of training and validation computation. Scores are percentages. Rare IoU is the mean over the eight categories fixed from training pixel frequency. Memory is peak allocated training memory in GiB.
System
Seed
mIoU
mF1
Rare IoU
Best epoch
EoMT DINOv3 L
42
57.970
68.576
31.336
26
EoMT DINOv3 L
123
58.731
69.353
30.846
36
EoMT DINOv3 L
2026
59.670
69.718
31.103
31
DINOv3 L linear
42
54.660
66.021
28.406
45
DINOv3 L linear
123
54.706
65.862
29.475
52
DINOv3 L linear
2026
54.098
65.523
27.448
58
Table 4: Independent confirmation runs with approximately 2400 seconds of training and validation computation each. The checkpoint is selected from ten scheduled validation evaluations per run.
Inference variant
mIoU
mF1
Seconds
GiB
Three EoMT seeds, scales and reflection
60.953
71.165
81.72
2.70
E, scales and reflection
60.671
70.828
16.77
2.70
Three EoMT seeds, native
59.852
70.093
51.82
2.32
E, native
59.670
69.718
7.85
2.32
0.75E+0.25D , native
59.652
69.624
34.72
2.32
Three EoMT seeds, reflection
59.510
69.595
56.95
2.39
Table 5: All eleven inference variants on 84 validation panoramas. E denotes the best EoMT checkpoint and D the best linear DINOv3 checkpoint. Ensemble weights apply to normalized semantic probabilities. Seconds measure the complete prediction loop, including transfers and PNG output.
Figure 2: Measured accuracy and prediction cost of four EoMT inference configurations. The selected ensemble yields the highest validation score, while the transformed single model is much faster. The graph begins at 59.2% on the vertical axis to show the small differences clearly.
Figure 3: Real validation examples nearest the lower quartile, median, and upper quartile of pixel agreement across all 84 images. This selection uses pixel agreement rather than class balanced IoU. Predictions are the saved outputs of the selected three seed ensemble with scales and reflection. Masks use the official PalmCity colors. Pink marks incorrect class identifiers, including Void.
Figure 4: The validation image with the lowest pixel agreement, GS__2750 . The blue box marks the detail shown below, selected as the 384×192 window with the most incorrect pixels outside reference Void regions. The shaded walkway is largely predicted as Road instead of Sidewalk, while the stairs are largely assigned to Operator and Shadow. Pink marks all incorrect class identifiers.
Rank
Participant
Submission
mIoU
mF1
1
yunusserhat
961441
57.08
67.96
2
onurcbayrak
945126
46.02
56.99
3
MaxwellEQ
950834
33.54
40.93
Table 6: PalmCity Test Task leaderboard snapshot dated 5 October 2026. Values are percentages. The table identifies the submitted system without inferring the training procedure of other participants.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
ID
Class
Single
Selected
Difference
0
Road
94.81
95.07
+0.26
1
Sidewalk
77.36
78.45
+1.10
2
Parking Lot
28.41
39.73
+11.31
3
Soil
76.21
78.94
+2.73
4
Pedestrian
65.88
69.22
+3.34
5
Driver
73.91
71.08
−2.84
Appendix
Table 7: Validation IoU for every official category. Single refers to EoMT seed 2026 at the native scale. Selected is the three seed ensemble with scales and reflection. Differences are percentage points calculated before rounding.
Panoptic segmentation requires the simultaneous recognition of countable thing instances and amorphous stuff regions, placing joint demands on long-range context modelling, multi-scale feature representation, and efficient dense prediction. Existing convolutional and transformer-based methods struggle to satisfy all three requirements concurrently: convolutional architectures are limited in their capacity to model long-range dependencies, while transformer-based methods incur quadratic computational cost that is prohibitive at high resolutions. In this paper, we propose MambaPanoptic, a fully Mamba-based panoptic segmentation framework that addresses these limitations through two principal contributions. First, we introduce MambaFPN, a top-down feature pyramid that leverages Mamba blocks to generate globally coherent, multi-scale feature representations with linear computational complexity. Second, we adopt a PanopticFCN-style kernel generator that produces unified thing and stuff kernels for proposal-free panoptic prediction, enhanced by a QuadMamba-based feature refinement module applied at multiple network stages. Experiments on the Cityscapes and COCO panoptic segmentation benchmarks demonstrate that MambaPanoptic consistently outperforms PanopticDeepLab and PanopticFCN under comparable model sizes, and matches or surpasses Mask2Former on Cityscapes in PQ and AP while requiring fewer parameters.
Qing Cheng, Damiano Bertolini, Wei Zhang +3
Technical University of Munich · Munich Center for Machine Learning (MCML) · Polytechnic University of Milan +3
We describe our winning entry to the Waymo Open Dataset 2D Video Panoptic Segmentation Challenge. The task asks for a semantic class at every pixel of every frame and, for countable objects, an identity that holds across 100 frames and across five overlapping cameras. We build on DVIS++, a cascade of a segmenter, a tracker, and a refiner, as our baseline. We propose hard region supervision (HRS) to improve the baseline. In particular, we use the baseline to define the hard region as where it makes mistakes, and design a loss and an auxiliary prediction head for this region. The auxiliary head is used only in training and removed at test time, so at inference the model trained with HRS has the same architecture as the baseline. In addition, we propose three test-time steps that further improve the results: a two-model ensemble, a merge of the segmenter's output into the final panoptic map, and cross-camera identity linking. On the challenge test set, our entry reaches 0.3547 wSTQ, 0.2071 wAQ, and 0.6075 mIoU, ranking first on all three metrics. It is 3.6 wSTQ points ahead of the second entry and 2.4 points ahead of our DVIS++ baseline.
This report presents our solution for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge, which requires parsing unstructured outdoor scenes from four camera platforms into 56 fine-grained categories. Our approach pairs foundation vision encoders (including DINOv3, SigLIP2, and InternImage) with a Mask2Former decoder, and trains them with a strong recipe including long training schedules, exponential moving average, a larger crop size, and multi-scale plus flip test-time augmentation. The three encoders, chosen for their complementary pretraining objectives, are combined into a pretraining-diverse ensemble through per-class validation-IoU weighting. Evaluated on the official GOOSE test set, our submission achieves 75.40% composite mIoU and wins the second place of the challenge. Our study further shows that the encoder's pretraining recipe, rather than its parameter count or the decoder design, is the dominant factor for accuracy on this benchmark.
Boyan Wang, Yongxi Huang, Wenjing Li +4
LION Lab, Hefei University of Technology, Hefei, China. · University of Macau, Macau, China.