Vision Transformer Ensembles for Panoramic Street Segmentation
Authors: Yunus Serhat Bıçakçı
Organizations: Department of Artificial Intelligence and Machine Learning, Faculty of Applied Sciences, Marmara University, Istanbul, Türkiye · Geospatial Data Science Group, School of Geographical and Earth Sciences, University of Glasgow, Glasgow, UK
Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Nine pretrained segmentation systems are compared using approximately equal computation budgets. The candidates include DeepLabV3+, SegFormer, UPerNet, Mask2Former, DINOv3 with a linear decoder, and an Encoder only Mask Transformer using DINOv3. The two leading candidates are trained independently with three random seeds and longer budgets. Equal averaging of class probabilities from the three Encoder only Mask Transformer models, evaluated at three image scales with horizontal reflection, produces 60.95% mean intersection over union and 71.16% mean F1 on the 84 image public validation split. The submitted predictions receive 57.08% mean intersection over union and 67.96% mean F1 on the hidden test leaderboard. Producing all 249 test masks takes 251.49 seconds including model initialization and provenance checks on one NVIDIA RTX 5090. Peak allocated GPU memory is 2.70 GiB. The study reports all eligible models, all inference variants, class level errors, source conditions, and reproducibility checks, providing a documented challenge workflow with existing architectures.
Figures & tables
Split
RGB panoramas
Available masks
Training
497
497
Validation
84
84
Hidden test
249
0
Table 1: Official split used in the experiments. Hidden test annotations were not available to the training or selection procedure.
Figure 1: The selection procedure uses public validation labels throughout. Hidden test labels are never used for training, checkpoint selection, or inference variant selection.
System
Base learning rate
Encoder factor
Decay
Clip
DeepLabV3+ ResNet50
10−4
1
0.01
1
SegFormer B2 and B5
6×10−4
0.1
0.01
1
DINOv3 linear B and L
6×10−4
0.1
0.01
1
EoMT DINOv3 L
6×10−5
1
0.01
1
UPerNet ConvNeXt L
10−4
0.3
0.05
1
UPerNet Swin L
6×10−5
1
0.01
1
Table 2: Eligible optimizer settings. The encoder learning rate equals the base rate multiplied by the encoder factor. Gradient clipping uses the recorded norm threshold.
System
mIoU
mF1
Rare IoU
Seconds
GiB
EoMT DINOv3 L
57.565
68.515
26.843
624.20
10.58
DINOv3 L linear
52.151
63.566
24.841
628.02
5.71
Mask2Former Swin L
51.926
62.885
15.591
609.14
9.36
DINOv3 B linear
49.110
60.300
21.979
603.15
1.99
UPerNet ConvNeXt L
48.784
59.043
13.314
609.52
9.59
UPerNet Swin L
46.876
56.323
5.802
605.99
13.05
Table 3: Eligible pilot results at approximately 600 seconds of training and validation computation. Scores are percentages. Rare IoU is the mean over the eight categories fixed from training pixel frequency. Memory is peak allocated training memory in GiB.
System
Seed
mIoU
mF1
Rare IoU
Best epoch
EoMT DINOv3 L
42
57.970
68.576
31.336
26
EoMT DINOv3 L
123
58.731
69.353
30.846
36
EoMT DINOv3 L
2026
59.670
69.718
31.103
31
DINOv3 L linear
42
54.660
66.021
28.406
45
DINOv3 L linear
123
54.706
65.862
29.475
52
DINOv3 L linear
2026
54.098
65.523
27.448
58
Table 4: Independent confirmation runs with approximately 2400 seconds of training and validation computation each. The checkpoint is selected from ten scheduled validation evaluations per run.
Inference variant
mIoU
mF1
Seconds
GiB
Three EoMT seeds, scales and reflection
60.953
71.165
81.72
2.70
E, scales and reflection
60.671
70.828
16.77
2.70
Three EoMT seeds, native
59.852
70.093
51.82
2.32
E, native
59.670
69.718
7.85
2.32
0.75E+0.25D , native
59.652
69.624
34.72
2.32
Three EoMT seeds, reflection
59.510
69.595
56.95
2.39
Table 5: All eleven inference variants on 84 validation panoramas. E denotes the best EoMT checkpoint and D the best linear DINOv3 checkpoint. Ensemble weights apply to normalized semantic probabilities. Seconds measure the complete prediction loop, including transfers and PNG output.
Figure 2: Measured accuracy and prediction cost of four EoMT inference configurations. The selected ensemble yields the highest validation score, while the transformed single model is much faster. The graph begins at 59.2% on the vertical axis to show the small differences clearly.
Figure 3: Real validation examples nearest the lower quartile, median, and upper quartile of pixel agreement across all 84 images. This selection uses pixel agreement rather than class balanced IoU. Predictions are the saved outputs of the selected three seed ensemble with scales and reflection. Masks use the official PalmCity colors. Pink marks incorrect class identifiers, including Void.
Figure 4: The validation image with the lowest pixel agreement, GS__2750 . The blue box marks the detail shown below, selected as the 384×192 window with the most incorrect pixels outside reference Void regions. The shaded walkway is largely predicted as Road instead of Sidewalk, while the stairs are largely assigned to Operator and Shadow. Pink marks all incorrect class identifiers.
Rank
Participant
Submission
mIoU
mF1
1
yunusserhat
961441
57.08
67.96
2
onurcbayrak
945126
46.02
56.99
3
MaxwellEQ
950834
33.54
40.93
Table 6: PalmCity Test Task leaderboard snapshot dated 5 October 2026. Values are percentages. The table identifies the submitted system without inferring the training procedure of other participants.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
ID
Class
Single
Selected
Difference
0
Road
94.81
95.07
+0.26
1
Sidewalk
77.36
78.45
+1.10
2
Parking Lot
28.41
39.73
+11.31
3
Soil
76.21
78.94
+2.73
4
Pedestrian
65.88
69.22
+3.34
5
Driver
73.91
71.08
−2.84
Appendix
Table 7: Validation IoU for every official category. Single refers to EoMT seed 2026 at the native scale. Selected is the three seed ensemble with scales and reflection. Differences are percentage points calculated before rounding.