Research
Vision Transformer Ensembles for Panoramic Street Segmentation
Overview Research area: Computer vision, specifically semantic segmentation of panoramic street-level imagery and model-selection methodology for a benchmark challenge. Technical level: Advanced. The

- arXiv
- 2610.06063
- Published
- 2026-10-05
- Authors
- Yunus Serhat Bıçakçı
AI summary
Overview
Research area: Computer vision, specifically semantic segmentation of panoramic street-level imagery and model-selection methodology for a benchmark challenge.
Technical level: Advanced. The paper assumes familiarity with transformer-based segmentation architectures, probability ensembling, and multi-scale test-time augmentation, though every recipe and budget is reported explicitly.
One-sentence scope: This paper documents the candidate-selection, training, ensembling, and inference procedure behind a first-place submission to the PalmCity panoramic street segmentation challenge in the leaderboard snapshot dated 5 October 2026.
What This Paper Is About
Segmenting street panoramas into 32 semantic categories can support detailed descriptions of urban environments, but the PalmCity labelled training set is small (497 panoramas) and different pretrained models consume very different amounts of computation. This makes fair model selection difficult: training everything for the same number of epochs gives slower models a larger computation allowance, while a very short shared schedule may hide the value of a large model. The paper's goal is to describe a documented selection procedure that compares nine pretrained systems under approximately equal computation budgets, confirms the two leaders with three random seeds each, and then chooses among eleven fixed inference variants using only the 84-image public validation split.
Key Contributions
-
A bounded, cost-aware candidate comparison: nine eligible pretrained segmentation systems evaluated at an estimated 600 seconds of training plus validation computation each, with no architecture expanded or dropped after seeing individual scores. The pilot produced 603.15 to 628.02 seconds of measured computation per eligible system.
-
An explicit recipe-correction record: four optimizer configurations (SegFormer encoder/head rate ratio, Mask2Former encoder rate, decay and clipping, ConvNeXt encoder rate and decay) were corrected and rerun under the same budget, giving 13 pilot executions plus six confirmation executions, or 19 training runs in total. The study retains both the corrected and uncorrected scores for those systems rather than selecting the larger result.
-
A uniform probability interface for dense and query-based segmentation: the paper defines how query models (EoMT, Mask2Former) are converted into the same 32-class semantic distribution as dense softmax predictors, discarding only the separate "no object" query category while retaining class 31 (Void).
-
A complete, reproducible inference sweep: eleven variants fixed before evaluation, covering single checkpoints, three ensemble weights, the three-seed ensemble, horizontal reflection, three scales, and overlapping windows, with per-variant runtime and memory reported alongside accuracy, and a public code, configuration, and model release.
Main Findings
-
EoMT led the pilot: EoMT with DINOv3 ViT L reached 57.565% mean IoU and 68.515% mean F1 in the pilot, ahead of DINOv3 ViT L with a linear head (52.151%) and Mask2Former Swin L (51.926%). The gap between second and third place was only 0.226 percentage points.
-
Model size did not determine ranking: SegFormer B5 (45.742%) trailed SegFormer B2 (46.221%), and DINOv3 ViT B with a linear decoder (49.110%) exceeded both UPerNet variants (ConvNeXt L at 48.784%, Swin L at 46.876%).
-
Recipe corrections changed pilot scores substantially: SegFormer B5 moved from 34.637% to 45.742% mean IoU, SegFormer B2 from 44.323% to 46.221%, and UPerNet ConvNeXt L from 47.747% to 48.784%. Mask2Former moved slightly downward, from 52.015% to 51.926%, and that lower corrected score remained the eligible result.
-
EoMT stayed ahead over three seeds: Across seeds 42, 123, and 2026, EoMT averaged 58.791% mean IoU versus 54.488% for the linear DINOv3 ViT L system, a mean difference of 4.303 percentage points. Every EoMT run exceeded every linear DINOv3 run. The population standard deviation across seeds was 0.695 percentage points for EoMT and 0.276 for the linear system.
-
The selected ensemble combined seeds and transforms: Equal averaging of the three EoMT seeds at scales 0.75, 1, and 1.25 with horizontal reflection produced 60.953% mean IoU and 71.165% mean F1 on the 84-image validation split. This is 1.283 percentage points above the best native single checkpoint (59.670%) and 0.282 points above the transformed single checkpoint (60.671%).
-
Ensembling three seeds alone helped only modestly: Native-scale averaging of the three EoMT seeds gave 59.852% mean IoU, an improvement of 0.181 percentage points over the strongest single checkpoint.
-
Added diversity was not automatically beneficial: Combining EoMT with the linear DINOv3 checkpoint reduced mean IoU at all three tested weights (0.75E+0.25D at 59.652%, 0.5E+0.5D at 58.364%, 0.25E+0.75D at 55.863%). Reflection alone (59.025%) and overlapping 256x512 windows with 0.5 overlap (59.070%) also scored below the native single checkpoint.
-
Accuracy cost time, not memory: Prediction time for 84 validation images rose from 16.77 seconds for the transformed single checkpoint to 81.72 seconds for the selected ensemble, a factor of 4.87. Peak allocated inference memory stayed near 2.70 GiB because models were moved to the GPU sequentially and probability sums were kept on the CPU.
-
Large scene regions were predicted well, rare classes remained weak: Sky reached 98.88% IoU, Road 95.07%, Building 94.38%, Void 92.65%, Car 91.69%, and Water Surface 91.24%. The average over the eight least frequent classes was 32.044% for the selected ensemble versus 31.103% for the best native single checkpoint. Parking Lot improved from 28.41% to 39.73% and Railing from 0.38% to 11.23%, while Motorcycle fell by 7.83 percentage points, Bus by 4.91, and Driver by 2.84. Bicycle remained at 15.86%, Stairs at 13.92%, and Overpass and Pruned Tree reached zero IoU.
-
Hidden test results were lower than validation: The submitted archive scored 57.08% mean IoU and 67.96% mean F1 on the hidden leaderboard, 3.88 and 3.20 percentage points below the public validation figures respectively. The leaderboard snapshot lists this submission first, ahead of 46.02% and 33.54% mean IoU entries.
-
Runtime and archive were fully measured: Producing all 249 test masks took 242.19 seconds, with total runtime of 251.49 seconds after checkpoint loading and provenance checks, on one NVIDIA RTX 5090. Peak allocated memory was 2.7015 GiB and peak reserved memory 3.6133 GiB. The submitted archive is 1.82 MiB of single-channel 1024x512 PNGs.
-
Total experiment cost was recorded: Complete training and validation computation was 6.3443 GPU hours, including four additional diagnostic pilot runs and excluding initialization and checkpoint input/output. The inference sweep took 392.63 seconds of elapsed time within a fixed 1800 second allowance.
Methodology in Plain English
The researchers started from the official PalmCity split (497 training panoramas, 84 validation panoramas, 249 hidden test panoramas, all 1024x512 with 32 classes) and changed nothing about it. They audited file counts, image modes, dimensions, class ranges, and duplicates across splits, and confirmed that all 32 categories appear in both training and validation.
Rather than giving every model the same number of epochs, they first calibrated each candidate on real data: one warmup optimizer update, three measured updates, one warmup inference image, and three measured validation images at the full 512x1024 resolution. That calibration converted an estimated 600-second budget into a fixed number of optimizer updates, a maximum epoch count, and a warmup length, with ten validation evaluations reserved inside the budget. All models used BF16 mixed precision and an effective batch size of four; DeepLabV3+ and UPerNet needed microbatches of two because their pooled batch normalization requires more than one sample, while the rest used microbatches of one with four accumulation steps.
The nine eligible systems were classified by where their weights came from: the ResNet50 and MiT B2 reference encoders used ImageNet weights with new segmentation heads; SegFormer B5, UPerNet, Mask2Former, and EoMT started from complete ADE20K semantic segmentation checkpoints; the DINOv3 linear systems used pretrained DINOv3 encoders with a new head. Every prediction layer was adapted to 32 PalmCity classes and every encoder was updated during PalmCity training.
The two leading pilots were retrained from their original pretrained sources, not continued from pilot weights, with three seeds each and an estimated 2400 seconds per run. To make the comparison fair, the pipeline then converted every model output into a probability distribution over the same 32 classes. Dense models use a softmax over logits. Query models combine the per-query class probability with the sigmoid mask value at each pixel, sum over queries, and normalize, discarding only the separate no-object category. Only probabilities are averaged, never integer class masks.
The final chosen configuration separately averages the three EoMT seeds over six transforms: scales 0.75, 1, and 1.25 (corresponding to heights 384, 512, and 640 pixels, preserving the 2:1 panorama aspect ratio), each in original and horizontally reflected form. Reflection is reversed and scaled outputs are aligned back to the native 512x1024 grid with bilinear interpolation before averaging. Class scores use a single pooled confusion matrix with mean IoU and mean F1 computed as arithmetic means over all 32 classes, including Void and including classes absent from truth or prediction (which score zero). The local scorer was checked against the published PalmCity scoring program on 103 synthetic confusion matrices with exact agreement.
Why This Matters
Impact on research. The paper's central claim is procedural rather than architectural: correct class semantics, suitable source initialization, budget calibration, and output verification matter alongside the architecture name. It shows that a weak short run can reflect an unsuitable adaptation recipe rather than a weak model, and it publishes 19 training runs including unsuccessful alternatives, giving other researchers a documented challenge workflow with existing architectures rather than a single headline score.
Real-world applications.
- Street-level urban description: identifying road, building, sky, vegetation, water, and sidewalk regions in pedestrian-height panoramas to complement remote sensing and conventional spatial data.
- Active mobility and accessibility analysis: measuring footpaths, barriers, and street furniture, though the paper cautions that classes relevant to bicycle infrastructure and small traffic controls remain weak.
- Urban typology and morphology studies: the paper cites prior work using 116696 Mapillary images from Istanbul's Fatih district to identify interpretable urban typologies without manual labels.
- Exposure measurement in environmental health research: street imagery has been used to study green and blue space associations with depression symptoms, and the Amsterdam follow-up found no statistically significant association, illustrating that measurement perspective and setting both matter.
Industry relevance. The work quantifies the real cost of accuracy: the selected ensemble is 4.87 times slower than the faster transformed single model for a 0.282 percentage point validation gain. Teams deploying segmentation at scale can use this tradeoff directly. The runtime figures (251.49 seconds for 249 masks, 2.70 GiB peak allocated memory on one NVIDIA RTX 5090) and the sequential model-loading strategy that keeps three checkpoints within one GPU's memory are practical deployment details. The released checkpoints, dependency lock, and configurations support direct reuse without retraining.
Future Directions
-
Confirming near-miss candidates: Mask2Former was only 0.226 percentage points behind the second pilot candidate and could merit longer training in a larger study, especially given that its corrected score was slightly lower than its uncorrected one.
-
Improving DINOv3 adaptation: The paper notes that larger or more suitable decoders for DINOv3 may narrow the gap with EoMT, and that isolating architectural contributions would require matched pretraining and additional controlled comparisons.
-
Rare-class and small-object performance: Overpass, Pruned Tree, Bicycle, and Stairs remain weak, and rare-object sampling, broader augmentation, and longer schedules are proposed as untested directions.
-
Geographic and temporal validation: Validation and test images may be visually similar across the official split, and route, GPS, and acquisition time metadata would be needed for a spatial evaluation. Extending encoder-focused mask prediction to connected panoramas or street video, as in recent VidEoMT work, is raised as a complementary direction.
Target Audience
This paper is most useful to computer vision practitioners and machine learning engineers working on semantic segmentation under constrained compute, particularly those entering benchmark challenges where model selection, budget calibration, and inference-variant choice must be justified rather than guessed. It also serves remote sensing and geospatial data scientists who need to know how reliable street-level segmentation outputs are before using them as inputs to urban, mobility, or environmental health analysis. Reproducibility-focused researchers will benefit from the retained unsuccessful alternatives, the recorded environment versions (Python 3.12.12, PyTorch 2.9.1 with CUDA 12.8, torchvision 0.24.1, Transformers 5.5.2, segmentation models pytorch 0.5.0, timm 1.0.30, SciPy 1.17.1), and the explicit note that strict CUDA determinism was disabled, so hardware and numerical differences can affect a rerun.
Authors’ abstract
Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Nine pretrained segmentation systems are compared using approximately equal computation budgets. The candidates include DeepLabV3+, SegFormer, UPerNet, Mask2Former, DINOv3 with a linear decoder, and an Encoder only Mask Transformer using DINOv3. The two leading candidates are trained independently with three random seeds and longer budgets. Equal averaging of class probabilities from the three Encoder only Mask Transformer models, evaluated at three image scales with horizontal reflection, produces 60.95% mean intersection over union and 71.16% mean F1 on the 84 image public validation split. The submitted predictions receive 57.08% mean intersection over union and 67.96% mean F1 on the hidden test leaderboard. Producing all 249 test masks takes 251.49 seconds including model initialization and provenance checks on one NVIDIA RTX 5090. Peak allocated GPU memory is 2.70 GiB. The study reports all eligible models, all inference variants, class level errors, source conditions, and reproducibility checks, providing a documented challenge workflow with existing architectures.