Organizations: Vilnius University, Institute of Computer Science, Artificial Intelligence Methods Lab · IRISA, Université Bretagne Sud · European Commission Joint Research Centre
This paper investigates AutoResearch, a protocol in which a coding language model edits a training program under a one-hour GPU budget and retains a change only if validation IoU improves. The protocol is applied to photovoltaic panel segmentation on a frozen real-image split, with DeepLabV3--ResNet-50 held fixed. Three campaigns of 24 experiments, using Gemma~4 12B, Qwen3-8B all improve their one-hour baselines, but retained modifications do not transfer across hardware. The Qwen3-8B configuration, trained on real images only, reaches a test IoU of 0.836 versus 0.833 for the reference GAN-augmented schedule. Research repository https://github.com/VU-AIML/automl4eo-autoresearch-segmentation.
Figures & tables
Loop A
Loop B
Loop C
Coding model
Gemma 4 12B
Qwen3-8B
Qwen3-8B
GPU
RTX 4070 12 GB
ASUS Ascent GX10
ASUS Ascent GX10
Batch / peak VRAM
2 / ≈ 9 GB
8 / 18–35 GB
8 → 16 / 35–36 GB
Baseline val. IoU
0.616
0.723
0.737
Best val. IoU
0.710
0.858
0.858
Δ val. IoU
+0.094
+0.135
+0.120
Table 1: Summary of the three AutoResearch campaigns. Architecture, split, seed, and a one-hour training budget are shared. Loops A and C are not evaluated on the held-out test set.
Figure 1: Validation IoU over three AutoResearch campaigns on the same frozen real split and one-hour budget. Solid lines show the running best of each campaign. Transparent markers correspond to rejected trials. Stars mark the selected best of each campaign. The Loop A clipping trial (IoU 0.349) lies below the plotted range.
Modification family
Loop A
Loop B
Loop C
Soft Dice (+BCE)
rejected
retained +0.103
retained +0.076
Cosine schedule
retained +0.092
rejected
rejected
Polynomial LR 0.9
rejected
retained +0.012
—
AdamW vs. Adam
rejected
retained +0.005
—
Auxiliary head ×0.4
rejected
retained +0.004
—
Gradient clip 1.0
rejected −0.359
retained +0.002
—
Table 2: Outcome of overlapping modification families. Deltas are computed against the configuration on which the change was stacked. Retention means that the commit became the new incumbent.
Setup
Data
Budget
IoU
F1
Paper no_aug
real
≤ 100 ep. (A100)
0.801
0.853
Paper basic_aug
real
≤ 100 ep. (A100)
0.813
0.865
Paper gan60 (paper best)
real + 60% GAN
≤ 100 ep. (A100)
0.833
0.880
Loop C best (val.)
real
3600 s (GX10)
0.858 †
—
Loop B best (val.)
real
3600 s (GX10)
0.858 †
0.908 †
Loop B best (test)
real
3600 s (GX10)
0.836
0.891
Table 3: Loop B compared with Table 5 of Lekavičius and Gružauskas (2024) on the same 256 test images. Loops A and C are validation-only.
Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Nine pretrained segmentation systems are compared using approximately equal computation budgets. The candidates include DeepLabV3+, SegFormer, UPerNet, Mask2Former, DINOv3 with a linear decoder, and an Encoder only Mask Transformer using DINOv3. The two leading candidates are trained independently with three random seeds and longer budgets. Equal averaging of class probabilities from the three Encoder only Mask Transformer models, evaluated at three image scales with horizontal reflection, produces 60.95% mean intersection over union and 71.16% mean F1 on the 84 image public validation split. The submitted predictions receive 57.08% mean intersection over union and 67.96% mean F1 on the hidden test leaderboard. Producing all 249 test masks takes 251.49 seconds including model initialization and provenance checks on one NVIDIA RTX 5090. Peak allocated GPU memory is 2.70 GiB. The study reports all eligible models, all inference variants, class level errors, source conditions, and reproducibility checks, providing a documented challenge workflow with existing architectures.
Yunus Serhat Bıçakçı
Department of Artificial Intelligence and Machine Learning, Faculty of Applied Sciences, Marmara University, Istanbul, Türkiye · Geospatial Data Science Group, School of Geographical and Earth Sciences, University of Glasgow, Glasgow, UK
We present \textbf{LlamaSeg}, a visual autoregressive framework that unifies multiple image segmentation tasks via natural language instructions. By reformulating segmentation as visual generation, LlamaSeg encodes masks as visual tokens and uses a LLaMA-style Transformer for direct next-token prediction, naturally fitting segmentation into autoregressive architectures. To support large-scale training, we introduce a data annotation pipeline and construct the \textbf{SA-OVRS} dataset, which contains \textbf{2M} segmentation masks annotated with over \textbf{5,800} open vocabulary labels or diverse textual descriptions, spanning diverse real-world scenarios. This enables our model to localize objects in images based on text prompts and to generate fine-grained masks. We further introduce the composite metric average Hausdorff Distance (dAHD) to evaluate mask contour fidelity for generative models better. Experiments show that LlamaSeg consistently outperforms existing generative approaches on multiple segmentation benchmarks and delivers finer, more accurate segmentation masks. Code and dataset are available at https://github.com/GML-FMGroup/llamaseg.
Jiru Deng, Tengjin Weng, Tianyu Yang +3
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) · Tsinghua University · Shenzhen University +2
Open-vocabulary segmentation identifies and segments objects from arbitrary textual descriptions. SAM 3 supports noun-phrase-guided segmentation and achieves competitive open-vocabulary performance through exhaustive vocabulary traversal, yet suffers from prohibitive computational overhead as target categories scale. In this paper, we propose an Efficient Open-Vocabulary segmentation framework with SAM 3 (EOVSAM), which adapts SAM 3 for single-pass prediction. EOVSAM removes prompt conditioning to turn SAM 3 into an efficient mask generator and introduces a new Attentional Aggregation strategy to optimize open-vocabulary classification end-to-end. This formulation avoids the multi-stage pipelines and post-processing heuristics commonly used by existing methods, while mitigating the closed-set collapse that can arise when classification is optimized directly. EOVSAM consistently improves segmentation accuracy over vanilla SAM 3 on all evaluated datasets and accelerates inference by up to 338×. Furthermore, EOVSAM maintains high accuracy at lower resolutions while achieving even more remarkable inference speeds. Experiments on standard semantic and panoptic segmentation benchmarks show that EOVSAM combines competitive or state-of-the-art accuracy with a substantial speed advantage over existing open-vocabulary segmentation models. Code and models are available at https://github.com/hustvl/EOVSAM.
Haomin Peng, Yongkang Li, Zhaoxiang Liu +4
Huazhong University of Science and Technology · Data Science & Artificial Intelligence Research Institute, China Unicom · Unicom Data Intelligence, China Unicom +1