cs.CVOct 6, 2026

Towards benchmarking Western Bluebird detection in the wild

Authors: Estela Monserrat Arriaga Santana, Julian Rosas Scull, Ibeth P. Alarcón, Bibiana Montoya, Aylin Sosa Mejía, Hugo Jair Escalante

Organizations: Universidad Nacional Aut´onoma de M´exico, Mexico · Universidad Aut´onoma de Tlaxcala, Mexico · The University of Texas at El Paso, USA and INAOE, Mexico

Abstract

Bird monitoring in natural environments is challenging due to the small size of some species of birds relative to the scene, background clutter, variability in illumination, and the observers' viewpoint. Progress is further limited by the scarcity of large-scale, realistic datasets, which are essential for understanding behavioral patterns. To address this gap, we introduce a new benchmark dataset for the detection and segmentation of Western bluebirds (Sialia Mexicana), comprising over 6,000 labeled images from 41 recording sessions. The dataset features high-resolution (4K) in-the-wild images in which birds occupy only a small fraction of the image. We evaluated supervised detectors, open-vocabulary models under zero-shot and fine-tuned settings, and segmentation approaches. Supervised detectors remain the most reliable overall, with Faster R-CNN achieving the highest detection mAP and RT-DETR offering the best precision-recall trade-off. Open-vocabulary models perform poorly in zero-shot settings; however, fine-tuning substantially improves their performance, with YOLO-World becoming competitive with supervised methods and achieving the highest precision, F1-score, and mAP@0.5. For segmentation, supervised methods significantly outperform Grounded-SAM and SAM 3: Mask R-CNN achieves the highest mask mAP, while YOLOv8-Seg provides the best precision and fastest inference. A diagnostic analysis further shows that failures are not explained by object size alone, but by a combination of apparent scale, brightness, contrast, clutter, blur, crowding, and recording-session variation. Overall, our findings highlight the difficulty of zero-shot bird detection in cluttered ecological scenes and underscore the importance of domain adaptation in small-object settings.

Figures & tables

Explore similar work

CardsList
  1. RareSpot+: A Benchmark, Model, and Active Learning Framework for Small and Rare Wildlife in Aerial Imagery

    Apr 21, 2026Bowen Zhang, Jesse T. Boulerice, Charvi Mendiratta +4Aerial ImagerySpecies

  2. Time-frequency localization of bird calls in dense soundscapes

    Jun 9, 2026Simen Hexeberg, Fanghui Tong, Hari Vishnu +1Animal VocalizationsPassive Acoustic Monitoring

  3. Overhead Wildlife Locator (OWL): Benchmarking Weakly Supervised Learning for Aerial Wildlife Surveys

    Jun 11, 2026Isai Daniel Chacón, Zhongqi Miao, Bruno Demuro +9SpeciesUnmanned Aerial Vehicles