cs.CLSep 28, 2026

OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations

Authors: Faadil Mustun, Chiara Semenzin, Roberto Dessi, Pablo Robin Guerrero, Pierre Orhan, Alexis Emanuelli, Emanuele Rossi, Yair Lakretz, +2 more

Organizations: Institut de Biologie de l’École normale supérieure, CNRS, INSERM, Université PSL, Paris, France · Earth Species Project, France · Not Diamond, San Francisco, USA · Institut du Cerveau, Paris, France · Sapienza University of Rome, Rome, Italy · École Normale Supérieure, Paris, France · Champalimaud Foundation, Lisbon, Portugal

Abstract

Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute for cetaceans: despite bottlenose dolphins (Tursiops truncatus) being a compelling case of complex vocal communication among non-human mammals, existing dolphin datasets are small, fragmented, and largely closed. We introduce OpenWhistle, the largest publicly available dataset of dolphin vocalizations. It comprises approximately 180,000 whistles (114 hours) recorded over five years from a stable pod of five individuals in a semi-natural environment, paired with a curated subset of 8,354 expert-annotated whistles and reproducible evaluation protocols for whistle-type detection and classification. We further release the full processing pipeline for whistle detection, segmentation, and categorization. To demonstrate its utility, we pretrain a Wav2Vec2.0 model adapted to dolphin acoustics on the OpenWhistle corpus and show that it learns effective representations, outperforming general-purpose bioacoustic models such as AVES and BioLingual on both tasks while leaving meaningful headroom for future work. By releasing the dataset, pipeline, and evaluation protocol, we provide the first open dolphin whistle dataset tailored for training self-supervised models, laying the groundwork for advancing dolphin communication research and developing models that capture fine-grained acoustic structure within species.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

    Jun 10, 2026Chiara Semenzin, Faadil Mustun, Roberto Dessi +5Passive Acoustic MonitoringSelf-Supervised Learning

  2. MyGardenBird: A Machine-Learning-Ready Bird Sound Dataset for Twelve Common Malaysian Birds

    Jun 5, 2026Muhammad Mun'im Ahmad Zabidi, Mohd Yamani Idna Idris, Norisma IdrisAnimal VocalizationsBird

  3. A strongly annotated passive acoustic dataset for tropical bird monitoring

    May 20, 2026Daniela Ruiz, Juan Sebastián Ulloa, Zhongqi Miao +11Animal VocalizationsBiodiversity Monitoring