OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations
Authors: Faadil Mustun, Chiara Semenzin, Roberto Dessi, Pablo Robin Guerrero, Pierre Orhan, Alexis Emanuelli, Emanuele Rossi, Yair Lakretz, +2 more
Organizations: Institut de Biologie de l’École normale supérieure, CNRS, INSERM, Université PSL, Paris, France · Earth Species Project, France · Not Diamond, San Francisco, USA · Institut du Cerveau, Paris, France · Sapienza University of Rome, Rome, Italy · École Normale Supérieure, Paris, France · Champalimaud Foundation, Lisbon, Portugal
Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute for cetaceans: despite bottlenose dolphins (Tursiops truncatus) being a compelling case of complex vocal communication among non-human mammals, existing dolphin datasets are small, fragmented, and largely closed. We introduce OpenWhistle, the largest publicly available dataset of dolphin vocalizations. It comprises approximately 180,000 whistles (114 hours) recorded over five years from a stable pod of five individuals in a semi-natural environment, paired with a curated subset of 8,354 expert-annotated whistles and reproducible evaluation protocols for whistle-type detection and classification. We further release the full processing pipeline for whistle detection, segmentation, and categorization. To demonstrate its utility, we pretrain a Wav2Vec2.0 model adapted to dolphin acoustics on the OpenWhistle corpus and show that it learns effective representations, outperforming general-purpose bioacoustic models such as AVES and BioLingual on both tasks while leaving meaningful headroom for future work. By releasing the dataset, pipeline, and evaluation protocol, we provide the first open dolphin whistle dataset tailored for training self-supervised models, laying the groundwork for advancing dolphin communication research and developing models that capture fine-grained acoustic structure within species.
Figures & tables
Dataset
# Whistles
Voc. hours
Time span (yrs)
Stable pod (# indiv.)
Setting
Seq. context
Open
OpenWhistle Pretraining
∼ 180,000*
114.3
5.0
✓ (5)
Semi-natural
✓
✓
OpenWhistle Expert subset
8,354
1.9
0.42
✓ (5)
Semi-natural
✗
✓
DOLPHINFREE [ 19 ]
4,600
7.3
2.0
✗
Wild
✗
✓
Di Nardo et al., 2025 [ 28 ]
3,111
0.6
0.003
✓ (7)
Captive
✗
✓
Watkins MMSD [ 38 ]
566
N/R
70+
✗
Wild
✗
✓
Korkmaz et al., 2023 [ 30 ]
∼ 29,000*
6.8
0.07
✗
Semi-natural
✓
∘
Table 1: Comparison of existing dolphin acoustic datasets. N/R = not reported; ∘ = available upon request. Time span is reported in years for consistency across datasets. "Seq. context" indicates whether the dataset preserves temporally contiguous sequences of multiple whistles, rather than only isolated whistle clips. All datasets are based on passive acoustic monitoring (PAM), except SDWD, which includes data collected through catch-and-release protocols. * Estimated from total vocalization duration and mean whistle duration.
Figure 1: Site Description and Whistle Repertoire. A) The unique recording site at Dolphin Reef, Eilat. Hydrophones (yellow microphones) are deployed at fixed locations to continuously capture underwater audio. Dolphins move freely within the area and can exit to the open sea. B) Representative spectrograms of whistle types. Left: Signature Whistles (SW) of resident dolphins, each showing individually distinctive frequency contours. Right: Whistles of past individuals and non-signature whistles (NSW), illustrating the diversity of vocalizations captured in the dataset.
Figure 2: OpenWhistle: Longitudinal Extent and Temporal Distribution of the Dataset. A) Cumulative recording hours over time, showing dataset growth and changes in pod composition. B) Distribution of recording hours across the day, indicating alignment with periods of human activity. C) Cumulative detected whistling hours over time, obtained by applying the whistle presence detection CNN to the raw recordings. D) Confusion matrix of the whistle presence detection CNN on the test set, indicating high reliability of the detected whistle segments used to derive panel C.
Figure 3: Analyses of Whistle Properties in OpenWhistle. A–B) Temporal structure: distributions of inter-whistle intervals (A) and whistle sequence durations (B). C–E) Expert-annotated subset: temporal coverage (C), class distribution (D), and whistle duration (E). F) Signal-to-noise ratio (SNR) for the full dataset and the expert-annotated subset.
Figure 4: Evaluation Tasks. The two evaluation tasks: (1) whistle type classification, where isolated whistle segments are assigned a predefined category, and (2) whistle-type detection, where fixed 0.5 s segments are labeled with a category if a whistle is present, or categorized as background otherwise.
Method
Pretraining
Classification (%)
Detection (mAP)
Chance level
–
16.7
8.3
Spectral features
–
34.9±2.2
26.3±1.1
MFCCs
–
45.6±2.4
33.6±1.8
Mean spectrogram
–
55.6±2.4
47.7±2.1
AVES-core
General audio (AudioSet, FSD50K)
68.0±2.2
57.4±2.1
BioLingual
Audio-text (AnimalSpeak)
71.3±2.1
66.5±2.2
Table 2: Performance on whistle-type classification (accuracy) and detection (mAP). Models are grouped by representation: classical acoustic features, off-the-shelf pretrained bioacoustic models, and a Wav2Vec2.0 model trained on OpenWhistle. All use linear probing, a logistic regression classifier on frozen embeddings. Uncertainty is estimated via bootstrap resampling (N=1000), results are reported as mean and standard deviation.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: A) Photographs of the five dolphins present during the recording period. B) Family tree of the pod, indicating sex and signature whistles (SW). Dolphins present during the recording period are highlighted in red. Dolphins not present but whose signature whistles appear in the dataset are shown in blue.
Figure 6: Dataset construction pipeline. Raw audio is converted to 224 × 224 spectrograms and processed by a VGG16-based detector. Positive detections are segmented and then used in two branches: one branch groups detections into whistle sequences to construct the pretraining corpus, while the other applies F0 estimation, ARTwarp categorization, and expert annotation to construct the expert-annotated dataset.
Figure 7: A) Confusion matrix of the Wav2Vec2.0 model trained on OpenWhistle for whistle-type classification (in %). B) Three example spectrograms for each of the six whistle classes in the classification task.
Self-supervised learning (SSL) has opened new opportunities in bioacoustics by enabling scalable modeling of animal vocalizations without the need for expensive manual annotation. However, current SSL models in this domain prioritize broad generalization across species and are not optimized for uncovering the fine-grained structure of individual communication systems. In this work, we collect and release a novel dataset of over five years of longitudinal recordings, from five known dolphins in a semi-naturalistic marine environment, an unprecedented resource for studying dolphin communication. We adapt the Wav2Vec2.0 Baevski et al. (2020) architecture to this domain and introduce Dolph2Vec, the first large-scale, species-specific SSL model trained exclusively on this data. We benchmark our model on two biologically relevant tasks: signature whistle classification and whistle detection. Dolph2Vec significantly outperforms general-purpose baselines in both tasks. Beyond performance, we show that learned embeddings and codebook structure capture interpretable acoustic units aligned with dolphin whistle categories and possibly sub-whistle structure, enabling fine-grained analysis of communication patterns. Our findings demonstrate how SSL can serve as both a model and a scientific tool to explore hypotheses in animal communication research.
Chiara Semenzin, Faadil Mustun, Roberto Dessi +5
École Normale Supérieure, Paris, France · Not Diamond, San Francisco, USA · Institut du Cerveau, Paris, France +1
Bioacoustic datasets from tropical regions remain limited, in part due to the absence of reproducible workflows for aggregating recordings from public archives. We present \textbf{MyGardenBird}, a curated dataset of bird vocalisations representing twelve common species across Peninsular Malaysia and the Indo-Malayan region. Recordings were sourced from Xeno-canto and processed through species-level filtering, manual spectrogram segmentation, and quality control checks. The primary release comprises 7,200 manually validated audio clips (16 kHz, 16-bit PCM mono WAV), balanced at 600 three-second clips per species (6.0 hours total) derived from 1,381 distinct recordings. Metadata includes geospatial coordinates, vocalisation categories, and signal-to-noise ratio (SNR) values (range: 0.83--59.18 dB; mean: 15.80 dB). A supplementary 44.1 kHz version is also provided. To mitigate data leakage, dataset partitions are defined at the source-recording level. Baseline classification experiments using convolutional neural networks on Mel-spectrograms achieved test accuracies of 92--96%, indicating strong interspecies separability. Limitations include reliance on single-annotator curation; however, validation with BirdNET confirmed label consistency. MyGardenBird is openly available at https://doi.org/10.5281/zenodo.20306877 under a CC BY-NC-SA 4.0 licence. Complete preprocessing code accompanies the release to support reproducibility and future expansion.
Muhammad Mun'im Ahmad Zabidi, Mohd Yamani Idna Idris, Norisma Idris
Faculty of Computer Science and Information Technology, Universiti Malaya, 50603 Kuala Lumpur, Malaysia · Faculty of Electrical Engineering, Universiti Teknologi Malaysia, 81310 Johor Bahru, Johor, Malaysia
Passive acoustic monitoring enables continuous, non-invasive biodiversity assessment across diverse ecosystems. The scale of these datasets has driven the adoption of machine learning, with supervised approaches showing strong performance. However, supervised methods require time-resolved annotated datasets, which remain scarce, especially in complex tropical soundscapes. We present PteroSet, a curated dataset of strongly annotated Neotropical bird vocalizations recorded in Puerto Asis (Putumayo) and Pivijay (Magdalena), Colombia, between 2023 and 2025. The dataset comprises 563 recordings (73.62 h) and 15,372 time-frequency annotations, including 6,702 events identified to the species level across 168 species. We release the annotations in a COCO-inspired JSON schema that unifies audio files, taxonomic categories, and labels for machine learning workflows. Beyond providing annotated data, PteroSet serves as a realistic benchmark that highlights key characteristics of tropical soundscapes, including acoustic co-occurrence and domain shift across recording sites. We provide a deep learning baseline for binary bird detection, demonstrating PteroSet's usability and the challenges it presents.
Daniela Ruiz, Juan Sebastián Ulloa, Zhongqi Miao +11
Microsoft AI for Good Research Lab, Redmond, Washington, United States · Universidad de Los Andes, Bogotá, Colombia, Center for Research and Formation in Artificial Intelligence · Instituto de Investigación de Recursos Biológicos Alexander von Humboldt, Bogotá, Colombia +2