BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics
Organizations: Earth Species Project · Queen Mary University of London
Abstract
Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models. However, progress in the field has so far been assessed on a narrow set of tasks, primarily centered on label-centric biological category recognition, such as species and call-type classification. In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning. Using BEANS-Next, we show that existing models exhibit limited performance beyond the task families emphasized by existing evaluations, constraining their usefulness for broader bioacoustic applications. To support progress on this broader task space, we also introduce ROOTS, a large-scale training resource built from expanded curated real-world data and previously underused behavioral and acoustic metadata, supplemented by audio-derived information and scalable synthetic generation where labeling is insufficient. We demonstrate that training on this dataset yields substantial progress across all task groups of BEANS-Next, moving audio-language models closer to their potential as general-purpose assistants for bioacoustics and ethology. To accelerate progress in the field, we open-source our benchmark, dataset, and data pipelines.
Figures & tables
| Model | CBI | HBDB | ESC50 | Call-type | Lifestage | N-indv |
|---|---|---|---|---|---|---|
| AF-Next | 0.006 | 0.004 | 0.768 | 0.078 | 0.048 | 0.579 |
| Qwen3-Omni | 0.040 | 0.038 | 0.797 | 0.018 | 0.054 | 0.531 |
| NatureLM-audio v1.0 | 0.778 | 0.114 | 0.820 | 0.870 | 0.794 | 0.658 |
| NatureLM-audio v1.1 | 0.808 | 0.062 | 0.360 | 0.882 | 0.753 | 0.501 |
| Model | HSN | NBP | NES | PER | POW | SNE | UHH | AVG |
|---|---|---|---|---|---|---|---|---|
| AF-Next | 0.000 | 0.000 | 0.000 | 0.000 | 0.002 | 0.002 | 0.000 | 0.001 |
| Qwen3-Omni | 0.003 | 0.014 | 0.000 | 0.000 | 0.067 | 0.000 | 0.000 | 0.012 |
| NatureLM-audio v1.0 | 0.273 | 0.601 | 0.435 | 0.557 | 0.891 | 0.656 | 0.339 | 0.536 |
| NatureLM-audio v1.1 | 0.325 | 0.644 | 0.446 | 0.503 | 0.953 | 0.715 | 0.404 | 0.570 |
| Acoustic description | SNR | Mean MAE (Hz) | |||||
|---|---|---|---|---|---|---|---|
| Model | Caption | Desc. MCQ | Crow | MCQ | MAE (dB) | Seen | Held-out |
| AF-Next | 0.5 | 50.5 | 48.0 | 36.7 | 31.0 1.5% | 5706 13.7% | 268 1.6% |
| Qwen3-Omni | 0.0 | 43.6 | 47.0 | 25.2 | 15.4 94.5% | 3749 99.9% | 208 100% |
| NatureLM-audio v1.0 | 0.3 | 28.4 | 33.5 | 34.6 | 22.6 92.2% | 4086 100% | 208 100% |
| NatureLM-audio v1.1 | 0.0 | 28.7 | 30.0 | 26.7 | 19.1 52.1% | 4334 84.0% | 360 91.1% |
| NatureLM-audio T1 | 2.8 | 97.9 | 27.5 | 60.8 | 10.4 100% | 253 100% | 125 100% |
| Taxon presence | Call-type presence | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | Bird | Mamm. | Amph. | Insect | Alarm | Flight | Beg. | Call type | Behav. | Caption |
| AF-Next | 71.2 | 67.1 | 81.2 | 75.5 | 56.2 | 46.5 | 54.9 | 44.4 | 56.5 | 0.1 |
| Qwen3-Omni | 68.4 | 65.0 | 85.7 | 73.4 | 66.8 | 57.0 | 65.7 | 25.6 | 56.7 | 0.1 |
| NatureLM-audio v1.0 | 55.7 | 63.4 | 68.0 | 56.0 | 50.3 | 51.7 | 50.2 | 48.2 | 39.9 | 1.7 |
| NatureLM-audio v1.1 | 54.7 | 65.2 | 53.2 | 59.3 | 58.8 | 71.5 | 62.5 | 39.5 | 33.7 | 1.1 |
| NatureLM-audio T2 | 91.5 | 87.9 | 96.4 | 93.4 | 69.9 | 82.2 | 74.6 | 59.0 | 87.1 | 1.2 |
| Counting MAE | Temporal | Scene / open | ||||||
|---|---|---|---|---|---|---|---|---|
| Model | #S | #V | #V / S | Referring | Overlap | Range | Sum. | Caption |
| AF-Next | 1.64 97.2% | 4.26 99.6% | 2.15 52.5% | 33.5 | 47.0 | 0.0 | 0.0 | 0.0 |
| Qwen3-Omni | 1.38 97.0% | 3.93 99.9% | 1.94 88.1% | 25.0 | 46.3 | 0.0 | 0.0 | 0.0 |
| NatureLM-audio v1.0 | 1.54 86.1% | 4.09 100% | 1.94 1.6% | 28.0 | 59.0 | 0.0 | 0.1 | 0.0 |
| NatureLM-audio v1.1 | 1.50 76.7% | 4.57 93.4% | 1.49 12.8% | 30.0 | 47.1 | 0.0 | 2.4 | 0.1 |
| NatureLM-audio T3 | 1.78 100% | 3.81 100% | 1.72 99.9% | 48.0 | 61.7 | 11.3 | 5.0 | 1.6 |
| Few-shot task | |||||
|---|---|---|---|---|---|
| Model | Gibbons | Otters | DCASE | Crows | Unseen |
| AF-Next | 0.007 | 0.255 | 0.119 | 0.225 | 0.230 |
| Qwen3-Omni | 0.175 | 0.535 | 0.327 | 0.620 | 0.480 |
| NatureLM-audio v1.1 | 0.176 | 0.200 | 0.048 | 0.240 | 0.270 |
| NatureLM-audio T4 | 0.243 | 0.645 | 0.452 | 0.715 | 0.650 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Example Prompt | Datasets |
|---|---|---|
| Species ID (common name) | What species is vocalizing in this audio recording? Common name? | Xeno-Canto, iNaturalist, Watkins, Animal Sound Archive |
| Species ID (scientific name) | What is the scientific name of the focal species in the audio? | Xeno-Canto, iNaturalist, Watkins, Animal Sound Archive |
| Species ID + context (common) | Given the context: ‘{{context}}’, what is the common name for the focal species? | Xeno-Canto, iNaturalist, Animal Sound Archive |
| Species ID + context (scientific) | Given the context: ‘{{context}}’, what is the scientific name for the focal species? | Xeno-Canto, iNaturalist, Animal Sound Archive |
| Full taxonomy + context | Given the context: ‘{{context}}’, provide the full taxonomic classification. | Xeno-Canto, iNaturalist, Animal Sound Archive |
| Genus classification | What is the genus of the focal species in the audio? | Xeno-Canto, iNaturalist, Animal Sound Archive |
| Task | Example Prompt | Datasets |
|---|---|---|
| SNR category (fixed-bin MCQ) | What is the signal-to-noise ratio (SNR) of the bird call in this recording? Options: Clean, Good, Fair, Poor. | Xeno-Canto, iNaturalist |
| SNR category (custom-bin MCQ) | How would you classify the SNR of the bird call in this recording? Options use LLM-defined dB ranges. | Xeno-Canto |
| SNR threshold (binary) | Is the SNR of the bird call in this recording above {{threshold}} dB? | Xeno-Canto |
| SNR prediction (open-ended) | Estimate the signal-to-noise ratio (SNR) in dB of the sound event in this recording. | Xeno-Canto |
| Vocalization description MCQ | Which of the following best describes the vocalization in this recording? Options: {{acoustic_descriptions}}. | AnimalSpeak Pseudovox |
| F0-aware vocalization description MCQ | Which description best matches the sound in this recording? Options include qualitative acoustic character and F0 information. | F0 Bioacoustic |
| Table label | Task | Format | Samples | Metric |
|---|---|---|---|---|
| Tier 1: Acoustic | ||||
| Acoustic description / Caption | Acoustic caption generation | Caption / OE | 213 | CIDEr |
| Acoustic description / Desc. MCQ | Field-note acoustic description matching | MCQ | 290 | Accuracy |
| Acoustic description / Crow | Carrion crow expert description matching | MCQ | 200 | Accuracy |
| SNR / MCQ | SNR category prediction | MCQ | 607 | Accuracy |
| SNR / MAE (dB) | SNR estimation | OE / regression | 397 | MAE (dB) |
| Capability | BEANS | BirdSet | WoW-Bench | BEANS-Next |
|---|---|---|---|---|
| Species recognition | Yes | Yes | No | Yes |
| Acoustic measurement | No | No | Yes | Yes |
| Temporal reasoning | No | No | No | Yes |
| Relational reasoning | No | No | No | Yes |
| Multi-audio reasoning | No | No | No | Yes |
| In-context learning | No | No | No | Yes |
| Acoustic description | SNR | Mean MAE (Hz) | |||||
| Model | Caption | Desc. MCQ | Crow | MCQ | MAE (dB) | Seen | Held-out |
| Real audio | |||||||
| AF-Next | 0.4 | 50.1 | 47.5 | 33.9 | 34.6 8.8% | 5652 13.3% | 249 1.8% |
| Qwen3-Omni | 0.0 | 47.0 | 45.5 | 7.7 | 20.0 82.6% | 3728 99.5% | 203 99.1% |
| NatureLM-audio v1.0 | 0.3 | 23.6 | 33.5 | 31.9 | 24.9 97.4% | 4087 100% | 208 100% |
| Gaussian noise | |||||||
| Counting MAE | ||||
| Model | #S | #V | #V / S | #V MCQ |
| Real audio | ||||
| AF-Next | 1.65 97.4% | 4.23 99.6% | 3.08 53.2% | 26.0 |
| Qwen3-Omni | 1.20 93.5% | 3.82 99.8% | 1.92 84.9% | 51.8 |
| NatureLM-audio v1.0 | 1.00 41.3% | 6.58 24.8% | 2.11 1.7% | 0.0 |
| Gaussian noise | ||||
| Open-ended | MCQ | |||||||||
| Model | Order | High | Low | Long | Freq. | Order | High | Low | Long | Freq. |
| Real audio | ||||||||||
| AF-Next | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 27.0 | 32.2 | 46.0 | 32.6 | 24.5 |
| Qwen3-Omni | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 25.7 | 33.6 | 19.7 | 19.2 | 13.4 |
| NatureLM-audio v1.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.1 | 0.0 | 0.0 | 0.0 | 0.0 |
| Gaussian noise | ||||||||||
| Few-shot task | ||||||
| Model | Gibbons | Otters | DCASE | Crows | Unseen | Zebra |
| Real audio | ||||||
| AF-Next | 0.028 | 0.174 | 0.141 | 0.225 | 0.046 | 0.300 |
| Qwen3-Omni | 0.196 | 0.492 | 0.323 | 0.580 | 0.486 | 0.200 |
| Gaussian noise | ||||||
| AF-Next | 0.000 | 0.286 | 0.000 | 0.250 | 0.050 | 0.275 |
| Split / task family | Format | Prompt | Reference output |
|---|---|---|---|
| T1 vocalization description | MCQ | Which description best matches the vocalization? Options include a slow trill, low drone, sharp whistle, and staccato notes with nasal rise. | a series of rapid, repetitive, staccato notes followed by a nasal, rising inflection |
| T1 vocalization description | MCQ | Which description best matches the sound? Options include impulsive bursts, a melodic whistle, sharp staccato notes, and a low nasal moan. | A series of dense, impulsive bursts followed by a quiet, rapid trill and punctuated by loud clicks |
| T1 acoustic caption | Caption | Describe the acoustic character of this vocalization, including its frequency range. | The vocalizations consist of high-pitched, whistling one-note sounds spanning roughly 2,500 to 7,000 Hz, as well as buzzing, nasal one-note sounds ranging from 2,000 to 18,000 Hz. |
| T1 acoustic caption | Caption | Describe the acoustic character of this vocalization, including its recording quality. | This vocalization consists of a series of rapid rattles that increase in frequency, concluding with a single prolonged rattle. The sound is difficult to hear over background noise but becomes more audible towards the end of the recording. |
| T2 behavior | MCQ | Which acoustic behavior is the Flame-crested Manakin performing? Options: wing snap, no vocalization/silent, song. | Wing snap |
| T2 behavior | MCQ | Which vocalizations can be heard from the blue duck? Options: rapid drumming, continuous song, male whistle followed by female growl. | A male whistle followed by a female growl |