BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics
Authors: Christos Plachouras, David Robinson, Marius Miron, Gagan Narula, Paul Laisné, Anthony L. T. Fine, Benno Weck, Ellen Gilsenan-McMahon, +12 more
Organizations: Earth Species Project · Queen Mary University of London
Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models. However, progress in the field has so far been assessed on a narrow set of tasks, primarily centered on label-centric biological category recognition, such as species and call-type classification. In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning. Using BEANS-Next, we show that existing models exhibit limited performance beyond the task families emphasized by existing evaluations, constraining their usefulness for broader bioacoustic applications. To support progress on this broader task space, we also introduce ROOTS, a large-scale training resource built from expanded curated real-world data and previously underused behavioral and acoustic metadata, supplemented by audio-derived information and scalable synthetic generation where labeling is insufficient. We demonstrate that training on this dataset yields substantial progress across all task groups of BEANS-Next, moving audio-language models closer to their potential as general-purpose assistants for bioacoustics and ethology. To accelerate progress in the field, we open-source our benchmark, dataset, and data pipelines.
Figures & tables
Figure 1 : Overview of BEANS-Next and ROOTS , along with their underlying data sources and the shared four-tier task taxonomy that grounds them.
Model
CBI
HBDB
ESC50
Call-type
Lifestage
N-indv
AF-Next
0.006
0.004
0.768
0.078
0.048
0.579
Qwen3-Omni
0.040
0.038
0.797
0.018
0.054
0.531
NatureLM-audio v1.0
0.778
0.114
0.820
0.870
0.794
0.658
NatureLM-audio v1.1
0.808
0.062
0.360
0.882
0.753
0.501
Table 1: BEANS-Zero classification results (accuracy). Higher is better.
Model
HSN
NBP
NES
PER
POW
SNE
UHH
AVG
AF-Next
0.000
0.000
0.000
0.000
0.002
0.002
0.000
0.001
Qwen3-Omni
0.003
0.014
0.000
0.000
0.067
0.000
0.000
0.012
NatureLM-audio v1.0
0.273
0.601
0.435
0.557
0.891
0.656
0.339
0.536
NatureLM-audio v1.1
0.325
0.644
0.446
0.503
0.953
0.715
0.404
0.570
Table 2: BirdSet results (top-1 accuracy). Higher is better.
Acoustic description
SNR
Mean f0 MAE (Hz) ↓
Model
Caption
Desc. MCQ
Crow
MCQ
MAE (dB) ↓
Seen
Held-out
AF-Next
0.5
50.5
48.0
36.7
31.0 1.5%
5706 13.7%
268 1.6%
Qwen3-Omni
0.0
43.6
47.0
25.2
15.4 94.5%
3749 99.9%
208 100%
NatureLM-audio v1.0
0.3
28.4
33.5
34.6
22.6 92.2%
4086 100%
208 100%
NatureLM-audio v1.1
0.0
28.7
30.0
26.7
19.1 52.1%
4334 84.0%
360 91.1%
NatureLM-audio T1
2.8
97.9
27.5
60.8
10.4 100%
253 100%
125 100%
Table 3 : BEANS-Next Tier 1 (acoustic perception) results. Captioning uses CIDEr ( ×100 ); description matching and SNR MCQ use accuracy (%); SNR regression and mean f0 use MAE in dB and Hz, respectively. Higher is better unless marked ↓ . Small grey percentages give the parse rate (share of answers with a parseable number); MAE is computed over parsed answers only, and grey MAEs rest on fewer than half of the answers and are not bolded.
Taxon presence
Call-type presence
Model
Bird
Mamm.
Amph.
Insect
Alarm
Flight
Beg.
Call type
Behav.
Caption
AF-Next
71.2
67.1
81.2
75.5
56.2
46.5
54.9
44.4
56.5
0.1
Qwen3-Omni
68.4
65.0
85.7
73.4
66.8
57.0
65.7
25.6
56.7
0.1
NatureLM-audio v1.0
55.7
63.4
68.0
56.0
50.3
51.7
50.2
48.2
39.9
1.7
NatureLM-audio v1.1
54.7
65.2
53.2
59.3
58.8
71.5
62.5
39.5
33.7
1.1
NatureLM-audio T2
91.5
87.9
96.4
93.4
69.9
82.2
74.6
59.0
87.1
1.2
Table 4 : BEANS-Next Tier 2 (biological category recognition) results. Presence and behavior use accuracy (%), fixed-vocabulary call type uses macro-F1 (%) averaged over the five labels, and captioning uses CIDEr ( ×100 ). Higher is better.
Counting MAE ↓
Temporal
Scene / open
Model
#S
#V
#V / S
Referring
Overlap
Range
Sum.
Caption
AF-Next
1.64 97.2%
4.26 99.6%
2.15 52.5%
33.5
47.0
0.0
0.0
0.0
Qwen3-Omni
1.38 97.0%
3.93 99.9%
1.94 88.1%
25.0
46.3
0.0
0.0
0.0
NatureLM-audio v1.0
1.54 86.1%
4.09 100%
1.94 1.6%
28.0
59.0
0.0
0.1
0.0
NatureLM-audio v1.1
1.50 76.7%
4.57 93.4%
1.49 12.8%
30.0
47.1
0.0
2.4
0.1
NatureLM-audio T3
1.78 100%
3.81 100%
1.72 99.9%
48.0
61.7
11.3
5.0
1.6
Table 5 : BEANS-Next Tier 3 (scene understanding) results. Classification and temporal tasks use top-1 accuracy (%); counting tasks use mean absolute error (MAE); the frequency-range task (Range) uses coverage-aware species-band IoU (%, mean over ground-truth species, missed species scored 0); summary and captioning use corpus CIDEr ( ×100 ). S : species, V : vocalization, # : number of. Higher is better unless marked ↓ . Small grey percentages give the parse rate (share of answers with a parseable number); MAE is computed over parsed answers only, and grey MAEs rest on fewer than half of the answers and are not bolded.
Few-shot task
Model
Gibbons
Otters
DCASE
Crows
Unseen
AF-Next
0.007
0.255
0.119
0.225
0.230
Qwen3-Omni
0.175
0.535
0.327
0.620
0.480
NatureLM-audio v1.1
0.176
0.200
0.048
0.240
0.270
NatureLM-audio T4
0.243
0.645
0.452
0.715
0.650
Table 6 : BEANS-Next Tier 4 (in-context learning) results. Gibbons and DCASE use macro-F1 averaged over reference labels (A–C and A–D, respectively), with no detection represented by an empty set; other tasks use accuracy. Scores are on a 0–1 scale; higher is better. NatureLM-audio v1.0 does not accept multiple audio inputs.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Example Prompt
Datasets
Species ID (common name)
What species is vocalizing in this audio recording? Common name?
Given the context: ‘{{context}}’, what is the common name for the focal species?
Xeno-Canto, iNaturalist, Animal Sound Archive
Species ID + context (scientific)
Given the context: ‘{{context}}’, what is the scientific name for the focal species?
Xeno-Canto, iNaturalist, Animal Sound Archive
Full taxonomy + context
Given the context: ‘{{context}}’, provide the full taxonomic classification.
Xeno-Canto, iNaturalist, Animal Sound Archive
Genus classification
What is the genus of the focal species in the audio?
Xeno-Canto, iNaturalist, Animal Sound Archive
Appendix
Table 7 : Datasets used for each training task, with a representative prompt.
Figure 2 : Overview of the LLM-based data synthesis process.
Task
Example Prompt
Datasets
SNR category (fixed-bin MCQ)
What is the signal-to-noise ratio (SNR) of the bird call in this recording? Options: Clean, Good, Fair, Poor.
Xeno-Canto, iNaturalist
SNR category (custom-bin MCQ)
How would you classify the SNR of the bird call in this recording? Options use LLM-defined dB ranges.
Xeno-Canto
SNR threshold (binary)
Is the SNR of the bird call in this recording above {{threshold}} dB?
Xeno-Canto
SNR prediction (open-ended)
Estimate the signal-to-noise ratio (SNR) in dB of the sound event in this recording.
Xeno-Canto
Vocalization description MCQ
Which of the following best describes the vocalization in this recording? Options: {{acoustic_descriptions}}.
AnimalSpeak Pseudovox
F0-aware vocalization description MCQ
Which description best matches the sound in this recording? Options include qualitative acoustic character and F0 information.
F0 Bioacoustic
Appendix
Table 8 : LLM-generated datasets used for training tasks, with a representative prompt.
Figure 3 : ROOTS data sources, with the number of audio-language pairs and unique audio clips contributed by each (log scale).
Figure 4 : Lexical diversity of LLM-synthesized acoustic-description and behavior MCQ options. Left: number of unique word bigrams as more options are sampled. Right: bigram rank–frequency distribution (log–log). Acoustic descriptions keep introducing new bigrams and have a long, flat tail, whereas behavior options reuse a smaller vocabulary.
Table label
Task
Format
Samples
Metric
Tier 1: Acoustic
Acoustic description / Caption
Acoustic caption generation
Caption / OE
213
CIDEr
Acoustic description / Desc. MCQ
Field-note acoustic description matching
MCQ
290
Accuracy
Acoustic description / Crow
Carrion crow expert description matching
MCQ
200
Accuracy
SNR / MCQ
SNR category prediction
MCQ
607
Accuracy
SNR / MAE (dB)
SNR estimation
OE / regression
397
MAE (dB)
Appendix
Table 9 : BEANS-Next evaluation task labels used in result tables. Labels include the higher-level table group followed by the abbreviated column name used in the main results.
Figure 5 : Taxonomic class distribution of BEANS-Next samples with full taxonomic metadata (Tiers 2 and 4).
Figure 6 : Species frequency distribution across BEANS-Next : number of species by samples per species. Of the 2,698 unique species, 1,183 (44%) appear only once, while a few focal species (e.g., Crocuta crocuta , Corvus corone ) contribute hundreds of samples, stressing few-shot and zero-shot generalization rather than memorization of frequent taxa.
Figure 7 : Top 30 species by sample count in BEANS-Next . High-count species are predominantly focal study subjects from controlled recording campaigns (e.g., Corvus corone from the crow call-type dataset, Tadarida teniotis and Pipistrellus spp. from bat monitoring). The rapid drop-off after the top few species illustrates the long-tail character of the overall species distribution.
Capability
BEANS
BirdSet
WoW-Bench
BEANS-Next
Species recognition
Yes
Yes
No
Yes
Acoustic measurement
No
No
Yes
Yes
Temporal reasoning
No
No
No
Yes
Relational reasoning
No
No
No
Yes
Multi-audio reasoning
No
No
No
Yes
In-context learning
No
No
No
Yes
Appendix
Table 10: Comparison of benchmark coverage across bioacoustic audio-language capabilities. Existing benchmarks primarily evaluate species recognition or isolated acoustic perception, while BEANS-Next evaluates a broader capability space required by ethological workflows.
Acoustic description
SNR
Mean f0 MAE (Hz) ↓
Model
Caption
Desc. MCQ
Crow
MCQ
MAE (dB) ↓
Seen
Held-out
Real audio
AF-Next
0.4
50.1
47.5
33.9
34.6 8.8%
5652 13.3%
249 1.8%
Qwen3-Omni
0.0
47.0
45.5
7.7
20.0 82.6%
3728 99.5%
203 99.1%
NatureLM-audio v1.0
0.3
23.6
33.5
31.9
24.9 97.4%
4087 100%
208 100%
Gaussian noise
Appendix
Table 11 : BEANS-Next acoustic and biological category recognition controls. Rows are grouped by input: real audio, duration-matched Gaussian noise, or no audio with ( informed ) or without ( silent ) the missing-audio instruction. Metrics and formatting follow Tables 3 and 4 : higher is better unless marked ↓ , small grey percentages give the parse rate, MAEs based on fewer than half of the answers are shown in grey, and – marks no parseable answer. Insect and begging-call presence use 636 and 586 examples, respectively. † Retained inference failures: three Llama 3.1 8B seen- f0 examples and one AF-Next caption per audio condition.
Counting MAE ↓
Model
#S
#V
#V / S
#V MCQ
Real audio
AF-Next
1.65 97.4%
4.23 99.6%
3.08 53.2%
26.0
Qwen3-Omni
1.20 93.5%
3.82 99.8%
1.92 84.9%
51.8
NatureLM-audio v1.0
1.00 41.3%
6.58 24.8%
2.11 1.7%
0.0
Gaussian noise
Appendix
Table 12 : BEANS-Next scene-understanding controls; inputs and formatting as in Table 11 , metrics as in Table 5 . #V MCQ (accuracy) and Presence (vocalization presence, accuracy) are additional Tier 3 tasks not shown in Table 5 . For Range, unmatched or unparseable examples score zero. NatureLM-audio v1.0 segment-wise responses are scored as returned.
Open-ended
MCQ
Model
Order
High
Low
Long
Freq.
Order
High
Low
Long
Freq.
Real audio
AF-Next
0.0
0.0
0.0
0.0
0.0
27.0
32.2
46.0
32.6
24.5
Qwen3-Omni
0.0
0.0
0.0
0.0
0.0
25.7
33.6
19.7
19.2
13.4
NatureLM-audio v1.0
0.0
0.0
0.0
0.0
0.0
0.1
0.0
0.0
0.0
0.0
Gaussian noise
Appendix
Table 13 : BEANS-Next species-ID controls; inputs, metrics, and formatting as in Tables 5 and 12 . † Qwen2.5 7B silent lowest-pitch identification uses 799 of 803 responses.
Few-shot task
Model
Gibbons
Otters
DCASE
Crows
Unseen
Zebra
Real audio
AF-Next
0.028
0.174
0.141
0.225
0.046
0.300
Qwen3-Omni
0.196
0.492
0.323
0.580
0.486
0.200
Gaussian noise
AF-Next
0.000
0.286
0.000
0.250
0.050
0.275
Appendix
Table 14 : BEANS-Next in-context learning controls and parse rates; inputs and formatting as in Table 11 . (a) Gibbons and DCASE use macro-F1; other tasks use accuracy, all on a 0–1 scale. Zebra is an additional matching task not shown in Table 6 . NatureLM-audio v1.0 does not accept multiple audio inputs. (b) Parse rate (%) for categorical answers, pooled over attempted examples: T1 MCQ ( n=3,468 , three tasks), T2 behavior ( 3,657 ), T3 MCQ ( 5,807 , seven), T4 matching ( 958 , four), call type ( 2,000 ), and T4 detection ( 4,026 , two). A parseable choice or label set need not be correct; an explicit no-detection answer counts where allowed. Binary presence and captioning are excluded. –: not applicable.
Split / task family
Format
Prompt
Reference output
T1 vocalization description
MCQ
Which description best matches the vocalization? Options include a slow trill, low drone, sharp whistle, and staccato notes with nasal rise.
a series of rapid, repetitive, staccato notes followed by a nasal, rising inflection
T1 vocalization description
MCQ
Which description best matches the sound? Options include impulsive bursts, a melodic whistle, sharp staccato notes, and a low nasal moan.
A series of dense, impulsive bursts followed by a quiet, rapid trill and punctuated by loud clicks
T1 acoustic caption
Caption
Describe the acoustic character of this vocalization, including its frequency range.
The vocalizations consist of high-pitched, whistling one-note sounds spanning roughly 2,500 to 7,000 Hz, as well as buzzing, nasal one-note sounds ranging from 2,000 to 18,000 Hz.
T1 acoustic caption
Caption
Describe the acoustic character of this vocalization, including its recording quality.
This vocalization consists of a series of rapid rattles that increase in frequency, concluding with a single prolonged rattle. The sound is difficult to hear over background noise but becomes more audible towards the end of the recording.
T2 behavior
MCQ
Which acoustic behavior is the Flame-crested Manakin performing? Options: wing snap, no vocalization/silent, song.
Wing snap
T2 behavior
MCQ
Which vocalizations can be heard from the blue duck? Options: rapid drumming, continuous song, male whistle followed by female growl.
A male whistle followed by a female growl
Appendix
Table 15 : BEANS-Next benchmark examples. Audio placeholders are omitted or compacted; each row shows the split/task family, prompt, and reference output used for evaluation.
Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features these models encode, nor which are useful for a given task. This hinders transparency and limits extension to rare species or data-scarce domains. Here we reveal which speech-like features are encoded in bioacoustic representations. Using the 88~eGeMAPS features across six taxonomic groups, we apply linear and nonlinear regression probes to quantify which acoustic properties each model captures. Results confirm a ``no free lunch'' pattern: no single model captures the full feature space. A concatenated embedding achieves the highest performance, suggesting complementary acoustic space coverage across models. Loudness features are best encoded (R2=0.76) while F0 is hardest to recover (R2=0.33). By cross-referencing recoverability with per-species feature salience (NMI), we derive data-driven model selection guidance for bioacoustics.
Training data for bioacoustics is scattered across taxa, regions, and institutions. Centralizing it all is often infeasible. We show that independently fine-tuned BEATs encoders can be composed into a unified 661-species classifier via task vector arithmetic without sharing data. We find that bioacoustic task vectors are near-orthogonal (cosine 0.01-0.09). Their separation aligns closely with spectral distribution distance, a gradient consistent with the acoustic niche hypothesis. This geometry makes simple averaging optimal while sign-conflict methods reduce accuracy by one to six percentage points. Composition also creates an asymmetric gap: species-rich groups lose accuracy relative to joint training while underrepresented taxa gain, a redistribution useful for equitable biodiversity monitoring. We verify linear mode connectivity across all taxonomic pairs, demonstrate zero-shot transfer to new regions, and identify domain negation as a boundary condition where composition fails. These results enable a collaborative paradigm for bioacoustics where institutions share only task vectors to assemble multi-taxa classifiers, preserving data privacy.
Ragib Amin Nihal, Benjamin Yen, Runwu Shi +2
Systems and Control Engineering, Institute of Science Tokyo, Japan · RIKEN BDR, Japan
Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating representation learning. Most bioacoustic benchmarks use a fixed, low-capacity probe, such as a linear layer on the final encoder layer. While this standardization enables model comparisons, it may bias results by overlooking the interaction between encoder features and probe design. In this work, we systematically study different probing strategies across two bioacoustic benchmarks, BEANs and BirdSet. We evaluate last- and multi-layer probing, across linear and attention probes. We show that larger probe heads that leverage time information have superior performance. Our results suggest that current benchmarks may misrepresent encoder quality when relying on a last-layer probing setup. Multi-layer probing improves downstream task performance across all tested models, while attention probing has superior performance to linear probing for transformer models.