Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.
Figures & tables
Figure 1: Four types of speaker reaction to an audio event. Each block shows the transcript (T, one box per word) and the speech signal (S) over time. The shaded box marks the audio event [ton,toff] . Colored elements are altered by the event. Pivot : persistent change of T and S. Verbal : local verbal reaction (T and S). Behavioral : local change of S only. Ambient : no impact. The speech signal, the audio event ( −23 LUFS) and the background ( −32 LUFS) are summed ( ⊕ ) to form the final mix.
Figure 2: The CARES construction pipeline: text generation and audio rendering. A deterministic allocator fixes the scene, relation, speaker genders, sound events, reaction registers, and composition. A language model fills only the conversation topic and the dialogue. The reaction-recovery filter sits between the two phases: rejected dialogues are regenerated until an independent evaluator recovers the intended reactions. utt. : utterances; ev. : acoustic events.
Reaction Classification
Acoustic Scene
Audio
Summary
Balanced
Precision(%) / Recall(%)
Classification
Tagging
Recall (%) ↑
Model
Accuracy(%) ↑
Pivot
Verbal
Behav.
Ambient
Accuracy (%) ↑
Accuracy (%) ↑
Pivot
Verbal
Behav.
Ambient
Audio Flamingo Next
29.7±2.0
50.0/0.3
30.6/65.0
50.0/0.7
41.1/52.7
62.5±3
34.8±2
9.5±3.7
3.3±1.7
0.5±1.0
2.5±1.5
Qwen3-Omni
39.0±2.0
27.0/89.5
33.6/41.7
41.5/24.2
75.0/0.5
89,6±2
57.7±2
22.2±4.9
6.0±2.2
1.6±1.4
3.3±1.7
Kimi-Audio
25.6±2.0
27.8/1.6
29.6/1.3
22.2/5.4
31.2/94.1
78.1±2
48.6±2
6.7±3.3
1.7±1.4
0.5±1.0
0.8±1.0
MiMo-Audio
32.8±2.0
40.6/40.0
30.4/91.2
—/0.0
—/0.0
81.7±2
46.5±2
12.0±4.0
4.0±2.0
1.0±1.0
2.0±1.2
Table 1: Benchmark on the 1,000 test scenes. Whisper + LLM is the text-only control. All scores in %, chance is 25 for all accuracies. Balanced accuracy is the mean of the four per-class accuracies. Recall and precision are provided for each reaction class. Acoustic scene classification and audio tagging are four-way multiple choice. Summary recall refers to the proportion of annotated sounds of this type that are mentioned in the 200-words summary generated by the model. ( ± ) give the 95 % confidence interval. Best per column in bold , except for per-class precision/recall. Results significantly higher than the text-only control are in blue . Classes with a Reaction Classification recall ≤7% are shown in red .
Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes. Real-world auditory interpretation requires Context-Aware Auditory Scene Understanding (CASU): the ability to comprehend the holistic scene by integrating sound layers. To evaluate this capability, we introduce the CASU benchmark, which assesses whether Audio LLMs can interpret auditory scenes composed of speech, acoustic events (e.g., announcements), and background environments (e.g., traffic), and reason about the logical relationships between these layers. We propose a scalable pipeline for constructing time-accurate, semi-synthetic audio streams by composing real-world scene sounds with synthetic speech. Building on this data, we design four tasks that probe scene understanding: contextual question answering, entity extraction from the scene, speaker role inference, and counterfactual reasoning where scene is manipulated. Experiments across multiple LALMs demonstrate that effective auditory scene understanding requires integration over all auditory layers, rather than reliance on speech or sound alone, underscoring the necessity of CASU for advancing complex audio understanding in LALMs.
Pengfei Zhang, Hoang H Nguyen, Kazi Shaharair Sharif +6
University of California Irvine · University of Illinois Chicago · Kennesaw State University
Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introduce Timestamped Audio Captioner (TAC), a model that produces temporally grounded audio descriptions at varying degrees of detail and resolution. TAC is trained with a synthetic data pipeline that constructs challenging and dynamic mixtures from real-world audio sources, enabling robust learning under realistic polyphonic conditions. Across event detection and dense captioning, TAC outperforms all competing methods, with a low hallucination rate and accurate temporal grounding. We also introduce TAC-V, an audio-visual pipeline to generate semantically rich audio-visual descriptions. We then show that TAC and TAC-V serves as a "semantic bridge" for a text-only reasoner: a simple TAC→LLM and TAC-V→LLM cascade achieves state-of-the-art scores on benchmarks for both audio (MMAU-Pro, MMSU, MMAR) and audio-visual (DailyOmni, VideoHolmes) understanding and reasoning respectively.
Sonal Kumar, Prem Seetharaman, Ke Chen +8
University of Maryland, College Park, USA · Adobe Research, USA · OpenAI, USA (work done while at Adobe).
Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.