CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound
Organizations: Inria, Paris, France · AMIAD, France
Abstract
Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.
Figures & tables
| Reaction Classification | Acoustic Scene | Audio | Summary | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Balanced | Precision(%) / Recall(%) | Classification | Tagging | Recall (%) | |||||||
| Model | Accuracy(%) | Pivot | Verbal | Behav. | Ambient | Accuracy (%) | Accuracy (%) | Pivot | Verbal | Behav. | Ambient |
| Audio Flamingo Next | 50.0/0.3 | 30.6/65.0 | 50.0/0.7 | 41.1/52.7 | |||||||
| Qwen3-Omni | 27.0/89.5 | 33.6/41.7 | 41.5/24.2 | 75.0/0.5 | |||||||
| Kimi-Audio | 27.8/1.6 | 29.6/1.3 | 22.2/5.4 | 31.2/94.1 | |||||||
| MiMo-Audio | 40.6/40.0 | 30.4/91.2 | —/0.0 | —/0.0 | |||||||