cs.SDOct 7, 2026

CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound

Authors: Marcel Gibier, Thomas Thebaud, Olivier Boëffard, Jean-François Bonastre

Organizations: Inria, Paris, France · AMIAD, France

Abstract

Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.

Figures & tables

Explore similar work

CardsList
  1. From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

    Jun 24, 2026Pengfei Zhang, Hoang H Nguyen, Kazi Shaharair Sharif +6Large Audio Language Models

  2. TAC: Timestamped Audio Captioning

    Feb 17, 2026Sonal Kumar, Prem Seetharaman, Ke Chen +8

  3. Logbook: Extremely Long-form Audio Event Understanding

    Oct 5, 2026Kwanghee Choi, Suwon Shon, Dmitriy Serdyuk +7Audio Understanding