cs.SDSep 20, 2026

Misrecognition or Abstraction? Rethinking Outputs of Sound Event Recognition

Authors: Naoya Tomida, Yuki Okamoto, Keisuke Imoto

Organizations: Kyoto University, Japan · The University of Tokyo, Japan

Abstract

Conventional general sound recognition systems typically output deterministic sound event labels, implicitly assuming that the target sound class can be correctly identified from the input audio. However, in real listening situations, the sound event class is not always clearly identifiable. Human listeners may nevertheless understand their surroundings from an ambiguous sound without identifying its exact sound event class. This motivates a discussion of how the outputs of sound recognition systems should be redesigned under such uncertainty. As a basis for this discussion, this paper proposes an output representation for sound event recognition that combines a sound event class, its confidence score, and an onomatopoeic description of the sound. The proposed representation preserves conventional class-based recognition while providing an additional onomatopoeic description of acoustic characteristics that can remain informative even when the class prediction is uncertain. Experiments using ESC-50 and ESC-50-Onomatopoeia show that the proposed method achieves sound recognition performance comparable to that of a conventional recognition-only system. In addition, an LLM-as-a-judge evaluation and subjective listening experiments indicate that the proposed output is preferred over conventional deterministic outputs based on the sound event label, particularly when used to support understanding of the surrounding environment. These results suggest that such output representations can make sound event recognition more informative and communicative under uncertainty.

Figures & tables

Explore similar work

May 5, 2026cs.SD

Towards Open World Sound Event Detection

Sound Event Detection (SED) plays a vital role in audio understanding, with applications in surveillance, smart cities, healthcare, and multimedia indexing. However, conventional SED systems operate under a closed-world assumption, limiting their effectiveness in real-world environments where novel acoustic events frequently emerge. Inspired by the success of open-world learning in computer vision, we introduce the Open-World Sound Event Detection (OW-SED) paradigm, where models must detect known events, identify unseen ones, and incrementally learn from them. To tackle the unique challenges of OW-SED, such as overlapping and ambiguous events, we propose a 1D Deformable architecture that leverages deformable attention to adaptively focus on salient temporal regions. Furthermore, we design a novel Open-World Deformable Sound Event Detection Transformer (WOOT) framework incorporating feature disentanglement to separate class-specific and class-agnostic representations, together with a one-to-many matching strategy and a diversity loss to enhance representation diversity. Experimental results demonstrate that our method achieves marginally superior performance compared to existing leading techniques in closed-world settings and significantly improves over existing baselines in open-world scenarios.
Sep 28, 2026cs.SD

Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions

Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a breakfast emerge compositionally from temporally distributed sound events, requiring abstraction beyond the event-centric granularity that dominates current training and evaluation paradigms. Our benchmark evaluates a wide set of LALMs under a principled framework that tests how language-based reasoning, grounded in acoustic perception, structures sound abstractions into higher-level understanding. By systematically varying exemplar typicality and distractor similarity, our evaluation exposes \added{that current models do not reliably perform compositional inference from atomic acoustic events to higher-level human activities solely from audio.} All data, taxonomies, and evaluation scripts are publicly available on our companion website: https://alm-sounding-actions.onrender.com/
Jun 5, 2026cs.SD

Towards Event-Robust Acoustic Scene Classification

This paper introduces the Event-Shifted Acoustic Scene (ESAS) dataset, a novel benchmark for evaluating the robustness of Acoustic Scene Classification (ASC) systems against unknown sound events. Existing ASC datasets typically contain recordings of clean and consistent audio, while real-world environments often include diverse and unexpected sound events. To bridge this gap, ESAS simulates real-world acoustic variability by injecting foreground sound events into background scenes with the assistance of large language models. In this work, we present the construction methodology, dataset statistics, and evaluation protocols. Furthermore, a comprehensive evaluation of state-of-the-art ASC systems is conducted using the ESAS benchmark. Experimental results reveal that existing ASC models suffer significant performance degradation when facing the event-shift challenge. The introduction of the ESAS dataset aims to drive future research toward event-robust ASC.