cs.SDFeb 23, 2026

AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration

Authors: Henry Zhong, Jörg M. Buchholz, Julian Maclaren, Simon Carlile, Richard F. Lyon

Organizations: Australian Hearing Hub, Macquarie University, Sydney, Australia · Google Research Australia, Sydney, Australia

Abstract

Manual annotation and clustering of audio datasets is labour intensive. We introduce AuditoryHuM, a training-free framework for the unsupervised discovery and clustering of auditory scene labels using human-Multimodal Large Language Model (MLLM) collaboration. Leveraging MLLMs, our framework generates contextually relevant labels for audio data. To ensure label quality and mitigate hallucinations, zero-shot learning (Human-CLAP) quantifies the alignment between generated text labels and raw audio. A targeted human-in-the-loop intervention, refines only the lowest aligned pairs. The discovered labels form an interpretable alignment vector to group audio into cohesive clusters. The framework was evaluated across three auditory scene datasets (ADVANCE, AHEAD-DS, and TAU 2019), achieving a 96.1% reduction in human labour for ADVANCE during testing. Downstream models trained on our clusters exhibit enhanced classification accuracy vs the baseline (0.79 to 0.84) in clean acoustic environments, though performance scales down in dense, complex soundscapes. The project page: https://github.com/Australian-Future-Hearing-Initiative

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

    Jul 22, 2026Siqian Tong, Xuan Li, Chaozhuo Li +5Audio UnderstandingAudio Understanding

  2. From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

    Jun 24, 2026Pengfei Zhang, Hoang H Nguyen, Kazi Shaharair Sharif +6Large Audio Language Models

  3. Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

    May 27, 2026Jiahao Mei, Heinrich Dinkel, Yadong Niu +7Modern Generative Audio ModelsAcoustic Latent Space