Manual annotation and clustering of audio datasets is labour intensive. We introduce AuditoryHuM, a training-free framework for the unsupervised discovery and clustering of auditory scene labels using human-Multimodal Large Language Model (MLLM) collaboration. Leveraging MLLMs, our framework generates contextually relevant labels for audio data. To ensure label quality and mitigate hallucinations, zero-shot learning (Human-CLAP) quantifies the alignment between generated text labels and raw audio. A targeted human-in-the-loop intervention, refines only the lowest aligned pairs. The discovered labels form an interpretable alignment vector to group audio into cohesive clusters. The framework was evaluated across three auditory scene datasets (ADVANCE, AHEAD-DS, and TAU 2019), achieving a 96.1% reduction in human labour for ADVANCE during testing. Downstream models trained on our clusters exhibit enhanced classification accuracy vs the baseline (0.79 to 0.84) in clean acoustic environments, though performance scales down in dense, complex soundscapes. The project page: https://github.com/Australian-Future-Hearing-Initiative
Figures & tables
Figure 1 : The dataflow diagram of the AuditoryHuM processes, showing the steps from MLLM label discovery, human-in-the-loop refinement, cleanup, alignment, and clustering.
μc
Gemma 3N E2B
Qwen 2 Audio 7B
Qwen 2.5 Omni 3B
ADVANCE
0.51
0.56
0.61
AHEAD-DS
0.42
0.51
0.55
TAU 2019
0.48
0.48
0.49
Table 1 : The μc values for each dataset and MLLM pair, rounded to two decimal places. μc ranges from 1 to −1 , a value of 1 represents perfect alignment and −1 is perfect misalignment. A higher value is better.
μ1%
Gemma 3N E2B
Qwen 2 Audio 7B
Qwen 2.5 Omni 3B
ADVANCE
-0.06
-0.14
0.13
AHEAD-DS
-0.19
-0.26
0.07
TAU 2019
0.07
0.00
0.10
Human Labelled
Human Labelled
Human Labelled
ADVANCE
0.30
0.31
0.30
AHEAD-DS
0.48
0.48
0.22
Table 2 : The μ1% values for each dataset and MLLM pair before and after relabelling by a human, rounded to two decimal places. μ1% represents the mean of the 1st percentile (bottom 1% ) of CLAP scores . μ1% ranges from 1 to −1 , a value of 1 represents perfect alignment and −1 is perfect misalignment. A higher value is better.
μc
AuditoryHuM
Zero-Shot
Captioning
ADVANCE
0.61
0.25
0.42
AHEAD-DS
0.55
0.40
0.46
TAU 2019
0.49
0.33
0.19
Table 3 : μc represents mean alignment of labels produced by AuditoryHum vs zero-shot using provided labels vs captioning. A higher value is better.
μ1%
Qwen 2.5 Omni 3B + Human-CLAP
Qwen 2.5 Omni 3B + LAION-CLAP
ADVANCE
0.13
0.07
AHEAD-DS
0.07
0.07
TAU 2019
0.10
0.13
Human Labelled + Human-CLAP
Human Labelled + LAION-CLAP
ADVANCE
0.30
0.10
AHEAD-DS
0.22
0.04
Table 4 : The μ1% values comparing the results from Human-CLAP and LAION-CLAP rounded to 2 decimal places. A higher value indicates better alignment between label and audio content.
Instances
Mean CLAP score Minimal Cleanup
Mean CLAP score Default Cleanup
ADVANCE Long Labels
578
0.60
0.57
ADVANCE Non-English Lab.
0
AHEAD-DS Long Labels
781
0.60
0.58
AHEAD-DS Non-Eng. Lab.
4
0.14
-0.07
TAU 2019 Long Labels
2726
0.50
0.49
TAU 2019 Non-English Lab.
2
0.13
0.18
Table 5 : The mean CLAP score values comparing the results with default or minimal label cleanup.
Figure 2 : Silhouette scores comparing AuditoryHuM vs zero-shot CLAP audio embeddings vs caption sentence transformer embeddings. A higher value is better.
Provided Labels mAP
Provided Labels Accuracy
AuditoryHuM Labels mAP
AuditoryHuM Labels Accuracy
ADVANCE
0.84
0.79
0.84
0.84
AHEAD-DS
0.97
0.98
0.83
0.89
TAU 2019
0.94
0.95
0.76
0.75
Table 6 : The mean average precision (mAP) and accuracy of the dataset provided and AuditoryHuM generated labels. Higher values are better.
Figure 3 : Comparison of CLAP scores between human and MLLM labels.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Prompt
Notes
Describe the auditory scene using word pairs. Separate each pair with a comma.
Initial prompt.
Provide a short sentence to describe this set of audio samples. The frequency distribution of individual labels for this set of audio samples is provided: <distribution>…</distribution> .
Prompt to generate a descriptive composite label for the entire cluster.
Describe this auditory scene.
Captioning prompt.
Appendix
Table 7 : The prompts for generating labels.
Step
Model or Algorithm
Params | ≈ MACs
Complexity
Notes
1
Label generation (Qwen 2.5 Omni 3B)
4.7035B | 971.995G
O(H⋅WlogW)+O(N⋅P+N2⋅D)+O(P+N⋅D)
Steps include STFT, encoder, and decoder. H is the number of hops. W is the window length. N is the total token length. P number of params. D is the hidden layer dimension.
Steps include STFT, audio encoder, label clean up on L labels, text encoder on L labels, and Cosine similarity on L labels. W is the window length. Ta is the audio token length. Tt is the text token length. D is the hidden layer dimension. E is the CLAP embedding length.
3
Select best aligned label
O(L)
Find the largest value in an unsorted sequence of length L .
Steps 1 to 3 are repeated A times. Where A is the number of audio samples.
Steps include STFT, audio encoder, label clean up on L labels, text encoder on L labels, and Cosine similarity on L labels. W is the window length. Ta is the audio token length. Tt is the text token length. D is the hidden layer dimension. E is the CLAP embedding length.
Step 4 is repeated up to A2 times to get alignment vectors.
Appendix
Table 8 : Breakdown of the computational cost and complexity of each step in AuditoryHuM, using multiply and accumulate computations (MACs) and Big-O.
Figure 4 : t-SNE visualisations using k∗ value for each dataset. Each unique colour represents a cluster of thematically related labels.
Dataset
Cluster Size / Total
Unique Labels
Composite Label for Cluster
ADVANCE
1582 / 5075
186
The audio samples consist mainly of ocean related sounds such as waves, thunder, and wind, with a few other elements like car, train, and aircraft sounds mixed in, along with various background noises and other natural sounds.
ADVANCE
1367 / 5075
375
The audio samples consist of a variety of sounds including thuds, frying noises, car sounds, choir voices, vibrations, horse trotting, train movement, drills, sewing machines, and other mechanical sounds, among others.
ADVANCE
983 / 5075
117
These audio samples contain various animal sounds, human activities, and natural elements, including bird vocalizations, insects, and weather sounds.
AHEADDS
4713 / 9968
198
This audio set contains a variety of sounds including clatter, child speaking, restaurant background noise, white noise, engine sounds, female conversations, dishes clinking, door slams, car passing, male laughter, crowd talk, and many other common everyday sounds, totaling 4713 samples.
AHEADDS
1668 / 9968
100
The audio set contains a diverse range of music genres and styles, including country, blues, heavy metal, opera, jazz, and more, with a focus on vocal performances and specific instruments like guitar and piano.
AHEADDS
1582 / 9968
93
The audio samples predominantly feature traffic noise, car passing, and various road-related sounds, with some additional elements like engine sounds, wind, and weather effects.
Appendix
Table 9 : Details of the 3 largest clusters for each dataset.
Component / Hyperparameter
Value / Strategy
Architecture
OpenYAMNet (Pretrained on AudioSet)
Optimiser
Adam
Loss Function
Focal Loss
Label Smoothing
0.1
Data Scheduling
Shuffled each epoch
Learning Rate Scheduler
Halved after 3 consecutive epochs without val improvement
Appendix
Table 10 : Downstream training configuration and hyperparameters.
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most players hear a reference audio, while one odd listener hears a subtle variant. The model first generates clues describing what it hears and then identifies the odd listener by reasoning over inconsistencies among clues. Since the odd listener is known by construction, the game provides verifiable rewards without any annotated answers. Experiments with Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA, MMAU Test-mini and MMAR show that Audio-Zero improves fine-grained audio reasoning while preserving broad audio understanding. Evolutionary and diagnostic analyses further reveal that increasingly fine-grained auditory descriptions emerge naturally from game pressure.
Siqian Tong, Xuan Li, Chaozhuo Li +5
Institute of Acoustics, Chinese Academy of Sciences · University of Chinese Academy of Sciences · 3Beijing Academy of Artificial Intelligence +3
Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes. Real-world auditory interpretation requires Context-Aware Auditory Scene Understanding (CASU): the ability to comprehend the holistic scene by integrating sound layers. To evaluate this capability, we introduce the CASU benchmark, which assesses whether Audio LLMs can interpret auditory scenes composed of speech, acoustic events (e.g., announcements), and background environments (e.g., traffic), and reason about the logical relationships between these layers. We propose a scalable pipeline for constructing time-accurate, semi-synthetic audio streams by composing real-world scene sounds with synthetic speech. Building on this data, we design four tasks that probe scene understanding: contextual question answering, entity extraction from the scene, speaker role inference, and counterfactual reasoning where scene is manipulated. Experiments across multiple LALMs demonstrate that effective auditory scene understanding requires integration over all auditory layers, rather than reliance on speech or sound alone, underscoring the necessity of CASU for advancing complex audio understanding in LALMs.
Pengfei Zhang, Hoang H Nguyen, Kazi Shaharair Sharif +6
University of California Irvine · University of Illinois Chicago · Kennesaw State University
Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a single description. The key obstacles are insufficient fine-grained supervision for real-world mixed audio and limited acoustic representations for modeling concurrent audio components. We present Dasheng AudioGen, a unified framework for generating general mixed-audio scenes from text. Dasheng AudioGen introduces structured multi-view captions, which explicitly decouple complex acoustic scenes into complementary description views, thereby enabling fine-grained control over audio layers. Furthermore, we employ a high-dimensional unified semantic-acoustic representation as the shared latent space. It injects semantic priors that facilitate cross-modal training convergence, while its high-dimensional feature space provides sufficient capacity to disentangle and fuse concurrent audio components effectively. With these designs, a simple flow-matching DiT achieves high-quality end-to-end audio scene generation. We also establish a comprehensive evaluation pipeline for audio scene generation. Experiments demonstrate that Dasheng AudioGen achieves performance approaching real-world recordings in mixed-audio categories, while remaining competitive with specialized models in single-type generation tasks. Demos are available at https://nieeim.github.io/Dasheng-AudioGen-Web/.
Jiahao Mei, Heinrich Dinkel, Yadong Niu +7
1X-LANCE Lab, Shanghai Jiao Tong University, Shanghai, China · 2MiLM Plus, Xiaomi Inc., Beijing, China