Manual annotation and clustering of audio datasets is labour intensive. We introduce AuditoryHuM, a training-free framework for the unsupervised discovery and clustering of auditory scene labels using human-Multimodal Large Language Model (MLLM) collaboration. Leveraging MLLMs, our framework generates contextually relevant labels for audio data. To ensure label quality and mitigate hallucinations, zero-shot learning (Human-CLAP) quantifies the alignment between generated text labels and raw audio. A targeted human-in-the-loop intervention, refines only the lowest aligned pairs. The discovered labels form an interpretable alignment vector to group audio into cohesive clusters. The framework was evaluated across three auditory scene datasets (ADVANCE, AHEAD-DS, and TAU 2019), achieving a 96.1% reduction in human labour for ADVANCE during testing. Downstream models trained on our clusters exhibit enhanced classification accuracy vs the baseline (0.79 to 0.84) in clean acoustic environments, though performance scales down in dense, complex soundscapes. The project page: https://github.com/Australian-Future-Hearing-Initiative
Figures & tables
Figure 1 : The dataflow diagram of the AuditoryHuM processes, showing the steps from MLLM label discovery, human-in-the-loop refinement, cleanup, alignment, and clustering.
μc
Gemma 3N E2B
Qwen 2 Audio 7B
Qwen 2.5 Omni 3B
ADVANCE
0.51
0.56
0.61
AHEAD-DS
0.42
0.51
0.55
TAU 2019
0.48
0.48
0.49
Table 1 : The μc values for each dataset and MLLM pair, rounded to two decimal places. μc ranges from 1 to −1 , a value of 1 represents perfect alignment and −1 is perfect misalignment. A higher value is better.
μ1%
Gemma 3N E2B
Qwen 2 Audio 7B
Qwen 2.5 Omni 3B
ADVANCE
-0.06
-0.14
0.13
AHEAD-DS
-0.19
-0.26
0.07
TAU 2019
0.07
0.00
0.10
Human Labelled
Human Labelled
Human Labelled
ADVANCE
0.30
0.31
0.30
AHEAD-DS
0.48
0.48
0.22
Table 2 : The μ1% values for each dataset and MLLM pair before and after relabelling by a human, rounded to two decimal places. μ1% represents the mean of the 1st percentile (bottom 1% ) of CLAP scores . μ1% ranges from 1 to −1 , a value of 1 represents perfect alignment and −1 is perfect misalignment. A higher value is better.
μc
AuditoryHuM
Zero-Shot
Captioning
ADVANCE
0.61
0.25
0.42
AHEAD-DS
0.55
0.40
0.46
TAU 2019
0.49
0.33
0.19
Table 3 : μc represents mean alignment of labels produced by AuditoryHum vs zero-shot using provided labels vs captioning. A higher value is better.
μ1%
Qwen 2.5 Omni 3B + Human-CLAP
Qwen 2.5 Omni 3B + LAION-CLAP
ADVANCE
0.13
0.07
AHEAD-DS
0.07
0.07
TAU 2019
0.10
0.13
Human Labelled + Human-CLAP
Human Labelled + LAION-CLAP
ADVANCE
0.30
0.10
AHEAD-DS
0.22
0.04
Table 4 : The μ1% values comparing the results from Human-CLAP and LAION-CLAP rounded to 2 decimal places. A higher value indicates better alignment between label and audio content.
Instances
Mean CLAP score Minimal Cleanup
Mean CLAP score Default Cleanup
ADVANCE Long Labels
578
0.60
0.57
ADVANCE Non-English Lab.
0
AHEAD-DS Long Labels
781
0.60
0.58
AHEAD-DS Non-Eng. Lab.
4
0.14
-0.07
TAU 2019 Long Labels
2726
0.50
0.49
TAU 2019 Non-English Lab.
2
0.13
0.18
Table 5 : The mean CLAP score values comparing the results with default or minimal label cleanup.
Figure 2 : Silhouette scores comparing AuditoryHuM vs zero-shot CLAP audio embeddings vs caption sentence transformer embeddings. A higher value is better.
Provided Labels mAP
Provided Labels Accuracy
AuditoryHuM Labels mAP
AuditoryHuM Labels Accuracy
ADVANCE
0.84
0.79
0.84
0.84
AHEAD-DS
0.97
0.98
0.83
0.89
TAU 2019
0.94
0.95
0.76
0.75
Table 6 : The mean average precision (mAP) and accuracy of the dataset provided and AuditoryHuM generated labels. Higher values are better.
Figure 3 : Comparison of CLAP scores between human and MLLM labels.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Prompt
Notes
Describe the auditory scene using word pairs. Separate each pair with a comma.
Initial prompt.
Provide a short sentence to describe this set of audio samples. The frequency distribution of individual labels for this set of audio samples is provided: <distribution>…</distribution> .
Prompt to generate a descriptive composite label for the entire cluster.
Describe this auditory scene.
Captioning prompt.
Appendix
Table 7 : The prompts for generating labels.
Step
Model or Algorithm
Params | ≈ MACs
Complexity
Notes
1
Label generation (Qwen 2.5 Omni 3B)
4.7035B | 971.995G
O(H⋅WlogW)+O(N⋅P+N2⋅D)+O(P+N⋅D)
Steps include STFT, encoder, and decoder. H is the number of hops. W is the window length. N is the total token length. P number of params. D is the hidden layer dimension.
Steps include STFT, audio encoder, label clean up on L labels, text encoder on L labels, and Cosine similarity on L labels. W is the window length. Ta is the audio token length. Tt is the text token length. D is the hidden layer dimension. E is the CLAP embedding length.
3
Select best aligned label
O(L)
Find the largest value in an unsorted sequence of length L .
Steps 1 to 3 are repeated A times. Where A is the number of audio samples.
Steps include STFT, audio encoder, label clean up on L labels, text encoder on L labels, and Cosine similarity on L labels. W is the window length. Ta is the audio token length. Tt is the text token length. D is the hidden layer dimension. E is the CLAP embedding length.
Step 4 is repeated up to A2 times to get alignment vectors.
Appendix
Table 8 : Breakdown of the computational cost and complexity of each step in AuditoryHuM, using multiply and accumulate computations (MACs) and Big-O.
Figure 4 : t-SNE visualisations using k∗ value for each dataset. Each unique colour represents a cluster of thematically related labels.
Dataset
Cluster Size / Total
Unique Labels
Composite Label for Cluster
ADVANCE
1582 / 5075
186
The audio samples consist mainly of ocean related sounds such as waves, thunder, and wind, with a few other elements like car, train, and aircraft sounds mixed in, along with various background noises and other natural sounds.
ADVANCE
1367 / 5075
375
The audio samples consist of a variety of sounds including thuds, frying noises, car sounds, choir voices, vibrations, horse trotting, train movement, drills, sewing machines, and other mechanical sounds, among others.
ADVANCE
983 / 5075
117
These audio samples contain various animal sounds, human activities, and natural elements, including bird vocalizations, insects, and weather sounds.
AHEADDS
4713 / 9968
198
This audio set contains a variety of sounds including clatter, child speaking, restaurant background noise, white noise, engine sounds, female conversations, dishes clinking, door slams, car passing, male laughter, crowd talk, and many other common everyday sounds, totaling 4713 samples.
AHEADDS
1668 / 9968
100
The audio set contains a diverse range of music genres and styles, including country, blues, heavy metal, opera, jazz, and more, with a focus on vocal performances and specific instruments like guitar and piano.
AHEADDS
1582 / 9968
93
The audio samples predominantly feature traffic noise, car passing, and various road-related sounds, with some additional elements like engine sounds, wind, and weather effects.
Appendix
Table 9 : Details of the 3 largest clusters for each dataset.
Component / Hyperparameter
Value / Strategy
Architecture
OpenYAMNet (Pretrained on AudioSet)
Optimiser
Adam
Loss Function
Focal Loss
Label Smoothing
0.1
Data Scheduling
Shuffled each epoch
Learning Rate Scheduler
Halved after 3 consecutive epochs without val improvement
Appendix
Table 10 : Downstream training configuration and hyperparameters.