Organizations: National Taiwan University · ASUS Open Cloud Infrastructure Software Center · NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)
Large audio-language models (LALMs) perform strongly on individual audio tasks, but whether these capabilities can be reliably composed remains underexplored. We conduct a controlled diagnostic study of capability composition in LALMs, requiring models to integrate audio-attribute recognition, cue-conditioned segment selection, and downstream ASR or question answering. We construct two-utterance inputs with distinct acoustic cues to evaluate composition across environmental sound, gender, and emotion cues, with ASR, Math QA, and Factual QA as downstream tasks. Across four open-source LALMs, compositional QA accuracy decreases in 39 of 40 model-task-cue settings, by an average of 26.7 percentage points. ASR exhibits a similarly consistent degradation, with WER increasing in 39 of 40 settings by an average of 28.5 percentage points, while the magnitude of degradation varies across models, cue types, and cue salience. We further probe these failures through output format, positional preference, and chain-of-thought (CoT) analyses. Our study reveals a systematic gap between possessing individual audio capabilities and reliably composing them.
Figures & tables
Figure 1: Illustration of our diagnostic setting. Despite succeeding at background sound recognition, atomic math QA, and segment selection, the model fails when these capabilities must be composed to answer the cue-matched question.
Task
Input
Output
Atomic recognition
X
(a1,a2)
Segment selection
(X,ci∗)
i∗
Atomic downstream
(X,T)
(y1,y2)
Compositional
(X,ci∗,T)
yi∗
Table 1: Task decomposition. T denotes the downstream task and i∗ the target segment index.
Math QA
Model
Task
E0
E10
E20
Gen.
Emo.
Qwen2.5-Omni
Recognition
58.33
50.33
41.75
95.42
35.50
Selection
87.33
84.17
75.50
76.67
72.67
MiniCPM-o
Recognition
50.25
44.50
37.33
99.42
45.50
Selection
67.17
63.67
58.50
94.83
68.33
MiMo-Audio
Recognition
42.83
38.42
30.75
83.17
40.83
Table 2: Atomic recognition and segment selection accuracy (%) on the math and factual QA subsets. E0/E10/E20 denote environmental-sound conditions at 0/10/20 dB SNR; Gen. and Emo. denote gender and emotion.
Math QA
ASR ( ↓ )
Math QA ( ↑ )
Model
Stage
E0
E10
E20
Gen.
Emo.
E0
E10
E20
Gen.
Emo.
Qwen2.5-Omni
Atomic
9.26/9.31
4.29/4.35
4.06/4.10
3.86/3.86
70.92/71.33
82.83/82.83
83.42/83.50
82.08/82.42
Comp.
29.05/29.05
31.81/31.81
42.59/42.59
18.62/18.67
48.74/48.79
34.33/42.83
34.83/47.33
28.67/40.50
42.00/47.67
29.33/40.17
MiniCPM-o
Atomic
30.07/15.30
21.77/9.82
19.72/8.92
18.95/9.19
75.33/75.33
86.08/86.17
86.83/86.92
88.50/88.67
Comp.
67.68/51.01
73.53/52.67
78.73/57.27
58.43/20.16
65.69/45.44
11.00/58.17
13.83/60.17
14.83/58.50
26.33/93.33
15.17/55.33
Table 3: Atomic and compositional performance. ASR is evaluated by WER (%, lower is better) and QA by accuracy (%); each entry reports tag-based / LLM-based parsing results. E0/E10/E20 denote environmental-sound conditions at 0/10/20 dB SNR, and Gen./Emo. denote gender/emotion conditions. Atomic Gen./Emo. values are shared because both attributes use the same audio.
Model
Stage
Fmt.
ΔASR
Δmath
Δqa
Qwen2.5-Omni
Atomic
99.98
-0.83
+0.15
+0.02
Comp.
100.00
-1.66
+9.87
+0.36
MiniCPM-o
Atomic
92.71
+7.20
+0.08
+0.02
Comp.
54.91
+12.67
+48.87
+8.57
MiMo-Audio
Atomic
88.88
+7.59
+4.81
+2.25
Comp.
83.43
+1.08
+3.73
+3.73
Table 4: Output-format compliance and parsing sensitivity. Fmt. denotes format-following rate (%). For QA, Δ is the accuracy difference between LLM-based and tag-based parsing; for ASR, it is the corresponding reduction in WER. Positive values indicate better scores with LLM-based parsing.
Model
E0
E10
E20
Gender
Emotion
Qwen2.5-Omni
54.67
55.42
56.25
63.42
52.50
MiniCPM-o
41.00
35.34
29.08
44.34
27.09
MiMo-Audio
75.83
79.17
83.07
52.17
84.50
Kimi-Audio
43.50
38.25
34.42
47.34
41.42
Table 5: Positional preference in segment selection. Entries show the percentage of predictions selecting the first segment, averaged over math and factual QA. With balanced target positions, 50% indicates no positional preference.
Math QA
Factual QA
Model
ΔTag
ΔLLM
ΔFmt
ΔTag
ΔLLM
ΔFmt
Qwen2.5-Omni
-1.56
-11.17
0.00
+2.10
+1.90
-0.03
MiniCPM-o
+41.67
-2.10
+68.50
+10.10
+14.63
-12.07
MiMo-Audio
+24.06
+20.53
+9.50
+10.53
+9.54
+25.43
Kimi-Audio
-4.37
+6.10
-15.17
+3.27
+12.71
+27.97
Table 6: Effect of CoT prompting on compositional QA. Values are changes from direct prompting, averaged across cue conditions; ΔTag , ΔLLM , and ΔFmt denote changes in tag-parsed accuracy, LLM-parsed accuracy, and format-following rate, all in percentage points.
While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evaluation pipeline that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability in the audio modality. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across LALMs: in-context demonstrations reliably improve format compliance but fail to improve the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from examples with audio, highlighting potential limitations in current cross-modal integration. We further probe how demonstrations are used through two complementary analyses, demonstration label shuffling and attention knockout on demonstration spans, both showing that LALMs leverage in-context examples primarily to establish the output label space and format rather than to learn a meaningful input-output correspondence.
Yen-Ting Piao, Jay Chiehen Liao, Wei-Tang Chien +5
Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in substantially different results. Existing MCQA frameworks do not account for this variability and report a single accuracy number per benchmark or category. We dive into the MCQA evaluation framework and conduct a systematic study spanning three benchmarks (MMAU, MMAR and MMSU) and four models: Audio Flamingo 2, Audio Flamingo 3, Qwen2.5-Omni-7B-Instruct, and Kimi-Audio-7B-Instruct. Our findings indicate that models are sensitive not only to the ordering of choices, but also to the paraphrasing of the question and the choices. Finally, we propose a simpler evaluation protocol and metric that account for subtle variations and provide a more detailed evaluation report of LALMs within the MCQA framework.
Fernando López, Santosh Kesiraju, Jordi Luque
Scientific Research, Telef´onica Innovaci´on Digital, Spain · Brno University of Technology, Czech Republic
Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a breakfast emerge compositionally from temporally distributed sound events, requiring abstraction beyond the event-centric granularity that dominates current training and evaluation paradigms. Our benchmark evaluates a wide set of LALMs under a principled framework that tests how language-based reasoning, grounded in acoustic perception, structures sound abstractions into higher-level understanding. By systematically varying exemplar typicality and distractor similarity, our evaluation exposes \added{that current models do not reliably perform compositional inference from atomic acoustic events to higher-level human activities solely from audio.} All data, taxonomies, and evaluation scripts are publicly available on our companion website: https://alm-sounding-actions.onrender.com/
Michel Olvera, Paraskevas Stamatiadis, Changhong Wang +1