Organizations: National Taiwan University · ASUS Open Cloud Infrastructure Software Center · NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)
Large audio-language models (LALMs) perform strongly on individual audio tasks, but whether these capabilities can be reliably composed remains underexplored. We conduct a controlled diagnostic study of capability composition in LALMs, requiring models to integrate audio-attribute recognition, cue-conditioned segment selection, and downstream ASR or question answering. We construct two-utterance inputs with distinct acoustic cues to evaluate composition across environmental sound, gender, and emotion cues, with ASR, Math QA, and Factual QA as downstream tasks. Across four open-source LALMs, compositional QA accuracy decreases in 39 of 40 model-task-cue settings, by an average of 26.7 percentage points. ASR exhibits a similarly consistent degradation, with WER increasing in 39 of 40 settings by an average of 28.5 percentage points, while the magnitude of degradation varies across models, cue types, and cue salience. We further probe these failures through output format, positional preference, and chain-of-thought (CoT) analyses. Our study reveals a systematic gap between possessing individual audio capabilities and reliably composing them.
Figures & tables
Figure 1: Illustration of our diagnostic setting. Despite succeeding at background sound recognition, atomic math QA, and segment selection, the model fails when these capabilities must be composed to answer the cue-matched question.
Task
Input
Output
Atomic recognition
X
(a1,a2)
Segment selection
(X,ci∗)
i∗
Atomic downstream
(X,T)
(y1,y2)
Compositional
(X,ci∗,T)
yi∗
Table 1: Task decomposition. T denotes the downstream task and i∗ the target segment index.
Math QA
Model
Task
E0
E10
E20
Gen.
Emo.
Qwen2.5-Omni
Recognition
58.33
50.33
41.75
95.42
35.50
Selection
87.33
84.17
75.50
76.67
72.67
MiniCPM-o
Recognition
50.25
44.50
37.33
99.42
45.50
Selection
67.17
63.67
58.50
94.83
68.33
MiMo-Audio
Recognition
42.83
38.42
30.75
83.17
40.83
Table 2: Atomic recognition and segment selection accuracy (%) on the math and factual QA subsets. E0/E10/E20 denote environmental-sound conditions at 0/10/20 dB SNR; Gen. and Emo. denote gender and emotion.
Math QA
ASR ( ↓ )
Math QA ( ↑ )
Model
Stage
E0
E10
E20
Gen.
Emo.
E0
E10
E20
Gen.
Emo.
Qwen2.5-Omni
Atomic
9.26/9.31
4.29/4.35
4.06/4.10
3.86/3.86
70.92/71.33
82.83/82.83
83.42/83.50
82.08/82.42
Comp.
29.05/29.05
31.81/31.81
42.59/42.59
18.62/18.67
48.74/48.79
34.33/42.83
34.83/47.33
28.67/40.50
42.00/47.67
29.33/40.17
MiniCPM-o
Atomic
30.07/15.30
21.77/9.82
19.72/8.92
18.95/9.19
75.33/75.33
86.08/86.17
86.83/86.92
88.50/88.67
Comp.
67.68/51.01
73.53/52.67
78.73/57.27
58.43/20.16
65.69/45.44
11.00/58.17
13.83/60.17
14.83/58.50
26.33/93.33
15.17/55.33
Table 3: Atomic and compositional performance. ASR is evaluated by WER (%, lower is better) and QA by accuracy (%); each entry reports tag-based / LLM-based parsing results. E0/E10/E20 denote environmental-sound conditions at 0/10/20 dB SNR, and Gen./Emo. denote gender/emotion conditions. Atomic Gen./Emo. values are shared because both attributes use the same audio.
Model
Stage
Fmt.
ΔASR
Δmath
Δqa
Qwen2.5-Omni
Atomic
99.98
-0.83
+0.15
+0.02
Comp.
100.00
-1.66
+9.87
+0.36
MiniCPM-o
Atomic
92.71
+7.20
+0.08
+0.02
Comp.
54.91
+12.67
+48.87
+8.57
MiMo-Audio
Atomic
88.88
+7.59
+4.81
+2.25
Comp.
83.43
+1.08
+3.73
+3.73
Table 4: Output-format compliance and parsing sensitivity. Fmt. denotes format-following rate (%). For QA, Δ is the accuracy difference between LLM-based and tag-based parsing; for ASR, it is the corresponding reduction in WER. Positive values indicate better scores with LLM-based parsing.
Model
E0
E10
E20
Gender
Emotion
Qwen2.5-Omni
54.67
55.42
56.25
63.42
52.50
MiniCPM-o
41.00
35.34
29.08
44.34
27.09
MiMo-Audio
75.83
79.17
83.07
52.17
84.50
Kimi-Audio
43.50
38.25
34.42
47.34
41.42
Table 5: Positional preference in segment selection. Entries show the percentage of predictions selecting the first segment, averaged over math and factual QA. With balanced target positions, 50% indicates no positional preference.
Math QA
Factual QA
Model
ΔTag
ΔLLM
ΔFmt
ΔTag
ΔLLM
ΔFmt
Qwen2.5-Omni
-1.56
-11.17
0.00
+2.10
+1.90
-0.03
MiniCPM-o
+41.67
-2.10
+68.50
+10.10
+14.63
-12.07
MiMo-Audio
+24.06
+20.53
+9.50
+10.53
+9.54
+25.43
Kimi-Audio
-4.37
+6.10
-15.17
+3.27
+12.71
+27.97
Table 6: Effect of CoT prompting on compositional QA. Values are changes from direct prompting, averaged across cue conditions; ΔTag , ΔLLM , and ΔFmt denote changes in tag-parsed accuracy, LLM-parsed accuracy, and format-following rate, all in percentage points.