While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evaluation pipeline that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability in the audio modality. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across LALMs: in-context demonstrations reliably improve format compliance but fail to improve the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from examples with audio, highlighting potential limitations in current cross-modal integration. We further probe how demonstrations are used through two complementary analyses, demonstration label shuffling and attention knockout on demonstration spans, both showing that LALMs leverage in-context examples primarily to establish the output label space and format rather than to learn a meaningful input-output correspondence.
Figures & tables
Fig. 1: Our three evaluation stages: Textual guidance progressively decreases from Stage 1 (Explicit Constraint) to Stage 3 (Audio Only), with removed elements grayed out.
Model
Task Group A: ASR/SER/GR/MMAU
Task Group B: ASR/SER/GR
Zero-shot
One-shot
Few-shot
Zero-shot
One-shot
Few-shot
Stage 1
Stage 1
Stage 2
Stage 1
Stage 2
Stage 1
Stage 1
Stage 2
Stage 3
Stage 1
Stage 2
Stage 3
Qwen2-Audio
40.08 ± 3.05
55.67 ± 3.09
22.57 ± 2.60
57.06 ± 3.08
40.02 ± 3.01
41.61 ± 3.59
56.17 ± 3.61
13.31 ± 2.48
27.60 ± 3.26
57.63 ± 3.60
32.11 ± 3.32
44.40 ± 3.57
DeSTA2.5-Audio
95.45 ± 1.31
98.48 ± 0.78
33.81 ± 2.94
98.94 ± 0.66
71.38 ± 2.62
95.70 ± 1.50
98.75 ± 0.85
32.87 ± 3.42
34.81 ± 3.47
98.79 ± 0.83
69.38 ± 3.15
67.11 ± 3.19
BLSP-Emo
66.80 ± 2.93
97.98 ± 0.90
72.57 ± 2.78
97.48 ± 0.99
88.35 ± 1.93
70.60 ± 3.32
98.20 ± 1.00
71.01 ± 3.30
60.75 ± 3.56
97.99 ± 1.05
90.60 ± 1.97
88.12 ± 2.13
Qwen2.5-Omni
82.79 ± 2.35
91.09 ± 1.78
58.50 ± 3.07
94.67 ± 1.40
82.19 ± 2.20
80.17 ± 2.91
89.32 ± 2.26
54.23 ± 3.63
67.96 ± 3.40
93.57 ± 1.78
79.66 ± 2.72
88.59 ± 2.13
TABLE I: Format Compliance Rate (FCR, %) reported with 95% confidence intervals. “Few-shot” is averaged over k≥1 . Shaded cells indicate the zero-shot baseline within each task group, and bold marks the highest FCR within each model × task-group block.
Fig. 2: Format Compliance Rate (FCR) of the Number of In-Context Examples k across the evaluation stages for Groups A and B. This resolves the aggregated “Few-shot” columns of Table I into the full per-shot trajectory. Please note that zero-shot ( k=0 ) works for Stage 1 only.
Fig. 3: Task Performance of CEQ. Performance vs. number of in-context examples k under the three-stage evaluation pipeline for closed-ended questions. Rows correspond to tasks (ASR in WER ↓ ; SER/GR/MMAU in accuracy ↑ ), and columns correspond to Stages 1–3.
Fig. 4: Task Performance of CoT. Performance vs. number of in-context examples k under the three-stage evaluation pipeline for chain-of-thought prompting. Rows correspond to tasks (ASR in WER ↓ ; SER/GR/MMAU in accuracy ↑ ), and columns correspond to Stages 1–3.
Fig. 5: Format Compliance Rate (FCR) and Task Performance of Qwen2.5-Omni and Phi-4-Multimodal in Stage 3 (Audio-Only) with Original and Shuffled Labels. WER ↓ ; accuracy ↑ .
Fig. 6: Stage 3 task performance under attention knockout , comparing the baseline (gray) against masking attention to the demonstration audio (blue) or label (orange) span across 1–8 shots (WER ↓ ; accuracy ↑ ).
Large audio-language models (LALMs) perform strongly on individual audio tasks, but whether these capabilities can be reliably composed remains underexplored. We conduct a controlled diagnostic study of capability composition in LALMs, requiring models to integrate audio-attribute recognition, cue-conditioned segment selection, and downstream ASR or question answering. We construct two-utterance inputs with distinct acoustic cues to evaluate composition across environmental sound, gender, and emotion cues, with ASR, Math QA, and Factual QA as downstream tasks. Across four open-source LALMs, compositional QA accuracy decreases in 39 of 40 model-task-cue settings, by an average of 26.7 percentage points. ASR exhibits a similarly consistent degradation, with WER increasing in 39 of 40 settings by an average of 28.5 percentage points, while the magnitude of degradation varies across models, cue types, and cue salience. We further probe these failures through output format, positional preference, and chain-of-thought (CoT) analyses. Our study reveals a systematic gap between possessing individual audio capabilities and reliably composing them.
Chien-Feng Liu, Chih-Kai Yang, Bo-Han Feng +3
National Taiwan University · ASUS Open Cloud Infrastructure Software Center · NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)
Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability central to human auditory perception. Understanding the causes of these failures remains challenging as existing benchmarks report performance gaps without probing underlying mechanisms. To address this, we introduce a benchmark with 1,657 questions across three foundational tasks designed specifically for mechanistic analysis. Examining model outputs across varying input settings (behavioral analysis) reveals that models often under-utilize audio when textual cues are available. We also provide the first causal mechanistic analysis of temporal reasoning failures in LALMs. Comparing attention upweighting against scaling, we find that redistributing attention across audio tokens is more effective than increasing audio attention. Targeting task-relevant tokens yields further gains. These findings suggest that modality imbalance alone cannot explain failures. Attention scaling at bottleneck layers improves accuracy from 55.9% to 59.1% without fine-tuning, demonstrating a promising direction for future work.
AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.