When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal
Organizations: National Taiwan University, Taiwan
Abstract
While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evaluation pipeline that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability in the audio modality. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across LALMs: in-context demonstrations reliably improve format compliance but fail to improve the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from examples with audio, highlighting potential limitations in current cross-modal integration. We further probe how demonstrations are used through two complementary analyses, demonstration label shuffling and attention knockout on demonstration spans, both showing that LALMs leverage in-context examples primarily to establish the output label space and format rather than to learn a meaningful input-output correspondence.
Figures & tables
| Model | Task Group A: ASR/SER/GR/MMAU | Task Group B: ASR/SER/GR | ||||||||||
| Zero-shot | One-shot | Few-shot | Zero-shot | One-shot | Few-shot | |||||||
| Stage 1 | Stage 1 | Stage 2 | Stage 1 | Stage 2 | Stage 1 | Stage 1 | Stage 2 | Stage 3 | Stage 1 | Stage 2 | Stage 3 | |
| Qwen2-Audio | 40.08 3.05 | 55.67 3.09 | 22.57 2.60 | 57.06 3.08 | 40.02 3.01 | 41.61 3.59 | 56.17 3.61 | 13.31 2.48 | 27.60 3.26 | 57.63 3.60 | 32.11 3.32 | 44.40 3.57 |
| DeSTA2.5-Audio | 95.45 1.31 | 98.48 0.78 | 33.81 2.94 | 98.94 0.66 | 71.38 2.62 | 95.70 1.50 | 98.75 0.85 | 32.87 3.42 | 34.81 3.47 | 98.79 0.83 | 69.38 3.15 | 67.11 3.19 |
| BLSP-Emo | 66.80 2.93 | 97.98 0.90 | 72.57 2.78 | 97.48 0.99 | 88.35 1.93 | 70.60 3.32 | 98.20 1.00 | 71.01 3.30 | 60.75 3.56 | 97.99 1.05 | 90.60 1.97 | 88.12 2.13 |
| Qwen2.5-Omni | 82.79 2.35 | 91.09 1.78 | 58.50 3.07 | 94.67 1.40 | 82.19 2.20 | 80.17 2.91 | 89.32 2.26 | 54.23 3.63 | 67.96 3.40 | 93.57 1.78 | 79.66 2.72 | 88.59 2.13 |