cs.SDMar 20, 2026

When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal

Authors: Yen-Ting Piao, Jay Chiehen Liao, Wei-Tang Chien, Toshiki Ogimoto, Shang-Tse Chen, Yun-Nung Chen, Chun-Yi Lee, Shao-Yuan Lo

Organizations: National Taiwan University, Taiwan

Abstract

While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evaluation pipeline that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability in the audio modality. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across LALMs: in-context demonstrations reliably improve format compliance but fail to improve the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from examples with audio, highlighting potential limitations in current cross-modal integration. We further probe how demonstrations are used through two complementary analyses, demonstration label shuffling and attention knockout on demonstration spans, both showing that LALMs leverage in-context examples primarily to establish the output label space and format rather than to learn a meaningful input-output correspondence.

Figures & tables

Explore similar work

CardsList
  1. When Capabilities Fail to Compose: Diagnosing the Compositionality Gap in Large Audio-Language Models

    Sep 29, 2026Chien-Feng Liu, Chih-Kai Yang, Bo-Han Feng +3Large Audio Language ModelsLarge Language Models Fail

  2. A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

    Jun 16, 2026Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3Large Audio Language ModelsAudio Understanding

  3. IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

    Jun 17, 2026Sakshi Joshi, Dhruv Subhash Rathi, Sanskar Singh +4Large Audio Language ModelsMultilingual Benchmark