cs.CLSep 27, 2026

In-Context Adaptation of Encoder-Decoder Models in Speech Recognition

Authors: Yen Meng, Sharon Goldwater, Hao Tang

Organizations: The Centre for Speech Technology Research, University of Edinburgh, United Kingdom

Abstract

In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, when providing interleaved speech-text demonstrations. In this work, we ask whether in-context adaptation is an inherent ability for all encoder-decoder models. We study two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures. We find that all tested models are able to perform in-context adaptation out of the box, achieving up to 30% relative improvement in the oracle experiments and up to 23% using first-pass hypotheses. Through controlled experiments on three English datasets, we show that lexical and speaker information both contribute to successful adaptation. While interleaved demonstration is effective in certain cases, collated demonstration brings consistent adaptation across the board. Our results suggest that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.

Figures & tables

Explore similar work

CardsList
  1. Encoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs

    Sep 16, 2026Mohan Shi, Zilai Wang, Natarajan Balaji Shankar +3Speech EncoderEncoders

  2. SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors

    May 30, 2026Yekaterina Yegorova, Argyrios Gerogiannis, Haolong Zheng +3Speech Language ModelsLinear Activation Steering