In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, when providing interleaved speech-text demonstrations. In this work, we ask whether in-context adaptation is an inherent ability for all encoder-decoder models. We study two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures. We find that all tested models are able to perform in-context adaptation out of the box, achieving up to 30% relative improvement in the oracle experiments and up to 23% using first-pass hypotheses. Through controlled experiments on three English datasets, we show that lexical and speaker information both contribute to successful adaptation. While interleaved demonstration is effective in certain cases, collated demonstration brings consistent adaptation across the board. Our results suggest that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.
Figures & tables
Fig. 1: Two demonstration approaches for in-context adaptation of encoder-decoder models: (a) Collated demonstration for AED models (top) and for decoder-only/LLM-based models (bottom) and (b) Interleaved demonstration for LLM-based models.
Model
Architecture
ICL
Whisper [ 25 ]
AED
not intended
Canary [ 26 ]
AED
not intended
Canary-Qwen [ 27 ]
LLM-based (ASR-only)
not intended
Qwen3-ASR [ 28 ]
LLM-based (ASR-only)
not intended
Phi-4-MM [ 29 ]
LLM-based (Omni)
intended
Qwen2.5-Omni [ 30 ]
LLM-based (Omni)
intended
TABLE I: A summary of the selected ASR models. Models that are exposed to interleaved input modality during training are marked as intended for ICL.
∅
Random
Gold text
Whisper
7.8
7.3
0.6
Canary
6.8
6.7
2.2
Canary-Qwen
6.7
6.6
0.2
Qwen3-ASR
6.0
5.9
0.1
Phi-4-MM
6.7
6.7
0.3
Qwen2.5-Omni
5.5
5.6
0.1
TABLE II: WERs of in-context adaptation in corner cases on L2-Arctic. The symbol ∅ denotes the baseline without demonstrations. The Random setting includes a single random demonstration from a different speaker, and the Gold text setting includes the target speech and its ground truth transcript as the demonstration.
Fig. 2: WERs when separating the demonstration and the target speech with silence in the Gold text setting on L2-Arctic. The results show the potential limit of a model’s ability to handle long-form audio.
Fig. 4: The absolute WER on L2-Arctic under the Random and Lexical settings as we increase the number of demonstrations from 1 to 4.
Fig. 5: Relative WER reduction of in-context adaptation against no adaptation on LibriSpeech test-other (top) and AMI test split from Open ASR Leaderboard (bottom). We use relative reduction because the WERs span a larger dynamic range across models compared to L2-Arctic.
Model
∅
Random
Lexical
first pass
oracle
Collated demonstration
Whisper
7.8
6.7
6.0
5.1
Canary
6.8
6.1
5.5
5.0
Canary-Qwen
6.7
6.2
5.5
5.0
Qwen3-ASR
6.0
5.6
5.0
4.5
TABLE III: WERs of second-pass in-context adaptation on L2-Arctic in the Random and Lexical settings using 4 demonstrations from the same speakers. The first pass column uses the first-pass transcription for selecting demonstrations, whereas the oracle column selects demonstrations with the ground truth transcripts.
Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR) performance on general-domain speech. However, adapting them to domain-shifted speech, such as child or dialectal speech, remains challenging under limited target-domain data. Given the dominant role of the LLM in Speech-LLMs, with cross-entropy loss applied only at the LLM output, the speech encoder may receive insufficient adaptation to new acoustic conditions. In this paper, we propose Encoder Awakening via Adapters (EAVA), a simple yet effective domain-adaptive fine-tuning method for Speech-LLM-based ASR. First, lightweight adapters are inserted into each encoder layer and trained exclusively, enabling target-domain acoustic knowledge to be incorporated into the encoder while preserving its pre-trained knowledge. Second, the full model is jointly fine-tuned on the target domain with LoRA applied to the LLM. Experiments on three domain-shifted ASR datasets, covering child and dialectal speech, show that EAVA consistently outperforms vanilla fine-tuning and other baselines, achieving new state-of-the-art performance.
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale language modeling. However, the cost of autoregressive decoding scales with decoder size, creating a fundamental trade-off between recognition quality and serving latency. We argue this trade-off is not inherent: unlike open-ended text generation, ASR outputs are strongly anchored to the input speech signal, providing a natural inductive bias toward high-parallelism decoding. Building on this, we introduce ParaASR, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step. Starting from a publicly available audio-language foundation, the model first establishes a robust autoregressive recognizer and then aligns five future-token branches through a staged optimization recipe. At inference, it proposes a six-token continuation per step and admits only the verified prefix into the transcript, preserving the safety of standard autoregressive decoding. The average accepted length reaches 5.0 out of 6 proposed tokens, confirming that the deterministic structure of speech makes ASR an especially natural setting for multi-token decoding. ParaASR further retains a native 32K-context window and transcribes up to 30 minutes of audio in a single pass. Across diverse benchmarks, it attains average error rates of 2.97%, 3.68%, and 3.70% on Chinese, English, and long-form evaluations, respectively, while reaching a real-time factor (RTF) as low as 0.0053. These results show that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal and guarded by autoregressive verification.
Speech-aware large language models often generalize poorly to out-of-domain settings. We propose SALSA (Speech-Aware LLM Adaptation via Learned Steering Activations), a lightweight adaptation method that learns layer-wise steering vectors. Unlike commonly used steering approaches that rely on contrastive activation differences, SALSA directly optimizes steering vectors using a supervised objective. Across children's speech, multilingual speech, and Mandarin-English code-switching benchmarks, SALSA substantially improves performance over zero-shot inference and speech in-context learning baselines, achieving up to 46.8% relative improvements over zero-shot. Analysis further demonstrates that steering the encoder, particularly the later layers, is more effective than steering the LLM backbone. These findings suggest that steering improves downstream ASR performance by adapting higher-level acoustic and phonetic representations to better align with the pretrained language model representation space, rather than by modifying the decoder itself.