eess.ASJun 15, 2026

Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition

Authors: Chengxi DengXurong XieShujie HuMengzhe GengTianzi WangYoujun ChenHuimeng WangHaoning Xu+2 more

Organizations: The Chinese University of Hong Kong, Hong Kong SAR, China · Institute of Software, Chinese Academy of Sciences, China · National Research Council Canada, Canada

Abstract

This paper proposes a novel cross-utterance audio-textual prompts based speaker adaptation approach for elderly speech recognition. It enables zero-shot, real-time adaptation to unseen speakers. Speech and text embeddings are extracted from the current and a few preceding utterances, before being fused in a cross-modal manner to produce compact speaker prompts that are more consistent than i/x-vectors and ECAPA-TDNN features. Experiments on the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets suggest that the proposed online adaptation outperforms the speaker-independent (SI) model by statistically significant word error rate (WER) or character error rate (CER) reductions of 0.61% and 1.22% absolute (2.99% and 4.48% relative). Real-time factor (RTF) speed-up ratios of up to 9.83 times are obtained over offline batch-mode adaptation.

Explore similar work

CardsList
  1. Text-only adaptation in LLM-based ASR through text denoising

    Jan 28, 2026Andrés Carofilis, Sergio Burdisso, Esaú Villatoro-Tello +8Llm-Based Automatic Speech Recognition

  2. Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

    Apr 7, 2026Thibault Bañeras-Roux, Sergio Burdisso, Esaú Villatoro-Tello +9