eess.ASMar 9, 2026

Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios

Authors: Pol Buitrago, Javier Hernando

Organizations: Barcelona Supercomputing Center (BSC), Spain · Universitat Polit`ecnica de Catalunya (UPC), Spain

Abstract

Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training. Synthetic visual data have been shown to be an effective augmentation strategy for addressing AV data scarcity. However, a more challenging scenario arises for languages such as Catalan, where no real audiovisual data are available for training. In this study, we investigate whether AVSR can be bootstrapped in such a zero-AV-resource setting, using synthetic visual data as the sole source of visual supervision. We synthesize over 700 hours of talking-head video and fine-tune a pre-trained AV-HuBERT model. On a manually annotated Catalan benchmark, our model achieves near state-of-the-art (SOTA) performance with much fewer parameters and training data than SOTA ASR systems such as Whisper-large-v3, outperforms an identically trained audio-only baseline, and preserves multimodal advantages under acoustic degradation. Scalable synthetic video thus offers a viable substitute for real recordings in zero-AV-resource AVSR.

Figures & tables

Explore similar work

CardsList
  1. M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

    Jun 4, 2026Fei Su, Cancan Li, Ming Li +1Audio-Visual ReasoningMulti-View

  2. Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

    Sep 9, 2026Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin +1Automatic Speech Recognition EvaluationConversational Datasets

  3. LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition

    Apr 30, 2026Doyeop Kwak, Jeongsoo Choi, Suyeon Lee +1Automatic Speech Recognition EvaluationAudio-Visual Reasoning