cs.CLJun 19, 2026

Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition

Authors: Raphaël BagatZhe ZhangJunichi YamagishiIrina IllinaEmmanuel Vincent

Organizations: Universit´e de Lorraine, CNRS, Inria, LORIA, F-54000 Nancy, France · National Institute of Informatics, Tokyo, Japan

Abstract

Automatic Speech Recognition (ASR) systems, despite achieving remarkable accuracy in general-purpose domains with native speech (L1), struggle in domains like Air Traffic Control (ATC) due to strong channel noise, a presence of non-native (L2) English accents, and data scarcity. We propose a synthetic data generation pipeline with acoustical properties simulations specifically designed to address this lack of real data to improve recognition accuracy in the ATC domain. Our approach leverages a combination of neural generation techniques, including Text-to-Speech, Voice Conversion, L2-to-L1 accent conversion, and a novel controllable L1-to-L2 accent conversion framework built to simulate accented speech. Our experiments with the Whisper model on the ATCO2 corpus demonstrate that fine-tuning with either synthetic data alone, or a mix of real and synthetic data, significantly improves the word error rate over out-of-the-box and real data only baselines respectively.

Explore similar work

CardsList
  1. How to Leverage Synthetic Speech for LLM-Based ASR Systems?

    Jun 27, 2026Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso +9