cs.CLAug 27, 2026

Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation

Authors: Zhen WangTianRui WuRongQi HanHao WuWei LiangWei Xu

Organizations: Shanghai Qi Zhi Institute, Shanghai, China · Megatronix (Beijing) Technology Co., Ltd. · Tsinghua University, Beijing, China

Abstract

Synthetic speech can provide additional supervision for automatic speech recognition (ASR), but constructing useful synthetic training data requires choosing both what to synthesize and how to synthesize it. We present a phoneme-guided text-to-speech (TTS) augmentation pipeline for ASR that connects multilingual speech generation with candidate-text selection and reference-speech quality control. Within this pipeline, we propose phoneme-frequency-guided selection (PFGS), which uses phoneme frequencies from real ASR training transcripts to prioritize candidate texts containing common phonetic content. Experiments with separate monolingual ASR systems cover four languages and 13 test sets. With random text selection, the pipeline improves recognition on 11 test sets at one or more synthesis ratios. PFGS further outperforms random selection on nine test sets, with relative word error rate (WER) reductions of up to 19.3%. An ablation with fixed target texts and synthesis counts further shows the benefit of reference-speech filtering. These results support using real-data phoneme statistics to guide the construction of effective synthetic supervision for ASR.

Explore similar work

CardsList
  1. How to Leverage Synthetic Speech for LLM-Based ASR Systems?

    Jun 27, 2026Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso +9