cs.CLOct 7, 2026

Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation

Authors: Shunsuke Mitsumori, Tomoya Mizumoto, Yusuke Fujita

Organizations: SB Intuitions · Waseda University, Tokyo, Japan

Abstract

Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.

Figures & tables

Explore similar work

CardsList
  1. DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

    Aug 8, 2026Yi Shu, Tianyu Peng, Yingzhuo Deng +5DialectsSelf-Supervised Speech Models

  2. Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models

    Jun 24, 2026Tomoya Mizumoto, Yusuke Fujita, Hao Shi +3DialectsSpeech Language Models

  3. PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

    May 31, 2026Sicheng Yang, Shulan Ruan, Shiwei Wu +4Multilingual Automatic Speech RecognitionDialects