cs.SDJan 19, 2026

Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

Authors: Seymanur Akti, Alexander Waibel

Organizations: KIT Campus Transfer GmbH (KCT) · Karlsruhe Institute of Technology (KIT) · Carnegie Mellon University (CMU)

Abstract

The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard-like speech in a zero-shot manner without requiring Lombard-specific training data. Our approach extends F5-TTS with a learned style embedding representation and analyzes the resulting latent space using principal component analysis (PCA) to identify directions associated with Lombard-related attributes. By manipulating these directions, we obtain interpretable control over vocal effort and articulation and generate speech at different Lombard levels. Experimental results show that the proposed method preserves speaker identity and naturalness, improves intelligibility under noisy conditions, and generalizes to previously unseen speakers. These findings demonstrate that style-embedding manipulation provides an effective and scalable framework for controllable zero-shot Lombard speech synthesis.

Figures & tables

Explore similar work

CardsList
  1. Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

    Jun 22, 2026Seymanur Akti, Alexander WaibelVocalizationsUtterances

  2. AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

    Jul 14, 2026Haowei Lou, Junda Wu, Chengkai Huang +4StyleSpeaker

  3. Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

    Jul 29, 2026Carlos Muñoz-Romero, Jose A. Gonzalez-LopezFlow-Matching Text-To-SpeechZero-Shot