The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard-like speech in a zero-shot manner without requiring Lombard-specific training data. Our approach extends F5-TTS with a learned style embedding representation and analyzes the resulting latent space using principal component analysis (PCA) to identify directions associated with Lombard-related attributes. By manipulating these directions, we obtain interpretable control over vocal effort and articulation and generate speech at different Lombard levels. Experimental results show that the proposed method preserves speaker identity and naturalness, improves intelligibility under noisy conditions, and generalizes to previously unseen speakers. These findings demonstrate that style-embedding manipulation provides an effective and scalable framework for controllable zero-shot Lombard speech synthesis.
Figures & tables
Fig. 1 : Training (left) and inference (right) pipelines of the proposed system. Blue blocks represent components inherited from the original F5-TTS model, while green blocks indicate modules introduced in the proposed version.
Fig. 2 : PCA analysis of style embeddings.
Condition
αE
αA
Speed
Soft
-0.5
-0.5
1.0
Normal
0.0
0.0
1.0
Loud
0.5
0.5
0.9
Very Loud
1.0
1.0
0.9
TABLE I : Lombard control parameters used for synthesis.
Prompt
Model
WER
SSIM
UTMOS
English
F5TTS-Base
2.11
95.90
3.84
F5TTS-Style
2.08
89.10
3.53
German
F5TTS-Base
7.04
96.30
3.18
F5TTS-Style
2.48
81.50
3.49
TABLE II : Comparison of baseline F5-TTS and F5-TTS Style.
Fig. 3 : Acoustic changes resulting from shifts along the PCA directions.
Fig. 4 : Evaluation of intelligibility robustness under Lombard-style speech generation and noise conditions.
Condition
No Noise
SNR=10
SNR=5
SNR=1
Proposed
3.09
3.24
3.67
6.52
w/o Articulation
3.98
4.06
4.14
6.76
w/o Vocal Effort
3.01
4.02
6.29
11.99
TABLE III : Ablation results measured by WER (%) across different SNR levels in the Very Loud condition.
Condition
Naturalness
Intelligibility
Clean
−0.75±0.55
1.48±0.38
Loud (SNR=10)
−1.29±0.46
0.92±0.34
Very Loud (SNR=5)
−1.60±0.33
0.21±0.40
Overall
−1.22±0.26
0.87±0.23
TABLE IV : CMOS scores with 95% confidence interval.
Universitat Oberta de Catalunya (UOC), Spain · Monoceros Labs, Spain · Dpt. of Signal Theory, Telematics and Communications, University of Granada, Spain