Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments
Authors: Seymanur Akti, Alexander Waibel
Organizations: Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany · KIT Campus Transfer (KCT), Karlsruhe, Germany · Carnegie Mellon University (CMU), Pittsburgh, USA
Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.
Figures & tables
Condition
WER ↓
ST ↑
VSA ↑
PR ↓
SSIM ↑
SLE [ 20 ]
Baseline
4.86
-19.89
5.51
15.97
0.86
Half scaling
0.70
-17.81
7.71
12.36
0.86
Full scaling
1.09
-14.70
7.30
10.63
0.84
Ours
Baseline
1.89
-18.87
4.18
15.14
0.90
Table 1: Performance on the Expresso dataset. Highlighted rows indicate proposed configurations.
Test Set
Condition
WER ↓
ST ↑
VSA ↑
PR ↓
SSIM ↑
VCTK
Baseline
0.51
-21.19
1.31
16.88
0.96
Steered
0.29
-17.41
1.93
14.93
0.95
German
Baseline
0.77
-18.18
9.05
14.22
0.96
Steered
1.26
-13.04
4.00
11.59
0.95
Spanish
Baseline
1.76
-22.91
5.09
16.88
0.98
Steered
1.80
-18.29
6.76
10.93
0.95
Table 2: Multilingual and multispeaker evaluation across VCTK, German, Spanish, and Japanese test sets.
Figure 1: WER results in background noise with different SNR levels.
Steering Mode
Mean TTFA (s)
Mean RTF
No Steering
0.469
1.078
Dynamic Steering
0.474
1.088
Constant Steering
0.475
1.086
Table 3: Streaming performance across evaluation modes.
Steering Mode
Before
During
After
No Steering
-20.84
-21.06
-20.96
Dynamic Steering
-20.51
-16.94
-20.68
Constant Steering
-17.75
-17.26
-17.65
Table 4: Segmented spectral tilt (dB) across steering modes.
Figure 2: Features of the dynamically steered speech. Highlighted part corresponds to active steering window.