Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments
Authors: Seymanur Akti, Alexander Waibel
Organizations: Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany · KIT Campus Transfer (KCT), Karlsruhe, Germany · Carnegie Mellon University (CMU), Pittsburgh, USA
Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.
Figures & tables
Condition
WER ↓
ST ↑
VSA ↑
PR ↓
SSIM ↑
SLE [ 20 ]
Baseline
4.86
-19.89
5.51
15.97
0.86
Half scaling
0.70
-17.81
7.71
12.36
0.86
Full scaling
1.09
-14.70
7.30
10.63
0.84
Ours
Baseline
1.89
-18.87
4.18
15.14
0.90
Table 1: Performance on the Expresso dataset. Highlighted rows indicate proposed configurations.
Test Set
Condition
WER ↓
ST ↑
VSA ↑
PR ↓
SSIM ↑
VCTK
Baseline
0.51
-21.19
1.31
16.88
0.96
Steered
0.29
-17.41
1.93
14.93
0.95
German
Baseline
0.77
-18.18
9.05
14.22
0.96
Steered
1.26
-13.04
4.00
11.59
0.95
Spanish
Baseline
1.76
-22.91
5.09
16.88
0.98
Steered
1.80
-18.29
6.76
10.93
0.95
Table 2: Multilingual and multispeaker evaluation across VCTK, German, Spanish, and Japanese test sets.
Figure 1: WER results in background noise with different SNR levels.
Steering Mode
Mean TTFA (s)
Mean RTF
No Steering
0.469
1.078
Dynamic Steering
0.474
1.088
Constant Steering
0.475
1.086
Table 3: Streaming performance across evaluation modes.
Steering Mode
Before
During
After
No Steering
-20.84
-21.06
-20.96
Dynamic Steering
-20.51
-16.94
-20.68
Constant Steering
-17.75
-17.26
-17.65
Table 4: Segmented spectral tilt (dB) across steering modes.
Figure 2: Features of the dynamically steered speech. Highlighted part corresponds to active steering window.
Humans tend to speak louder and clearer in challenging environments, such as noisy conditions or when addressing hearingimpaired listeners, which is called Lombard effect. To simulate this behavior in speech synthesis systems, we introduce a flow-matching based text-to-speech (TTS) model trained with vocal effort and articulation pseudo-labels. The proposed model achieves continuous and disentangled control of vocal effort and articulation, while also enabling word-level emphasis for clarifying specific segments of an utterance. Experimental results show that these control mechanisms effectively improve clarityrelated acoustic features. Furthermore, speech-in-noise experiments demonstrate that our model successfully simulates the intelligibility gains of human clear speech in noisy conditions.
Seymanur Akti, Alexander Waibel
Karlsruhe Institute of Technology (KIT), Germany · KIT Campus Transfer (KCT), Germany · Carnegie Mellon University (CMU), USA
Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations. For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen. To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens. At inference, a target-emotion steering direction is optimized with its respective transform and injected into the base TTS model. Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation. SteerSpeech achieves 1.08x-7.12x baseline target-emotion scores and for a representative emotion subjectively, it receives 78.1%-96.8% intensity preference and 1.43x-1.46x speaker-identity preservation at high steering strengths.
Afsara Benazir, Darius Pétermann, Felix Xiaozhu Lin +1
Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tokens share a single residual stream. We train BatchTopK sparse autoencoders on the LM backbone of CosyVoice3 and introduce a modality-aware auto-interp pipeline that labels each feature from where it fires-text-prefix context, 1-second speech clips, or both. The recovered features are interpretable, spanning phonemes, laughter, accent prompts and speaker gender. Steering through the SAE latent space shows these features are causal rather than merely descriptive: targeted interventions raise laughter probability from 0.02 to 0.79, flip perceived speaker gender, and control speech rate while preserving spoken content. SAE features thus serve both as interpretability objects and as control directions for TTS synthesis.
Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1