cs.SDOct 6, 2026

Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

Authors: Seymanur Akti, Alexander Waibel

Organizations: Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany · KIT Campus Transfer (KCT), Karlsruhe, Germany · Carnegie Mellon University (CMU), Pittsburgh, USA

Abstract

Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.

Figures & tables

Explore similar work

CardsList
  1. Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

    Jun 22, 2026Seymanur Akti, Alexander WaibelTTS SynthesisSpeech Prosody

  2. Steerspeech: Activation Steering For Emotion Control In Generated Speech

    Oct 7, 2026Afsara Benazir, Darius Pétermann, Felix Xiaozhu Lin +1Controllable Speech GenerationEmotional Speech Synthesis

  3. Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

    Jun 8, 2026Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1TTS SynthesisLanguage Model Steering