eess.ASAug 13, 2020

Textual Echo Cancellation

Authors: Shaojin Ding, Ye Jia, Ke Hu, Quan Wang

Organizations: Google LLC, USA

Abstract

In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can largely improve speech recognition performance and user experience for intelligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS playback as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to enhancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC).

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

    Jun 20, 2026Muyang Du, Jason Roche, Junjie LaiText-To-Speech SynthesisF5-Tts

  2. Taming Long-form Text-to-Speech

    Sep 15, 2026Rongxiang Wang, Berkin Durmus, Aysegul Orhon +2Autoregressive Text-To-SpeechVoice Conversion

  3. DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

    Jul 5, 2026Junwon Moon, Seungbeom Kim, Yejin Lee +4Autoregressive Text-To-SpeechDiffusion Language Models