Textual Echo Cancellation
Organizations: Google LLC, USA
Abstract
In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can largely improve speech recognition performance and user experience for intelligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS playback as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to enhancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC).
Figures & tables
| Spectral analysis | frame length: 50 ms; frame shift: 12.5 ms; | |
| 128 Mel-filterbanks | ||
| Audio encoder | Conv layers 2 | 32 3 3 kernel with 2 2 stride; |
| ReLU; batch norm | ||
| Bi-CLSTM 1 | 256 units per direction; | |
| 1 3 kernel with 1 1 stride | ||
| Bi-LSTM 3 | 256 units per direction | |
| Condition | Synthetic subset | User’s speech | Interfering speech |
|---|---|---|---|
| Single interfering voice | training | LibriTTS training | LJ Speech training |
| test-clean | LibriTTS test-clean | LJ Speech test | |
| test-other | LibriTTS test-other | LJ Speech test | |
| Multiple interfering voices | training | LibriTTS training | VCTK training |
| test-clean | LibriTTS test-clean | VCTK test | |
| test-other | LibriTTS test-other | VCTK test |
| Condition | Method | WER (%) | MCD | MOS | Side input (KB) | GFLOPS | |||
| test- | test- | test- | test- | test-clean | test-other | ||||
| clean | other | clean | other | ||||||
| Ground-truth LibriTTS | - | 2.30 | 4.50 | 0.00 | 0.00 | 4.43 0.04 | 3.82 0.06 | - | - |
| Single interfering voice | Microphone signal | 89.9 | 120.5 | 18.83 | 21.44 | - | - | - | - |
| AEC-NLMS | 48.6 | 60.1 | 12.26 | 12.57 | 1.95 0.10 | 1.28 0.09 | 310 | 0 | |
| Vanilla-Seq2seq | 25.4 | 54.0 | 7.85 | 8.84 | 1.99 0.06 | 1.47 0.05 | 0 | 6.32 | |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Original Internal Implementation (Main Paper) | Open-Source Reproduction ( wq2012/tec ) |
|---|---|---|
| Modeling Framework | Internal Google Lingvo / TensorFlow trained on Cloud TPU slices (global batch size 32, 50,000 steps). | Standalone open-source lingvo ( ) and tensorflow ( ) supporting CPU, GPU, and TFLite execution. |
| Acoustic Frontend | 24 kHz sample rate, 50 ms frame length (1,200 samples), 12.5 ms frame shift (300 samples), 128 log-Mel filterbanks. | WaveformProcessor ( tec/waveform_processor.py ): 24 kHz, 50 ms Hann window, 12.5 ms hop, 2,048-point FFT, 128 log-Mel bins (20–12,000 Hz, natural-log compression with magnitude floor ). |
| Audio Encoder | 2 strided Conv2D layers ( stride, 32 channels), 1 Bi-CLSTM ( kernel, 256 units/dir), 3 Bi-LSTM layers (256 units/dir). | SpeechEncoderV1 ( tec/encoder_speech.py ): Subclasses lingvo.tasks.asr.encoder.AsrEncoder with identical Conv2D pyramid ( time reduction), configurable Bi-CLSTM, and 3 Bi-LSTM layers (512-dim output). |
| Text Tokenizer & Encoder | Internal TTS text normalization and pre-trained 512-dim phoneme embedding table, followed by 3 Conv1D layers ( , 512 channels, BatchNorm, ReLU) and 1 Bi-LSTM (256 units/dir). | CharTokenizer ( tec/tokenizer.py ) + TtsEncoderV2 ( tec/encoder_text.py ): Self-contained 96-symbol printable ASCII character vocabulary ( <pad>=0 , </s>=1 , <unk>=2 + 93 ASCII characters) with learned 512-dim embedding, 3 Conv1D layers ( , 512 channels, BatchNorm, ReLU, dropout 0.5), and 1 Bi-LSTM (256 units/dir, zoneout 0.1). |
| Multi-Source Attention | Independent 5-mixture Gaussian Mixture Model (GMM) monotonic attention [ 13 ] per encoder stream with context summation . | MultiSourceAttentionStep ( tec/multi_source_attention_steps.py ): Wraps GmmMonotonicAttention (5 mixtures, 128-dim hidden state) or AdditiveAttention for each source stream ( source_0 , source_1 ) and sums context vectors. |
| Spectrogram Decoder | Tacotron-2 [ 43 ] Pre-Net ( ReLU, dropout 0.5), 2-layer LSTM (256 units), linear Mel and stop-token heads, and 5-layer Conv1D Post-Net ( , 512 channels, BatchNorm, ). | MultiSourceFbeDecoderV1 and FbeDecoderV1 ( tec/decoder.py ) + PostEditConvNet ( tec/layers.py ): Identical layer dimensions with frame reduction factor , masked loss, and step-aligned spectral residual skip connections for fast convergence. |
| File Path | Key Classes / Functions | Paper Reference & Functionality |
|---|---|---|
| tec/waveform_processor.py | WaveformProcessor , compute_fft_size | Section 2 .1 & Table 1 : 24 kHz STFT, 128-bin log-Mel filterbank extraction ( WaveformsToSpectrograms ), and Griffin-Lim / phase-preserving waveform synthesis ( SpectrogramsToWaveforms ). |
| tec/tokenizer.py | CharTokenizer | Section 2 .3: Maps TTS source text strings to 96-symbol ASCII character token IDs and padding masks ( text_to_ids , ids_to_text , batch_encode ). |
| tec/encoder_speech.py | SpeechEncoderV1 | Section 2 .2, Eq. (1) & Table 1 : Encodes 128-bin log-Mel spectrograms into 512-dim hidden sequences at reduced frame rate. |
| tec/encoder_text.py | TtsEncoderV2 | Section 2 .3, Eq. (2) & Table 1 : Encodes TTS character IDs into 512-dim hidden representations via 512-dim embeddings, 3 Conv1D layers ( ), and 1 Bi-LSTM. |
| tec/multi_source_attention_steps.py | MultiSourceAttentionStep | Section 2.5 , Eq. (8)–(10): Computes per-source GMM monotonic attention contexts and and sums them into . |
| tec/layers.py | PostEditConvNet | Section 2 .4, Eq. (7) & Table 1 : 5-layer Conv1D residual Post-Net ( kernel, 512 channels, BatchNorm, on first 4 layers, linear final projection). |
| # | Clean User Query (LibriTTS test-clean ) | Interfering TTS Source Text (LJ Speech test ) | ASR on Original Mixture (Before TEC) | ASR on TEC Enhanced Audio (After TEC) |
|---|---|---|---|---|
| 1 | “I can’t see you at all, anywhere.” | “a table showing the figures for the year ending Michaelmas eighteen oh two.” | “I can’t see you at all anywhere for the year ending, Michael Miss 1802. ” ( 100.0% WER ) | “ I can’t see you at all anywhere. ” ( 0.0% WER ) |
| 2 | “Because the thing had been such a scare?” | “and to approve or disapprove the public policy written into these laws.” | “Because the thing had been such a scare. The public policy written into these laws. ” ( 87.5% WER ) | “ because the thing had been such a scare. ” ( 0.0% WER ) |
| 3 | “I must know about you.” | “was living in the city while the walls were still standing, though in a ruinous condition.” | “I must know about you. The yellow walls were still standing, though in a row it is conditioned. ” ( 260.0% WER ) | “ I must know about you. ” ( 0.0% WER ) |
| 4 | “I should much prefer that you called in the aid of the police.” | “a subsequent bullet, which was lethal, shattered the right side of his skull.” | “I should much prefer that you called in the aid of a police. Shattered the right side of his skull. ” ( 61.5% WER ) | “ I should much prefer that you call them the aid of the police. ” ( 15.4% WER ) |
| Condition | Split | Source Manifests (User + Interfering) | User Pool | Interfering Pool | Prepared TFRecord Shard |
|---|---|---|---|---|---|
| Single interfering voice | train | LibriTTS train-clean-100 + LJSpeech train (90%) | 33,236 | 11,790 | 1,000 mixtures (1.3 GB) |
| test-clean | LibriTTS test-clean + LJSpeech test (10%) | 4,837 | 1,310 | 100 mixtures (131 MB) | |
| test-other | LibriTTS test-other + LJSpeech test (10%) | 5,120 | 1,310 | 100 mixtures (126 MB) | |
| Multiple interfering voices | train | LibriTTS train-clean-100 + VCTK train (90%/spk) | 33,236 | 39,859 | 1,000 mixtures (1.1 GB) |
| test-clean | LibriTTS test-clean + VCTK test (10%/spk) | 4,837 | 4,424 | 100 mixtures (106 MB) | |
| test-other | LibriTTS test-other + VCTK test (10%/spk) | 5,120 | 4,424 | 100 mixtures (105 MB) |
| Open-Source Reproduction ( wq2012/tec + Qwen3-ASR-0.6B ) | Paper Ref. (Table 3 ) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| WER (%) | MCD (dB) | Side Input (KB) | WER (%) | MCD (dB) | ||||||
| Condition | Method | Hugging Face Repo ( wq2012/... ) | clean | other | clean | other | clean | other | clean / other | clean / other |
| Single interfering voice (LibriTTS + LJ Speech) | GroundTruth | — | 3.52 (7/199) | 6.78 (12/177) | 0.00 | 0.00 | 0.000 | 0.000 | 2.30 / 4.50 | 0.00 / 0.00 |
| MicrophoneSignal | — | 90.45 (180/199) | 114.12 (202/177) | 12.86 | 14.61 | 0.000 | 0.000 | 89.9 / 120.5 | 18.83 / 21.44 | |
| NlmsAec (AEC-NLMS) | — | 88.44 (176/199) | 107.34 (190/177) | 12.80 | 14.48 | 243.465 | 209.085 | 48.6 / 60.1 | 12.26 / 12.57 | |
| Vanilla-Seq2seq | vanilla_seq2seq_single_interfering | 45.23 (90/199) | 91.53 (162/177) | 9.58 | 11.34 | 0.000 | 0.000 | 25.4 / 54.0 | 7.85 / 8.84 | |