eess.ASSep 3, 2026

Brain2Speech-Net: Fast and Intelligible Brain-to-Speech Synthesis Without Text Decoding

Authors: Shreeram Suresh ChandraZexin CaiYu TsaoSimon KingBerrak Sisman

Abstract

The loss of speech limits communication for individuals with paralysis. Direct neural-to-speech synthesis is challenging due to the limited availability of neural data for training speech brain-computer interfaces. Most existing systems rely on cascaded neural-to-text-to-speech pipelines, which increase inference latency and propagate errors across stages. We present Brain2Speech-Net, a single-stage neural-to-speech generation framework without intermediate text decoding. We use a differentiable phoneme bottleneck and a deep-HMM alignment mechanism to map long neural recordings into the latent space of a text-to-speech (TTS) model, enabling high-quality speech synthesis. Brain2Speech-Net is the only system in our comparison that produces intelligible speech while generating faster than real time.

Explore similar work

CardsList
  1. CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment

    Feb 23, 2026Hanwen Liu, Saierdaer Yusuyin, Hao Huang +1F5-Tts