cs.SDAug 12, 2026

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

Authors: Rong ChaoSung-Feng HuangMoreno La QuatraSabato Marco SiniscalchiWen-Huang ChengSzu-Wei FuYu Tsao

Organizations: 1Academia Sinica, Taiwan · 2National Taiwan University, Taiwan · 5NVIDIA · 3Kore University of Enna, Italy · University of Palermo, Italy

Abstract

We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size recurrent state per layer, enabling memory- and bandwidth-efficient long-form inference. We further introduce a progressive knowledge distillation (KD) strategy that compresses an 8-layer teacher into a shallow 1-layer student by jointly distilling complex spectral outputs and intermediate representations. On Voicebank-DEMAND, the 8-layer RT-SEMamba achieves 3.32 PESQ with a 25 ms algorithmic latency constraint, and the distilled 1-layer student improves over a naive 1-layer baseline from 3.06 to 3.18 PESQ while preserving the same steady-state RTF, delivering a 2.64x speedup over the teacher. These results demonstrate that state-space models with progressive KD provide a competitive quality-latency trade-off for real-time SE.

Explore similar work

CardsList
  1. SE-MSB: End-to-End Unpaired Speech Enhancement using Mamba Schrödinger Bridges

    Sep 22, 2026Andreas Bagge, Andreas Nymand, Michael Riis Andersen +1Speech Enhancement