cs.SDSep 3, 2026

Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

Authors: Yoto FujitaSimon LeglaiveLaurent Girin

Organizations: CentraleSup´elec, IETR (UMR CNRS 6164), France · Univ. Grenoble Alpes, CNRS, Grenoble-INP, GIPSA-lab, France

Abstract

Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate a set of different decoding policies, ceteris paribus, that is, using the same DNN (a Conformer model), the same NAC (the DAC codec) and the same training setup. The results show that MARSE enables a flexible trade-off between SE performance and computational cost. Audio examples and code are available online.

Explore similar work

CardsList
  1. DriftSE: Speech Enhancement with Generative Drifting

    Sep 14, 2026Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn +2Speech EnhancementAcoustic Representation