Domain-Incremental Learning for Generative Speech Enhancement
Authors: Manjunath Mulimani, Annamaria Mesaros, Minje Kim, Jesper Rindom Jensen
Organizations: Aalborg University, Department of Electronic Systems, Denmark · Tampere University, Signal Processing Research Centre, Finland · University of Illinois Urbana-Champaign, Siebel School of Computing and Data Science, USA
We propose a domain-incremental learning framework for generative speech enhancement (SE) that learns from a sequence of datasets or domains recorded under diverse acoustic conditions. Fine-tuning a pretrained model on continuously evolving domains leads to catastrophic forgetting of previously acquired knowledge, while zero-shot generalization often fails to adequately adapt to unseen domains. To address these challenges, we first develop a novel language model-based generative SE model that we then use as a pretrained backbone and incrementally adapt it to acoustically mismatched domains using lightweight domain-specific Low-Rank Adaptation. The proposed framework enables the model to acquire enhancement capabilities for new domains while preserving performance on previously learned domains. Evaluated on four heterogeneous speech datasets, our approach effectively adapts to new domains without forgetting previously learned domains.
Figures & tables
Figure 1: An overview of the proposed DIL for generative speech enhancement. (a) The base model M is trained on domain D0 . (b) LoRA is added to M for each incremental domain Dt .
Method
D1 DNS
D2 EARS
D3 LibriTTS-R
D4 VoiceBank
GRL
0.89
0.82
0.84
0.83
FT
0.85
0.78
0.74
0.73
Joint FT
0.85
0.80
0.82
0.83
DIL-GenSE
0.90
0.87
0.88
0.88
Table 1: Average SBSt ( ↑ ) across the current domain Dt and all previously seen domains D1,…,t−1 under the domain-aware setup.
Figure 2: Comparison of the proposed DIL-GenSE with zero-shot generalization and fine-tuning. (a) SBS ( ↑ ) of the generalization and the DIL-GenSE method in the current domain Dt . (b) SBS ( ↑ ) at the Dt and average forgetting FRt over the previously encountered domains D1,…,t−1 learned for the FT.
D1 DNS
D2 EARS
D3 LibriTTS-R
D4 VoiceBank
0.90
0.79
0.80
0.76
Table 2: Average SBSt ( ↑ ) across the current domain Dt and all previously seen domains D1,…,t−1 under the domain-agnostic setup.
DNSMOS ↑
Method
SIG
BAK
OVL
SBS ↑
GenSE [ 31 ]
2.60
3.22
2.14
0.52
DIL-GenSE Agnostic
3.39
3.99
3.16
0.74
DIL-GenSE Aware
3.56
4.05
3.31
0.84
Table 3: Comparison of proposed DIL-GenSE with GenSE [ 31 ] on the EARS domain.
State-of-the-art speech enhancement models benefit from large-scale labeled datasets, whereas singing voice separation models suffer from limited available training data. To address this limitation, we formulate singing voice separation as domain adaptation from speech enhancement to singing voice separation. We investigate two fine-tuning strategies: full fine-tuning and parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA) on a discriminative and a generative model. Models with either adaptation strategy outperform the same architectures trained from scratch by 0.29-1.8 dB in Signal-to-Distortion-Ratio. Full fine-tuning yields the highest singing voice separation performance, but catastrophic forgetting degrades speech enhancement performance. LoRA fine-tuning achieves competitive singing voice separation performance while preserving the original speech enhancement capability with only 6-12% additional parameters compared to the base speech enhancement model. Furthermore, the generative model shows improved generalization to an unseen test set. The results demonstrate that adapting pretrained speech enhancement models is an effective strategy for training singing voice separation models in data-scarce scenarios.
Paul A. Bereuter, Mark D. Plumbley, Alois Sontacchi
Institute of Electronic Music and Acoustics, University of Music and Performing Arts, Graz, Austria · Department of Informatics, King’s College London, London, United Kingdom
We propose DriftSE, a novel one-step generative framework for speech enhancement formulated as a latent distribution equilibrium problem. During training, the drifting field aligns the generator's pushforward distribution with the clean speech manifold through drifting in a latent domain. During inference, the drifting process is discarded, enabling one-step generation. We establish that its enhancement quality depends fundamentally on the choice of latent representation. Semantic latents preserve phonetic structure but fail to capture physical acoustic cues, whereas acoustic latents reconstruct the physical signal but risk linguistic hallucination. Therefore, we introduce dual-latent drifting, performing parallel drifting in both semantic and acoustic latents to simultaneously preserve phonetic intelligibility and acoustic fidelity. Additionally, we demonstrate that DriftSE enables fully unpaired training by aligning latent distributions rather than exact point-wise targets. Consequently, DriftSE facilitates cross-dataset learning in the absence of paired noisy-clean samples. Moreover, DriftSE exhibits broad architectural flexibility across different generator backbones. Extensive evaluations on additive denoising and convolutive dereverberation demonstrate robust one-step enhancement across both offline and real-time causal settings. Notably, DriftSE achieves state-of-the-art word error rates across all four evaluated datasets while strictly operating at 1 NFE. Code and audio examples are available online.
Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn +2
Victoria University of Wellington, New Zealand · GN Advanced Science, Denmark · Lincoln University
We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). Starting from a pretrained SE model, RAM tilts the model's conditional distribution toward outputs with higher reward. During training, the current model generates enhanced speech on-policy, evaluates each generated endpoint with a potentially non-differentiable reward, and analytically re-noises the endpoint to construct inputs for a reward-guided regression objective. This enables post-training directly on real recordings using weak supervision, such as text transcripts, without requiring paired clean speech targets or reward gradients. We investigate word error rate (WER)-based post-training and whether recognition performance can be improved without compromising perceptual speech quality. Experiments on real CHiME-4 recordings reduce WER by 5.08 percentage points relative to pretrained FlowSE without reducing any of the reported non-intrusive speech quality metrics. A subjective listening test at the default reward scale finds no statistically significant preference between the post-trained and pretrained models.
Julius Richter, Christoph Boeddeker, Yoshiki Masuyama +4
Mitsubishi Electric Research Laboratories (MERL), USA · Mitsubishi Electric Corporation, Japan