Domain-Incremental Learning for Generative Speech Enhancement
Authors: Manjunath Mulimani, Annamaria Mesaros, Minje Kim, Jesper Rindom Jensen
Organizations: Aalborg University, Department of Electronic Systems, Denmark · Tampere University, Signal Processing Research Centre, Finland · University of Illinois Urbana-Champaign, Siebel School of Computing and Data Science, USA
We propose a domain-incremental learning framework for generative speech enhancement (SE) that learns from a sequence of datasets or domains recorded under diverse acoustic conditions. Fine-tuning a pretrained model on continuously evolving domains leads to catastrophic forgetting of previously acquired knowledge, while zero-shot generalization often fails to adequately adapt to unseen domains. To address these challenges, we first develop a novel language model-based generative SE model that we then use as a pretrained backbone and incrementally adapt it to acoustically mismatched domains using lightweight domain-specific Low-Rank Adaptation. The proposed framework enables the model to acquire enhancement capabilities for new domains while preserving performance on previously learned domains. Evaluated on four heterogeneous speech datasets, our approach effectively adapts to new domains without forgetting previously learned domains.
Figures & tables
Figure 1: An overview of the proposed DIL for generative speech enhancement. (a) The base model M is trained on domain D0 . (b) LoRA is added to M for each incremental domain Dt .
Method
D1 DNS
D2 EARS
D3 LibriTTS-R
D4 VoiceBank
GRL
0.89
0.82
0.84
0.83
FT
0.85
0.78
0.74
0.73
Joint FT
0.85
0.80
0.82
0.83
DIL-GenSE
0.90
0.87
0.88
0.88
Table 1: Average SBSt ( ↑ ) across the current domain Dt and all previously seen domains D1,…,t−1 under the domain-aware setup.
Figure 2: Comparison of the proposed DIL-GenSE with zero-shot generalization and fine-tuning. (a) SBS ( ↑ ) of the generalization and the DIL-GenSE method in the current domain Dt . (b) SBS ( ↑ ) at the Dt and average forgetting FRt over the previously encountered domains D1,…,t−1 learned for the FT.
D1 DNS
D2 EARS
D3 LibriTTS-R
D4 VoiceBank
0.90
0.79
0.80
0.76
Table 2: Average SBSt ( ↑ ) across the current domain Dt and all previously seen domains D1,…,t−1 under the domain-agnostic setup.
DNSMOS ↑
Method
SIG
BAK
OVL
SBS ↑
GenSE [ 31 ]
2.60
3.22
2.14
0.52
DIL-GenSE Agnostic
3.39
3.99
3.16
0.74
DIL-GenSE Aware
3.56
4.05
3.31
0.84
Table 3: Comparison of proposed DIL-GenSE with GenSE [ 31 ] on the EARS domain.
Institute of Electronic Music and Acoustics, University of Music and Performing Arts, Graz, Austria · Department of Informatics, King’s College London, London, United Kingdom