Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Model weights are available for download at: https://huggingface.co/nvidia/RE-USE.
Figures & tables
Figure 1 : Motivated by the distortion–perception tradeoff theory, the proposed two-stage framework integrates a frozen regression model with a residual generative model.
Figure 2 : Histogram of VQScore for URGENT 2025 Challenge Track 1 subsets. Dashed lines indicate median scores.
Figure 3 : Learning curves of UTMOS scores on the validation set under (a) different VQScore filtering thresholds and (b) different learning targets.
Team /
Non-intrusive SE metrics
Intrusive SE metrics
Task-ind.
Task-dep.
Rank
DNSMOS
NISQA
UTMOS
PESQ
ESTOI
SDR
MCD ↓
LSD ↓
SBERT
LPS
SpkSim
CAcc
Noisy
1.84
1.69
1.56
1.37
0.61
2.53
7.92
5.51
0.75
0.62
0.63
81.29
Baseline
2.94
2.89
2.11
2.43
0.80
11.29
3.32
2.85
0.86
0.79
0.80
84.96
Rank 3
3.00
3.45
2.31
2.74
0.84
13.06
3.30
3.08
0.89
0.84
0.83
87.94
Rank 2
3.01
3.21
2.30
2.79
0.85
13.11
2.93
2.94
0.90
0.85
0.84
88.05
Rank 1
3.01
3.41
2.40
2.95
0.86
14.33
3.01
2.83
0.91
0.86
0.85
88.92
Table 1 : Non-blind test set results of the URGENT 2025 Challenge. All metrics are “higher is better”, except MCD and LSD. Rank N denotes the system ranked Nth in the challenge. Note that shaded metrics are not directly comparable across the two learning targets due to mismatches in the definition of the “clean” reference. See Tables 5 to 7 in the Appendix for the ablation studies using the original anechoic clean speech as the reference.
Team /
Non-intrusive SE metrics
Intrusive SE metrics
Task-ind.
Task-dep.
Rank
DNSMOS
NISQA
UTMOS
PESQ
ESTOI
SDR
MCD ↓
LSD ↓
SBERT
LPS
SpkSim
CAcc
48k Hz
Noisy
2.04
1.83
1.99
1.28
0.56
2.29
7.77
5.50
0.78
0.77
0.69
90.60
ClearerVoice
2.97
3.38
3.02
2.09
0.72
11.55
5.08
5.15
0.85
0.87
0.63
89.90
Proposed
3.31
4.41
3.55
2.65
0.77
12.79
3.93
3.19
0.89
0.91
0.87
92.50
44.1k Hz
Table 2 : Comparison with other open-source USE models on the subsets of the URGENT 2025 non-blind test set. All metrics are “higher is better,” except MCD and LSD.
DNSMOS
SpkSim
CAcc
Italian (it _ it)
Original
3.12
-
97.28
FLEURS-R
3.37
0.87
97.69
Proposed
3.20
0.98
97.00
Proposed (EARS)
3.27
0.97
98.09
Dutch (nl _ nl)
Table 3 : Speech enhancement results for unseen languages from the FLEURS dataset.
Language
Context audio
Train audio
CER (%)
WER (%)
SpkSim
FCD
Dutch
original
original
14.28 ± 0.98
19.60 ± 0.76
0.6064 ± 0.0080
0.2444 ± 0.0155
enhanced
enhanced
7.75 ± 0.83
13.66 ± 0.71
0.6603 ± 0.0047
0.1837 ± 0.0086
Italian
original
original
11.13 ± 0.94
19.20 ± 0.94
0.6004 ± 0.0034
0.1846 ± 0.0042
enhanced
enhanced
8.30 ± 0.52
15.98 ± 0.53
0.6006 ± 0.0032
0.1373 ± 0.0021
Table 4 : Zero-shot TTS evaluation after training data cleaning using our USE model on unseen languages. We report the 95% confidence intervals based on standard errors calculated from 10 independent runs per dataset.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Method
DNSMOS
NISQA
UTMOS
PESQ
STOI
SBERT
LPS
SpkSim
CACC
Unfiltered
3.06
3.18
2.24
2.36
0.69
0.87
0.81
0.79
87.29
Filtered
3.06
3.24
2.26
2.40
0.70
0.88
0.82
0.80
87.70
p -value
0.13
0.0007
2.36×10−7
1.77×10−11
4.23×10−31
2.01×10−8
5.24×10−8
2.21×10−6
0.46
Appendix
Table 5 : Comparison of speech enhancement performance with and without training data filtering.
Method
DNSMOS
NISQA
UTMOS
PESQ
STOI
SBERT
LPS
SpkSim
CACC
Early-reflected
3.06
3.24
2.26
2.40
0.70
0.88
0.82
0.80
87.70
Shifted anechoic
3.23
3.82
2.76
2.70
0.78
0.89
0.85
0.81
89.16
p -value
1.69×10−76
6.89×10−98
2.40×10−107
2.34×10−84
8.30×10−103
1.59×10−60
1.19×10−46
2.82×10−6
9.99×10−6
Appendix
Table 6 : Comparison between early-reflected and shifted anechoic training targets.
Method
DNSMOS
NISQA
UTMOS
PESQ
STOI
SBERT
LPS
SpkSim
CACC
Single-stage
3.23
3.82
2.76
2.70
0.78
0.89
0.85
0.81
89.16
Two-stage (+ GAN)
3.26
4.12
2.80
2.71
0.78
0.89
0.86
0.84
89.88
p -value
3.83×10−9
3.20×10−68
9.58×10−27
7.28×10−60
4.59×10−30
2.23×10−10
0.25
2.40×10−23
0.03
Appendix
Table 7 : Comparison between single-stage and two-stage training with GAN correction.
Figure 4 : An example of a room impulse response, highlighting the time shift n0 introduced by the direct path.
Type
Corpus
Condition
Sampling (kHz)
Duration (h)
Speech
LibriVox (DNS5)
Audiobook
8–48
350
LibriTTS
Audiobook
8–24
200
VCTK
Newspaper-style
48
80
WSJ
WSJ news
16
85
EARS
Studio recording
48
100
Multilingual Librispeech (de, en, es, fr)
Audiobook
8–48
450
Appendix
Table 8 : Dataset Composition for URGENT 2025 Challenge
Figure 5 : Example illustrating that GANs can focus on correcting over-smoothed regions while leaving other parts unchanged. The noisy speech is bandwidth-limited in the green box, corresponding to a less informative region.
Figure 6 : Example illustrating that GANs can focus on correcting over-smoothed regions while leaving other parts unchanged. The noisy speech contains strong noise in the green box, corresponding to a less informative region.
Figure 7 : Histogram of VQScore across different speech sources in the URGENT 2025 Challenge Track 1. The median of each data source is indicated by a dashed vertical line.
Figure 8 : Enhanced spectrogram comparison between using time-shifted anechoic clean speech and early-reflected speech as learning targets. (a) and (b) correspond to the same noisy input, and (c) and (d) correspond to another noisy input. Both samples are drawn from the blind-test set.
Figure 9 : Learning curves comparison on validation-set between pre-training with a regression loss followed by adversarial fine-tuning and our two-stage GAN correction. (a) Magnitude loss, (b) Phase loss, (c) Time loss, and (d) PESQ score.
Figure 10 : Spectrogram comparison of a Japanese utterance (9997427445140542468.wav) from the FLEURS dataset. The original speech contains some very low-level stationary noise, which is commonly found in non-curated ’clean’ training data.
Team
DNSMOS
NISQA
UTMOS
PESQ
ESTOI
SDR
MCD
LSD
SBERT
LPS
SpkSim
CAcc
Overall ranking
Bobbsun
3.01 (8)
3.41 (6)
2.40 (3)
2.95 (1)
0.86 (1)
14.33 (1)
3.01 (4)
2.83 (5)
0.91 (1)
0.86 (1)
0.85 (1)
88.92 (1)
2.516
Our Early reflected + GAN
3.04 (5)
3.53 (3)
2.30 (6)
2.78 (4)
0.84 (4)
12.25 (5)
2.97 (3)
2.75 (4)
0.90 (2)
0.84 (3)
0.85 (1)
88.13 (2)
3.166
rc
3.01 (8)
3.21 (9)
2.30 (6)
2.79 (3)
0.85 (2)
13.11 (2)
2.93 (2)
2.94 (8)
0.90 (2)
0.85 (2)
0.84 (3)
88.05 (3)
4.016
Our Early reflected
3.06 (4)
3.23 (8)
2.26 (8)
2.81 (2)
0.85 (2)
12.28 (4)
2.87 (1)
2.66 (1)
0.90 (2)
0.84 (3)
0.82 (5)
87.62 (5)
4.041
Xiaobin
3.00 (10)
3.45 (4)
2.31 (5)
2.74 (5)
0.84 (4)
13.06 (3)
3.30 (6)
3.08 (11)
0.89 (5)
0.84 (3)
0.83 (4)
87.94 (4)
5.033
subatomicseer
3.02 (6)
3.28 (7)
2.34 (4)
2.63 (7)
0.82 (6)
12.18 (6)
3.90 (12)
3.06 (10)
0.88 (7)
0.82 (6)
0.82 (5)
86.15 (7)
6.591
Appendix
Table 9 : The overall ranking leaderboard for the non-blind test set of the URGENT 2025 Challenge (considering only the early reflected learning target for consistency).
Language
Context Audio
Train Audio
CER (%)
WER (%)
SpkSim
FCD
Dutch
original
original
14.28 ± 0.98
19.60 ± 0.76
0.6064 ± 0.0080
0.2444 ± 0.0155
enhanced
original
12.93 ± 0.48
19.21 ± 0.63
0.6643 ± 0.0096
0.2282 ± 0.0135
original
enhanced
8.24 ± 0.55
14.31 ± 0.69
0.6359 ± 0.0050
0.1761 ± 0.0070
enhanced
enhanced
7.75 ± 0.83
13.66 ± 0.71
0.6603 ± 0.0047
0.1837 ± 0.0086
Italian
original
original
11.13 ± 0.94
19.20 ± 0.94
0.6004 ± 0.0034
0.1846 ± 0.0042
enhanced
original
12.18 ± 0.52
20.41 ± 0.75
0.6135 ± 0.0047
0.1321 ± 0.0046
Appendix
Table 10 : Zero-shot TTS evaluation after training data cleaning using our USE model on unseen languages. We report the 95% confidence intervals based on standard errors calculated from 10 independent runs per dataset.
Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhancement models for each scenario. In this paper, we propose a one-for-all, real-time universal speech enhancement model that provides explicit control over both algorithmic and computational latency. Algorithmic latency is flexibly adjusted via configurable look-ahead frames. To avoid learning inefficiency caused by varying padding configurations, we introduce parallel convolutional layers corresponding to different look-ahead settings. Computational latency is controlled through an early-exit mechanism, enabling inference at different network depths. To narrow the performance gap between specialized and flexible models, we propose a two-stage training strategy with a shared-to-multiple decoder transition. Overall, the proposed framework enables a single model to be deployed across diverse latency budgets without retraining separate models. Model weights are available for download at: https://huggingface.co/nvidia/Real-time_RE-USE
Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive prior trained on clean speech latent representations extracted from a neural audio codec. Adaptation is performed by minimizing the Kullback-Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments across multiple noisy speech datasets show consistent improvements in speech quality, particularly under training-testing noise mismatch conditions. Code and audio examples are available online.
Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda +1
CentraleSup´elec, IETR (UMR CNRS 6164), France · Inria at Univ. Grenoble Alpes, CNRS, LJK, France · Signal Processing Group, University of Hamburg, Germany
Speech enhancement language models achieve strong results when trained on discrete audio tokens, but their optimization relies on token-level cross-entropy rather than the perceptual metrics used for evaluation. We introduce a post-training stage for autoregressive speech enhancement language models using Group Sequence Policy Optimization (GSPO) with multi-metric perceptual rewards. Our method directly optimizes non-differentiable quality metrics (DNSMOS, WER, and UTMOS) as reward signals, without learned surrogates or offline preference pairs. Applied to two autoregressive base models, UniSE and GenSE, our approach achieves state-of-the-art results on the DNS2020 benchmark. A human evaluation ablation further shows that the composite multi-metric reward is preferred over any single-metric variant, confirming that multi-reward optimization avoids the reward hacking observed with single-metric training.
Frédéric Berdoz, Luca A. Lanzendörfer, Antonis Asonitis +1