This paper addresses unsupervised speech enhancement in the unpaired setting using drifting methods, where training relies on separate collections of degraded and clean audio without corresponding pairs. While recent drifting approaches enable unpaired training, they do so at a heavy cost: because the objective optimizes only a marginal prior over clean speech, the enhancer gradually loses the input's linguistic content and speaker identity. To fix this, we introduce input-conditioned drifting. We preserve the pull of the clean corpus while re-tethering the output to the degraded input via two mechanisms: an anchor encoder supplies the missing likelihood by pulling toward the input's features, and a key encoder conditions the prior by re-weighting retrieved frames. Neither requires labels or paired data. Using a training-free encoder selection criterion, Word Error Rate on VoiceBank-DEMAND falls to 10.1% (unprocessed: 11.7%), speaker similarity recovers from 0.490 to 0.879, and the recipe transfers in part to dereverberation on WSJ0-REVERB: content improves, rendering quality does not.
Figures & tables
diagnostics (training-free)
outcome
latent
family
phone-sep
recov. R2
Jac. align
WER ↓
PESQ ↑
PANNs
tagger
0.16
0.46
+0.01
17.7
1.61
BEATs
tagger
0.36
0.70
+0.06
9.9
1.84
DistilHuBERT
SSL
0.60
0.86
+0.11
5.5
2.03
WavLM-large
SSL
0.68
0.87
+0.06
5.3
2.18
WavCube
vocoder
0.65
0.83
+0.11
5.0
2.30
Table 1: Screening encoders before training (WSJ): phone separability (linear probe at PCA-128), recoverability (ridge R2 of the clean log-spectrum), and the alignment a(ϕ) of ( 5 ) (chance ∼0.06 ), against the WER and PESQ of training one system per row with that latent as its bank, as reported in [ 17 ] . Phone separability orders WER, recoverability PESQ; alignment neither.
system
trains on
mech.
steps
params (M)
WER ↓ (D/I)
ECAPA ↑
DNSMOS ↑
SCOREQ ↑
PESQ ↑
STOI ↑
noisy input
—
—
—
—
11.7
0.888
2.70
3.31
1.97
0.921
Masking-based
MetricGAN-U [ 7 ]
mixtures
mask
1
1.9
16.1 (116/58)
0.667
2.81
3.37
2.13
0.889
SelfSE [ 9 ]
mix+noise
mask
1
1.5
9.3 (74/37)
0.860
3.18
4.13
2.98
0.943
Resynthesis-based
DiffUSEEN [ 6 ]
clean+noise
diff.
30
5.2
10.7 (83/37)
0.865
3.05
3.79
2.67
0.933
Table 2: Denoising. VoiceBank–DEMAND, complete 824-utterance test set. We rescore all methods in our pipeline from one render (external rows: authors’ released artefacts). D/I are deletions/insertions; DNSMOS is OVRL; SIG/BAK are quoted in Sec. 3.4 . “Steps”: sampler inference steps (a corrector is a separate net); “params”: inference-time parameters. Bold: best within each block.
system
WER ↓
ESTOI ↑
SRMR ↑
SCOREQ ↑
ECAPA ↑
reverberant input (unprocessed)
32.71
0.452
3.45
2.74
0.814
ours: 2 kernels + 2 anchors
23.81
0.586
12.93
2.37
0.745
no anchors
81.69
0.433
4.16
3.06
0.215
Table 3: Dereverberation. WSJ0-REVERB, complete 651-utterance test set, same pipeline and judge as Table 2 (wav2vec2-base-960h); WER in percent. The full system is two banked kernels plus two conditioning anchors (Sec. 3.1 ); the ablation removes only the anchors.
We propose DriftSE, a novel one-step generative framework for speech enhancement formulated as a latent distribution equilibrium problem. During training, the drifting field aligns the generator's pushforward distribution with the clean speech manifold through drifting in a latent domain. During inference, the drifting process is discarded, enabling one-step generation. We establish that its enhancement quality depends fundamentally on the choice of latent representation. Semantic latents preserve phonetic structure but fail to capture physical acoustic cues, whereas acoustic latents reconstruct the physical signal but risk linguistic hallucination. Therefore, we introduce dual-latent drifting, performing parallel drifting in both semantic and acoustic latents to simultaneously preserve phonetic intelligibility and acoustic fidelity. Additionally, we demonstrate that DriftSE enables fully unpaired training by aligning latent distributions rather than exact point-wise targets. Consequently, DriftSE facilitates cross-dataset learning in the absence of paired noisy-clean samples. Moreover, DriftSE exhibits broad architectural flexibility across different generator backbones. Extensive evaluations on additive denoising and convolutive dereverberation demonstrate robust one-step enhancement across both offline and real-time causal settings. Notably, DriftSE achieves state-of-the-art word error rates across all four evaluated datasets while strictly operating at 1 NFE. Code and audio examples are available online.
Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn +2
Victoria University of Wellington, New Zealand · GN Advanced Science, Denmark · Lincoln University
We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Rather than relying on iterative sampling, DriftSE natively achieves one-step inference by evolving the pushforward distribution of a mapping function to directly match the clean speech distribution. This evolution is driven by a Drifting Field, a learned correction vector that guides samples toward the high-density regions of the clean distribution, which naturally facilitates training on unpaired data by matching distributions rather than paired samples. We investigate the framework under two formulations: a direct mapping from the noisy observation, and a stochastic conditional generative model from a Gaussian prior. Experiments on the VoiceBank-DEMAND benchmark demonstrate that DriftSE achieves high-fidelity enhancement in a single step, outperforming multi-step diffusion baselines and establishing a new paradigm for speech enhancement.
Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn +2
Victoria University of Wellington, 3 Lincoln University, New Zealand · GN Advanced Science, Denmark
Speech enhancement (SE) models typically rely on supervised learning with paired data examples where clean speech is synthetically degraded. This paradigm limits performance in real-world scenarios where the target environment's specific acoustic characteristics are unknown. We propose a fully unpaired SE framework that uses principled Diffusion Schrödinger Bridges (DSB) to learn a stochastic transport process between a clean and a degraded speech distribution. Algorithms for learning transport maps are computationally heavy since they require simulating differential equations during training, usually at each training step. Therefore, we propose using a high-efficiency Mamba Diffusion Model designed for end-to-end waveform processing. We compare against state-of-the-art methods for speech enhancement, both paired and unpaired, as well as a classical signal processing algorithm. Experimental results show that we are on par or better than the baselines while being orders of magnitude faster during inference. Furthermore, we show that the flexibility of the DSB formulation allows our model to generalize across SE tasks, offering a robust and efficient solution for real-world speech restoration.
Andreas Bagge, Andreas Nymand, Michael Riis Andersen +1