Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framework that addresses hallucination through complementary modality and inference designs. First, a double-stream-to-single-stream multimodal Diffusion Transformer (MMDiT) allows transcript and acoustic representations to interact while preserving a dedicated pathway for the degraded input. Second, we find that the best decoder configuration, governed by guidance scale, sampling temperature, and step count, varies substantially across utterances. This observation motivates Inference Policy Optimization (IPO), an online, on-policy preference optimization method. IPO generates multiple candidates from the current model under different inference configurations, ranks them with a multi-objective reward, and learns from their relative preferences. AuraSE-IPO ranks first on 11 of 12 metrics across the synthetic test sets and obtains the highest DNSMOS and blind-listening scores among the evaluated systems on the real DNS blind test set. At deployment, it uses a fixed 10-step ODE decoder without classifier-free guidance (CFG) or per-utterance configuration search.
Figures & tables
Figure 1: Overview of AuraSE . (left) A double-stream-to-single-stream MMDiT enhances the degraded mel conditioned on the transcript, predicting the flow velocity. (center) A double-stream block: mel and text tokens attend jointly but keep separate projections. (right) IPO decodes each input under diverse inference policies, scores the candidates, forms self-generated preference pairs in a buffer, and updates the model in a closed loop.
Property
DPO ( Zhang et al. 2026 )
GRPO ( Wang et al. 2026 )
IPO (Ours)
Preference data
Static off-policy corpus
On-policy SDE rollouts
Refreshed on-policy buffer
Policy regime
Off-policy
On-policy
On-policy
Sampler (train / deploy)
ODE / matched
SDE / mismatched
ODE / matched
Multi-policy candidates
✗
✗
✓
Trajectory likelihood
Not needed
Per-step Gaussian
Not needed
Learning signal
Pairwise contrastive
Policy gradient
Pairwise contrastive
Table 1: Reinforcement-learning setups compared in this work; DPO and GRPO denote our reproductions. IPO retains ODE sampling, builds pairs from multi-configuration outputs, and periodically refreshes its on-policy buffer.
Figure 2: Flow-matching inference is sample-dependent and a dense preference source. Each of 300 utterances is decoded under 8 policies (CFG ∈{0,1} , temp. ∈{0.7,1.0} , steps ∈{10,20} ) and scored by the reward. (a) Best-of-eight selection improves on every tested fixed policy; (b) no policy dominates the per-utterance winner; (c) reward gaps supply many potential preference pairs.
DPO
GRPO
IPO
Reward spread
0.53
0.54
0.58
Mean reward
7.32
7.25
7.38
Table 2: Inference-policy variation yields a larger reward spread and higher mean reward than stochastic reseeding or SDE perturbation (randomly sampled 20 -utterance post-hoc diagnostic subset, 8 candidates/utterance, reward max 10 ; reward spread = mean per-utterance best − worst reward gap).
With Reverb
Without Reverb
Model
Type
SIG
BAK
OVRL
WER
SIM
SBS
SIG
BAK
OVRL
WER
SIM
SBS
Noisy Input
-
2.279
2.185
1.817
10.82%
0.763
0.840
3.063
2.690
2.371
9.02%
0.859
0.862
Baselines
VoiceFixer
-
3.393
4.010
3.099
18.80%
0.564
0.881
3.425
4.047
3.142
13.10%
0.655
0.891
PASE
-
3.231
3.854
2.865
10.39%
0.716
0.882
3.492
4.066
3.209
8.05%
0.803
0.903
LLaSE-G1
AR
3.019
3.504
2.570
28.78%
0.386
0.845
3.135
3.671
2.716
17.62%
0.465
0.856
Table 3: Main results on the synthetic benchmark ( 824 reverberant +824 non-reverberant). Higher is better except WER; bold / underline = best/second among enhancement systems, excluding the noisy-input reference.
Model
SIG
BAK
OVRL
Human ↑
Noisy Input
3.053
2.509
2.255
-
Baselines
VoiceFixer
3.291
3.959
2.991
-
PASE
3.490
4.057
3.203
3.33±0.15
LLaSE-G1
3.257
4.014
2.987
-
AnyEnhance
3.488
3.977
3.161
3.43±0.16
Table 4: Real-world DNS Challenge blind test set. SIG/BAK/OVRL are non-intrusive DNSMOS scores; Human reports the input-referenced blind-listening composite (mean ± 95% CI half-width).
Model
OVRL
WER
SIM
SBS
Noisy
2.101
9.91%
0.817
0.854
Single-stream design
w/ ASR text
3.202
10.97%
0.734
0.894
w/ GT text
3.193
10.97%
0.735
0.894
w/o text
3.196
11.24%
0.727
0.892
Double-stream design (AuraSE)
Table 5: Architecture and text-conditioning ablation ( 300 -utterance subset): single-stream vs. AuraSE’s double-stream design, each with ground-truth text, ASR text, and no text.
Method
OVRL
WER
SIM
SBS
AuraSE (base)
3.204
8.81%
0.758
0.902
Oracle ‡
3.271
5.87%
0.772
0.900
Offline RL
GT-DPO
3.113
13.35%
0.636
0.889
DPO
3.251
9.42%
0.742
0.905
IPO off†
3.280
7.97%
0.764
0.909
Table 6: RL ablation ( 300 -utterance subset; shared AuraSE base). DPO ( Zhang et al. 2026 ) (offline) and GRPO ( Wang et al. 2026 ) (online) are prior baselines; † marks ablation-only variants (online DPO, offline IPO) isolating the preference source from the offline/online axis; GT-DPO is a ground-truth-reference control; ‡ is the non-deployable per-sample oracle (best of 8 base policies).
Figure 3: Online training dynamics vs. compute (GPU-hours, EMA-smoothed). (a) Mean rollout reward for five online schemes; (b) reward spread among rollout policies over training.
We propose DriftSE, a novel one-step generative framework for speech enhancement formulated as a latent distribution equilibrium problem. During training, the drifting field aligns the generator's pushforward distribution with the clean speech manifold through drifting in a latent domain. During inference, the drifting process is discarded, enabling one-step generation. We establish that its enhancement quality depends fundamentally on the choice of latent representation. Semantic latents preserve phonetic structure but fail to capture physical acoustic cues, whereas acoustic latents reconstruct the physical signal but risk linguistic hallucination. Therefore, we introduce dual-latent drifting, performing parallel drifting in both semantic and acoustic latents to simultaneously preserve phonetic intelligibility and acoustic fidelity. Additionally, we demonstrate that DriftSE enables fully unpaired training by aligning latent distributions rather than exact point-wise targets. Consequently, DriftSE facilitates cross-dataset learning in the absence of paired noisy-clean samples. Moreover, DriftSE exhibits broad architectural flexibility across different generator backbones. Extensive evaluations on additive denoising and convolutive dereverberation demonstrate robust one-step enhancement across both offline and real-time causal settings. Notably, DriftSE achieves state-of-the-art word error rates across all four evaluated datasets while strictly operating at 1 NFE. Code and audio examples are available online.
Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn +2
Victoria University of Wellington, New Zealand · GN Advanced Science, Denmark · Lincoln University
Whisper, a widely adopted ASR model, is known to suffer from hallucinations - coherent transcriptions generated for non-speech audio entirely disconnected from the input. We investigate whether hallucinations can be detected and mitigated through Whisper's internal representations. We extract audio encoder activations and evaluate two representation spaces: raw Whisper activations and Sparse AutoEncoder (SAE) latents. We show that both spaces encode linearly separable hallucination-related information, with discriminative power concentrated in a sparse feature subset and increasing toward deeper encoder layers. We propose two steering strategies: activation-space steering and SAE latent-space steering. SAE-based steering reduces hallucination rate from 72.63% to 14.11% for Whisper small and from 86.88% to 27.33% for Whisper large-v3 on the full non-speech test set, with small WER degradation on speech data, approaching the performance of fine-tuning-based methods.
Georgii Aparin, Vadim Popov, Tasnima Sadekova +1
AI Foundation and Algorithm Lab · National Research University Higher School of Economics
Attention encoder-decoder (AED) Speech Foundation Models achieve strong ASR performance but can generate acoustically unsupported text when inputs contain no speech, weak acoustic evidence, or unreliable transcription. We propose AURA: Activation-editing with Uncertainty-Routed Adaptation, an ultra-efficient representation-editing method that freezes the pretrained model and applies sparse scale-and-shift edits to decoder cross-attention heads. AURA dynamically routes edits using cross-attention uncertainty features that capture over-concentration, diffuse attention, and abrupt frame shifts. We evaluate AURA on four datasets spanning non-speech hallucination and speech grounding stressors, including imperfect-label child speech, imperfect-label adult speech, and disfluent speech. On non-speech audio, AURA reduces hallucination rate from 89.18% to 1.94% without prior hallucination-head identification. On imperfect-label corpora, AURA approaches LoRA WER while using roughly 500x fewer trainable parameters. Sensitivity analysis and qualitative cross-attention examples are consistent with AURA's uncertainty-routed editing behavior, supporting dynamic activation editing as a practical path for grounding AED speech models.
Natarajan Balaji Shankar, Zilai Wang, Zihan Wang +3
Department of Electrical and Computer Engineering University of California Los Angeles Los Angeles, USA