Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framework that addresses hallucination through complementary modality and inference designs. First, a double-stream-to-single-stream multimodal Diffusion Transformer (MMDiT) allows transcript and acoustic representations to interact while preserving a dedicated pathway for the degraded input. Second, we find that the best decoder configuration, governed by guidance scale, sampling temperature, and step count, varies substantially across utterances. This observation motivates Inference Policy Optimization (IPO), an online, on-policy preference optimization method. IPO generates multiple candidates from the current model under different inference configurations, ranks them with a multi-objective reward, and learns from their relative preferences. AuraSE-IPO ranks first on 11 of 12 metrics across the synthetic test sets and obtains the highest DNSMOS and blind-listening scores among the evaluated systems on the real DNS blind test set. At deployment, it uses a fixed 10-step ODE decoder without classifier-free guidance (CFG) or per-utterance configuration search.
Figures & tables
Figure 1: Overview of AuraSE . (left) A double-stream-to-single-stream MMDiT enhances the degraded mel conditioned on the transcript, predicting the flow velocity. (center) A double-stream block: mel and text tokens attend jointly but keep separate projections. (right) IPO decodes each input under diverse inference policies, scores the candidates, forms self-generated preference pairs in a buffer, and updates the model in a closed loop.
Property
DPO ( Zhang et al. 2026 )
GRPO ( Wang et al. 2026 )
IPO (Ours)
Preference data
Static off-policy corpus
On-policy SDE rollouts
Refreshed on-policy buffer
Policy regime
Off-policy
On-policy
On-policy
Sampler (train / deploy)
ODE / matched
SDE / mismatched
ODE / matched
Multi-policy candidates
✗
✗
✓
Trajectory likelihood
Not needed
Per-step Gaussian
Not needed
Learning signal
Pairwise contrastive
Policy gradient
Pairwise contrastive
Table 1: Reinforcement-learning setups compared in this work; DPO and GRPO denote our reproductions. IPO retains ODE sampling, builds pairs from multi-configuration outputs, and periodically refreshes its on-policy buffer.
Figure 2: Flow-matching inference is sample-dependent and a dense preference source. Each of 300 utterances is decoded under 8 policies (CFG ∈{0,1} , temp. ∈{0.7,1.0} , steps ∈{10,20} ) and scored by the reward. (a) Best-of-eight selection improves on every tested fixed policy; (b) no policy dominates the per-utterance winner; (c) reward gaps supply many potential preference pairs.
DPO
GRPO
IPO
Reward spread
0.53
0.54
0.58
Mean reward
7.32
7.25
7.38
Table 2: Inference-policy variation yields a larger reward spread and higher mean reward than stochastic reseeding or SDE perturbation (randomly sampled 20 -utterance post-hoc diagnostic subset, 8 candidates/utterance, reward max 10 ; reward spread = mean per-utterance best − worst reward gap).
With Reverb
Without Reverb
Model
Type
SIG
BAK
OVRL
WER
SIM
SBS
SIG
BAK
OVRL
WER
SIM
SBS
Noisy Input
-
2.279
2.185
1.817
10.82%
0.763
0.840
3.063
2.690
2.371
9.02%
0.859
0.862
Baselines
VoiceFixer
-
3.393
4.010
3.099
18.80%
0.564
0.881
3.425
4.047
3.142
13.10%
0.655
0.891
PASE
-
3.231
3.854
2.865
10.39%
0.716
0.882
3.492
4.066
3.209
8.05%
0.803
0.903
LLaSE-G1
AR
3.019
3.504
2.570
28.78%
0.386
0.845
3.135
3.671
2.716
17.62%
0.465
0.856
Table 3: Main results on the synthetic benchmark ( 824 reverberant +824 non-reverberant). Higher is better except WER; bold / underline = best/second among enhancement systems, excluding the noisy-input reference.
Model
SIG
BAK
OVRL
Human ↑
Noisy Input
3.053
2.509
2.255
-
Baselines
VoiceFixer
3.291
3.959
2.991
-
PASE
3.490
4.057
3.203
3.33±0.15
LLaSE-G1
3.257
4.014
2.987
-
AnyEnhance
3.488
3.977
3.161
3.43±0.16
Table 4: Real-world DNS Challenge blind test set. SIG/BAK/OVRL are non-intrusive DNSMOS scores; Human reports the input-referenced blind-listening composite (mean ± 95% CI half-width).
Model
OVRL
WER
SIM
SBS
Noisy
2.101
9.91%
0.817
0.854
Single-stream design
w/ ASR text
3.202
10.97%
0.734
0.894
w/ GT text
3.193
10.97%
0.735
0.894
w/o text
3.196
11.24%
0.727
0.892
Double-stream design (AuraSE)
Table 5: Architecture and text-conditioning ablation ( 300 -utterance subset): single-stream vs. AuraSE’s double-stream design, each with ground-truth text, ASR text, and no text.
Method
OVRL
WER
SIM
SBS
AuraSE (base)
3.204
8.81%
0.758
0.902
Oracle ‡
3.271
5.87%
0.772
0.900
Offline RL
GT-DPO
3.113
13.35%
0.636
0.889
DPO
3.251
9.42%
0.742
0.905
IPO off†
3.280
7.97%
0.764
0.909
Table 6: RL ablation ( 300 -utterance subset; shared AuraSE base). DPO ( Zhang et al. 2026 ) (offline) and GRPO ( Wang et al. 2026 ) (online) are prior baselines; † marks ablation-only variants (online DPO, offline IPO) isolating the preference source from the offline/online axis; GT-DPO is a ground-truth-reference control; ‡ is the non-deployable per-sample oracle (best of 8 base policies).
Figure 3: Online training dynamics vs. compute (GPU-hours, EMA-smoothed). (a) Mean rollout reward for five online schemes; (b) reward spread among rollout policies over training.