cs.SDOct 5, 2026

AuraSE: Low-Hallucination Generative Speech Enhancement via Multimodal Flow Matching and Inference Policy Optimization

Authors: Yingda Shen, Yao Qian, Yuxuan Hu, Junan Zhang, Yuxiang Wang, Hardik Hansrajbhai Chauhan, Yudong Li, Yufei Xia, +2 more

Organizations: The Chinese University of Hong Kong, Shenzhen · Microsoft Research

Abstract

Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framework that addresses hallucination through complementary modality and inference designs. First, a double-stream-to-single-stream multimodal Diffusion Transformer (MMDiT) allows transcript and acoustic representations to interact while preserving a dedicated pathway for the degraded input. Second, we find that the best decoder configuration, governed by guidance scale, sampling temperature, and step count, varies substantially across utterances. This observation motivates Inference Policy Optimization (IPO), an online, on-policy preference optimization method. IPO generates multiple candidates from the current model under different inference configurations, ranks them with a multi-objective reward, and learns from their relative preferences. AuraSE-IPO ranks first on 11 of 12 metrics across the synthetic test sets and obtains the highest DNSMOS and blind-listening scores among the evaluated systems on the real DNS blind test set. At deployment, it uses a fixed 1010-step ODE decoder without classifier-free guidance (CFG) or per-utterance configuration search.

Figures & tables

Explore similar work

CardsList
  1. DriftSE: Speech Enhancement with Generative Drifting

    Sep 14, 2026Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn +2Speech EnhancementAcoustic Latent Space

  2. Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

    Jun 5, 2026Georgii Aparin, Vadim Popov, Tasnima Sadekova +1Citation Hallucination DetectionObject Hallucination

  3. AURA: Uncertainty-Routed Activation Editing for Acoustic Grounding in Speech Foundation Models

    Sep 21, 2026Natarajan Balaji Shankar, Zilai Wang, Zihan Wang +3Class Activation Mapping