cs.SDAug 13, 2026

HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement

Authors: Zhengyi Lu, Aswini Sivakumar, Jie Hu, Yao Qiang

Organizations: Department of Computer Science and Engineering, Oakland University, Rochester, MI 48309

Abstract

Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schrödinger Bridges (SB) shorten transport from noise to clean speech but leave inference cost only loosely tied to training. We propose HybridSB-MoE, a dual-domain framework that fills these gaps through three contributions unified by a single asymmetric design principle. (i) Asymmetric uncertainty fusion: The spectral path captures epistemic uncertainty via expert disagreement, while the waveform bridge models aleatoric variance through stochastic dynamics. We fuse them asymmetrically, allowing the mixing weight to adapt to distinct error regimes rather than average predictions. (ii) Heterogeneous MoE with top-k=2 routing across five distinct architectural archetypes, where architectural diversity makes the epistemic signal indicate which inductive bias fails rather than small perturbations among similar experts. (iii) Discretization bound (Theorem 1): path-consistency and trajectory regularizers together bound the K-step bridge sampling error in 2-Wasserstein distance at rate K-alpha, making small-K inference an objective-level guarantee rather than an empirical claim. On VoiceBank+DEMAND, HybridSB-MoE outperforms diffusion- and SB-based baselines at their step budgets while remaining competitive with consistency-distilled few-step methods.

Explore similar work

CardsList
  1. SE-MSB: End-to-End Unpaired Speech Enhancement using Mamba Schrödinger Bridges

    Sep 22, 2026Andreas Bagge, Andreas Nymand, Michael Riis Andersen +1Speech Enhancement

  2. Speech Enhancement Based on Drifting Models

    Apr 27, 2026Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn +2Speech EnhancementGenerative Framework