eess.ASOct 5, 2026

SEAL: Mixture-Closed Additive Reconstruction and Refinement-Aware Expert Routing for Efficient Speech Separation

Authors: Shao-Chun Hu, Zi-Xiang Lin, Jeih-Weih Hung, Hung-Shin Lee

Abstract

Compact time-frequency separators that mask the mixture and refine through a shared cell face two limits. First, a bounded multiplicative mask only scales a mixture bin, so where overlapping components cancel, the estimate stays small. Second, a shared cell applies the same weights to every time-frequency token at every step, so enlarging it adds compute everywhere. We present SEAL (Sparse Expert routing with Additive Latent reconstruction) to address both. For reconstruction, a zero-sum additive residual bounded by the local mixture amplitude lets estimates be nonzero where components cancel yet still sum to the mixture. For routing, a query built from acoustic and inter-step evidence sends each token to one of six residual experts, and a norm cap keeps the step cue from overriding clear acoustic evidence. On EchoSet, SEAL (small) surpasses TIGER (small) by 0.31 dB SI-SDRi with 28% fewer parameters and 2.9 times fewer MACs, and SEAL (large) is within 0.07 dB SI-SDRi of TIGER (large) at 3.1 times fewer MACs.

Explore similar work

CardsList
  1. TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

    Jun 28, 2026Qinzhe Hu, Chenda Li, Wangyou Zhang +3Sparse Mixture-of-ExpertsSpeech Processing

  2. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    Sep 15, 2025Adhiraj Banerjee, Vipul AroraAudio Source Separation