eess.ASJul 21, 2026

Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution

Authors: Zhenglong Liu, Wangyou Zhang, Chenda Li, Yanmin Qian

Organizations: Auditory Cognition and Computational Acoustics Lab Shanghai Jiao Tong University, Shanghai, China · VUI Labs

Abstract

Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnostic SE methods address variable microphone numbers and permutations, they largely fail to exploit explicit array geometry priors when available, missing a crucial cue for optimal spatial filtering. A Geometry-Aware Dynamic Convolution (Geo-DConv) framework is proposed, which explicitly leverages microphone coordinates to transform standard fixed-array SE models into robust array-invariant systems. Experiments are conducted on the recent real-recorded RealMAN multi-channel speech dataset. Results demonstrate that the proposed architecture enables two widely used fixed-array models to adapt to array-invariant settings, with consistent performance improvements across diverse array topologies.

Explore similar work

CardsList