Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recently, diffusion-based approaches have emerged as a promising alternative by leveraging powerful generative priors. Among them, ArrayDPS formulates BSS problem as a posterior sampling problem, and utilizes a pretrained speech diffusion model to guide the recovery of clean source signals. A key factor behind its separation capability is the multi-channel consistency (MC) objective, which enforces the estimated source signals to reconstruct the observed microphone mixtures through the estimated acoustic transfer functions. However, the number of microphones in the array is often limited, which constrains the performance of ArrayDPS. To address this issue, we propose VM-ArrayDPS, a novel method that augments the microphone array with virtual microphones with higher-SNR, these microphones can offer extra MC constraints to enhance the separation performance. Experimental results demonstrate that VM-ArrayDPS significantly outperforms ArrayDPS on both 2-speaker and 3-speaker datasets, showcasing the effectiveness of virtual microphone augmentation in improving BSS performance. We also did ablation studies to show the influence of the number of virtual microphones and weight of the MC objective brought by virtual microphones.
Figures & tables
Figure 1: Gradient computation of VM-ArrayDPS at a diffusion step. Given the current source estimate sτ , the physical and virtual microphone observations are reconstructed and used to compute the corresponding likelihood gradients.
Method
SDR (dB) ↑
SI-SDR (dB) ↑
PESQ ↑
eSTOI ↑
Mixture
0.1
0.0
1.87
0.603
Spatial Clustering †
9.5
8.5
2.52
0.759
IVA-Laplace †
12.0
10.7
2.67
0.802
IVA-Gaussian †
13.4
12.2
2.82
0.834
UNSSOR †
15.4
14.4
3.20
0.875
ArrayDPS-AVG
15.8
15.0
3.37
0.864
Table 1: Results on 3-channel 2-speaker SMS-WSJ test set. Results marked with † are obtained from existing studies. All VM-ArrayDPS results use two virtual microphones with α=0.5 .
Method
SDR (dB) ↑
SI-SDR (dB) ↑
PESQ ↑
eSTOI ↑
Mixture
−3.0
−3.2
1.57
0.458
ArrayDPS-AVG
11.5
10.5
2.82
0.763
VM-ArrayDPS-AVG
13.2
12.1
2.95
0.798
ArrayDPS-Max
13.1
12.2
3.00
0.798
VM-ArrayDPS-Max
14.1
13.2
3.07
0.822
ArrayDPS-ML
12.8
11.9
3.00
0.792
Table 2: Results on 4-channel 3-speaker SMS-WSJ test set.
#virtual microphones
α
SDR (dB) ↑
SI-SDR (dB) ↑
PESQ ↑
eSTOI ↑
0
–
15.8
15.0
3.37
0.864
2
1.0
17.0
16.1
3.44
0.883
2
0.5
16.9
16.1
3.45
0.881
6
1.0
17.0
16.1
3.43
0.882
6
0.5
17.0
16.2
3.45
0.881
6
0.1
16.3
15.5
3.41
0.868
Table 3: Ablation study of number of virtual microphones and virtual-microphone loss weight α on 3-channel 2-speaker SMS-WSJ test set. All results use AVG sample aggregation.
Audio spotforming is a technique for extracting target speech from noisy mixtures by utilizing multiple microphone arrays. Conventional methods estimate a shared target speech component from linearly separated signals obtained by each array using low-rank approximations and apply post filtering (PF) based on this estimated low-rank representation. However, owing to the mismatch between low-rank models and the complex structure of speech signals, directly relying on low-rank approximations for PF can degrade the speech extraction performance. In this study, we leverage the observation that non-target components located in the target speech direction from the perspective of one array can be spatially separated when viewed from other arrays. This insight motivates a new spotforming method for efficient post-filter estimation using non-target estimates across arrays instead of relying on low-rank approximations. Experiments demonstrate that the proposed method outperforms conventional spotforming methods.
Yuto Ishikawa, Li Li, Shogo Seki +1
CyberAgent, Inc., Japan · The University of Tokyo, Japan
This paper proposes a general framework for stable and effective iterative audio separation with mixture consistency by extending source separation models to a multi-input multi-output (MIMO) configuration. In the field of audio separation, mixture consistency is an essential property for many applications that require accurate phase and timbral information of target sources. While iterative approaches such as diffusion models achieve perceptually superior results in speech enhancement or user-guided target source separation tasks, most existing methods focus on single-step separation with a single-input single-output (SISO) or single-input multi-output (SIMO) configuration through architectural improvements, since mixture-consistent audio separation is generally regarded as a regression problem that admits a unique solution. By extending these architectures to a MIMO configuration, we introduce iterative prediction without compromising the architectural advantages or the characteristics of mixture consistency. We conduct a comprehensive ablation study of combining the framework with discriminators and extending it to a generative model. Experimental results demonstrate significant performance improvements when applying the proposed framework to state-of-the-art separation models.
Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-N candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.
Anastasia Zorkina, Alexandr Anikin, Nikita Khmelev +5
ITMO University, Speech Processing Group, Russia · Speech Technology Center Ltd., R&D department, Russia.