Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recently, diffusion-based approaches have emerged as a promising alternative by leveraging powerful generative priors. Among them, ArrayDPS formulates BSS problem as a posterior sampling problem, and utilizes a pretrained speech diffusion model to guide the recovery of clean source signals. A key factor behind its separation capability is the multi-channel consistency (MC) objective, which enforces the estimated source signals to reconstruct the observed microphone mixtures through the estimated acoustic transfer functions. However, the number of microphones in the array is often limited, which constrains the performance of ArrayDPS. To address this issue, we propose VM-ArrayDPS, a novel method that augments the microphone array with virtual microphones with higher-SNR, these microphones can offer extra MC constraints to enhance the separation performance. Experimental results demonstrate that VM-ArrayDPS significantly outperforms ArrayDPS on both 2-speaker and 3-speaker datasets, showcasing the effectiveness of virtual microphone augmentation in improving BSS performance. We also did ablation studies to show the influence of the number of virtual microphones and weight of the MC objective brought by virtual microphones.
Figures & tables
Figure 1: Gradient computation of VM-ArrayDPS at a diffusion step. Given the current source estimate sτ , the physical and virtual microphone observations are reconstructed and used to compute the corresponding likelihood gradients.
Method
SDR (dB) ↑
SI-SDR (dB) ↑
PESQ ↑
eSTOI ↑
Mixture
0.1
0.0
1.87
0.603
Spatial Clustering †
9.5
8.5
2.52
0.759
IVA-Laplace †
12.0
10.7
2.67
0.802
IVA-Gaussian †
13.4
12.2
2.82
0.834
UNSSOR †
15.4
14.4
3.20
0.875
ArrayDPS-AVG
15.8
15.0
3.37
0.864
Table 1: Results on 3-channel 2-speaker SMS-WSJ test set. Results marked with † are obtained from existing studies. All VM-ArrayDPS results use two virtual microphones with α=0.5 .
Method
SDR (dB) ↑
SI-SDR (dB) ↑
PESQ ↑
eSTOI ↑
Mixture
−3.0
−3.2
1.57
0.458
ArrayDPS-AVG
11.5
10.5
2.82
0.763
VM-ArrayDPS-AVG
13.2
12.1
2.95
0.798
ArrayDPS-Max
13.1
12.2
3.00
0.798
VM-ArrayDPS-Max
14.1
13.2
3.07
0.822
ArrayDPS-ML
12.8
11.9
3.00
0.792
Table 2: Results on 4-channel 3-speaker SMS-WSJ test set.
#virtual microphones
α
SDR (dB) ↑
SI-SDR (dB) ↑
PESQ ↑
eSTOI ↑
0
–
15.8
15.0
3.37
0.864
2
1.0
17.0
16.1
3.44
0.883
2
0.5
16.9
16.1
3.45
0.881
6
1.0
17.0
16.1
3.43
0.882
6
0.5
17.0
16.2
3.45
0.881
6
0.1
16.3
15.5
3.41
0.868
Table 3: Ablation study of number of virtual microphones and virtual-microphone loss weight α on 3-channel 2-speaker SMS-WSJ test set. All results use AVG sample aggregation.