Identity-Assisted Association of Unordered DOA Estimates for Neural Speech Source Tracking
Authors: Bing Yang, Di Liang, Xiaofei Li
Organizations: School of Engineering, Westlake University, Hangzhou, China · School of Artificial Intelligence, Tianjin University, Tianjin, China · Zhejiang University, Hangzhou, China · Westlake Institute for Advanced Study, Hangzhou, China
Tracking speech sources remains a challenge due to ambiguous data association arising from intermittent speech, close spatial proximity, and complex acoustic conditions. To address these issues, we propose an identity-assisted association that maps unordered direction-of-arrival (DOA) estimates to speaker-consistent source trajectories for reliable speech source tracking. Specifically, speaker identity embeddings are directly integrated into the model input as a complementary cue to spatial features. This enables maintaining identity consistency by combining long-term time-invariant vocal identity characteristics with the short-term continuity of spatial cues. To effectively process these heterogeneous inputs while accommodating their distinct characteristics, we design a unified neural tracker. Within this model, time self-attention modules capture the temporal evolution of each source, while source self-attention modules distinguish between competing source tracks. Experimental results demonstrate the superiority of the proposed neural tracker in mitigating association confusion for speech source tracking.
Figures & tables
Fig. 1: Block diagram of the proposed neural speech source tracker.
Method
MDR ↓
FAR ↓
MAE ↓
AssA ↑
DetA ↑
HOTA ↑
[ % ]
[ % ]
[ ∘ ]
[ % ]
[ % ]
[ % ]
CST-Former [ 13 ]
21.0
16.0
3.8
61.7
69.3
64.8
C-Conformer [ 15 ]
11.0
10.1
3.3
75.7
81.9
78.3
RNN [ 4 ]
12.7
8.3
2.9
60.5
81.0
69.1
Neural-SRP [ 7 ]
17.9
14.6
3.3
63.1
72.9
67.2
TR (proposed)
7.9
5.9
2.1
79.9
87.2
83.1
TABLE I: Performance comparison on synthetic data
Method
MDR ↓
FAR ↓
MAE ↓
AssA ↑
DetA ↑
HOTA ↑
[ % ]
[ % ]
[ ∘ ]
[ % ]
[ % ]
[ % ]
CST-Former [ 13 ]
24.2
24.2
4.2
44.8
62.9
52.4
C-Conformer [ 15 ]
20.1
18.1
4.2
55.5
69.1
60.9
RNN [ 4 ]
10.5
8.4
3.1
61.5
83.0
69.9
Neural-SRP [ 7 ]
28.8
25.6
3.9
45.0
56.6
49.4
TR (proposed)
9.3
5.9
2.3
70.0
85.8
76.4
TABLE II: Performance comparison on the LOCATA dataset
Method
MDR ↓
FAR ↓
MAE ↓
AssA ↑
DetA ↑
HOTA ↑
[ % ]
[ % ]
[ ∘ ]
[ % ]
[ % ]
[ % ]
Unordered DOA estimates (IPDnet [ 20 ] )
9.0
6.1
2.2
54.2
86.0
67.4
TR w/o spk self-attention
8.3
6.5
2.1
79.4
86.4
82.5
TR w/o time self-attention
10.1
6.6
2.1
60.2
84.7
70.5
TR w/o spk emb.
9.5
6.1
2.1
73.2
85.7
78.5
TR (proposed)
7.9
5.9
2.1
79.9
87.2
83.1
TABLE III: Ablation study on proposed tracker
Fig. 2: Illustration of DOA trajectory for (a) ground truth, (b) CST-Former [ 13 ] , (c) C-Conformer [ 15 ] , (d) RNN [ 4 ] , (e) Neural-SRP [ 7 ] , (f) proposed tracker w/o speaker embedding, and (g) proposed tracker w/ speaker embedding. Each row corresponds to a recording example. Different colors indicate different speaker identities.
University of Michigan, USA · National Taiwan University, Taipei, Taiwan · Signal Analysis and Interpretation Laboratory (SAIL), University of Southern California, USA +1