Identity-Assisted Association of Unordered DOA Estimates for Neural Speech Source Tracking
Authors: Bing Yang, Di Liang, Xiaofei Li
Organizations: School of Engineering, Westlake University, Hangzhou, China · School of Artificial Intelligence, Tianjin University, Tianjin, China · Zhejiang University, Hangzhou, China · Westlake Institute for Advanced Study, Hangzhou, China
Tracking speech sources remains a challenge due to ambiguous data association arising from intermittent speech, close spatial proximity, and complex acoustic conditions. To address these issues, we propose an identity-assisted association that maps unordered direction-of-arrival (DOA) estimates to speaker-consistent source trajectories for reliable speech source tracking. Specifically, speaker identity embeddings are directly integrated into the model input as a complementary cue to spatial features. This enables maintaining identity consistency by combining long-term time-invariant vocal identity characteristics with the short-term continuity of spatial cues. To effectively process these heterogeneous inputs while accommodating their distinct characteristics, we design a unified neural tracker. Within this model, time self-attention modules capture the temporal evolution of each source, while source self-attention modules distinguish between competing source tracks. Experimental results demonstrate the superiority of the proposed neural tracker in mitigating association confusion for speech source tracking.
Figures & tables
Fig. 1: Block diagram of the proposed neural speech source tracker.
Method
MDR ↓
FAR ↓
MAE ↓
AssA ↑
DetA ↑
HOTA ↑
[ % ]
[ % ]
[ ∘ ]
[ % ]
[ % ]
[ % ]
CST-Former [ 13 ]
21.0
16.0
3.8
61.7
69.3
64.8
C-Conformer [ 15 ]
11.0
10.1
3.3
75.7
81.9
78.3
RNN [ 4 ]
12.7
8.3
2.9
60.5
81.0
69.1
Neural-SRP [ 7 ]
17.9
14.6
3.3
63.1
72.9
67.2
TR (proposed)
7.9
5.9
2.1
79.9
87.2
83.1
TABLE I: Performance comparison on synthetic data
Method
MDR ↓
FAR ↓
MAE ↓
AssA ↑
DetA ↑
HOTA ↑
[ % ]
[ % ]
[ ∘ ]
[ % ]
[ % ]
[ % ]
CST-Former [ 13 ]
24.2
24.2
4.2
44.8
62.9
52.4
C-Conformer [ 15 ]
20.1
18.1
4.2
55.5
69.1
60.9
RNN [ 4 ]
10.5
8.4
3.1
61.5
83.0
69.9
Neural-SRP [ 7 ]
28.8
25.6
3.9
45.0
56.6
49.4
TR (proposed)
9.3
5.9
2.3
70.0
85.8
76.4
TABLE II: Performance comparison on the LOCATA dataset
Method
MDR ↓
FAR ↓
MAE ↓
AssA ↑
DetA ↑
HOTA ↑
[ % ]
[ % ]
[ ∘ ]
[ % ]
[ % ]
[ % ]
Unordered DOA estimates (IPDnet [ 20 ] )
9.0
6.1
2.2
54.2
86.0
67.4
TR w/o spk self-attention
8.3
6.5
2.1
79.4
86.4
82.5
TR w/o time self-attention
10.1
6.6
2.1
60.2
84.7
70.5
TR w/o spk emb.
9.5
6.1
2.1
73.2
85.7
78.5
TR (proposed)
7.9
5.9
2.1
79.9
87.2
83.1
TABLE III: Ablation study on proposed tracker
Fig. 2: Illustration of DOA trajectory for (a) ground truth, (b) CST-Former [ 13 ] , (c) C-Conformer [ 15 ] , (d) RNN [ 4 ] , (e) Neural-SRP [ 7 ] , (f) proposed tracker w/o speaker embedding, and (g) proposed tracker w/ speaker embedding. Each row corresponds to a recording example. Different colors indicate different speaker identities.
Natural conversations make both speech recognition and speaker attribution challenging for ASR, as speakers take turns, overlap, and reappear over time. We propose STAM-ASR, Speaker-Temporal Anchoring with Memory, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR. Without relying on an external diarization system, STAM-ASR learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features. Hence providing explicit who and when cues to modulate the AudioLLM's semantic representation without explicit speech separation. STAM-ASR further maintains fixed-size speaker and conversational memories to carry complementary context across turns. We evaluate STAM-ASR on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions. Our reported results shows that speaker-temporal conditioning and memory provide complementary benefits, while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
Victor Tolulope Olufemi, Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury
Qatar Computing Research Institute (QCRI), Doha, Qatar
In long-form multi-party conversations, highly imbalanced speaker activity and frequent overlap make it difficult to identify "who spoke when and what". Sliding-window continuous speech separation (CSS) mitigates sparse supervision, but often suffers from cross-window speaker inconsistency and residual crosstalk, which in practice requires diarization for reliable speaker attribution. Motivated by the stability of speakers' directions of arrival (DOAs) in meetings, we propose PATSE, a multi-channel Position-Aware Target Speaker Extraction front-end that uses DOA as a spatial prior to directly extract the speech of each target speaker. PATSE combines a DOA-guided spatial encoder and conditioner to generate speaker-attributed streams, from which speaker activity can be inferred via simple post-processing (e.g., VAD) without explicit diarization. Experiments on both replayed and real conversations show consistent ASR gains outperforming CSS and diarization-based pipelines.
Yichi Wang, Junzhe Chen, Wangjin Zhou +1
Graduate School of Informatics, Kyoto University, Kyoto, Japan
As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (NVVs) to enhance naturalness, reliable speaker verification (SV) becomes essential to objectively assess identity consistency across both verbal and non-verbal segments. Yet current SV systems generalize poorly to NVVs, and fine-tuning on NVV data causes catastrophic forgetting of speech performance. We present the first systematic study across 10 NVV types and propose a framework combining frozen Data2Vec self-supervised features with ECAPA-TDNN, enhanced by a Mixture of Experts (MoE) module with learned domain-aware routing. A conditional distillation loss on speech inputs via a pretrained teacher retains speech-to-speech accuracy, while a contrastive loss bridges the speech-NVV domain gap. Our method reduces speech-NVV EER from 38.93% to 22.66% over a pretrained baseline, and improves speech EER from 13.17% to 9.24% via distillation.
Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou +4
University of Michigan, USA · National Taiwan University, Taipei, Taiwan · Signal Analysis and Interpretation Laboratory (SAIL), University of Southern California, USA +1