cs.SDSep 27, 2026

Identity-Assisted Association of Unordered DOA Estimates for Neural Speech Source Tracking

Authors: Bing Yang, Di Liang, Xiaofei Li

Organizations: School of Engineering, Westlake University, Hangzhou, China · School of Artificial Intelligence, Tianjin University, Tianjin, China · Zhejiang University, Hangzhou, China · Westlake Institute for Advanced Study, Hangzhou, China

Abstract

Tracking speech sources remains a challenge due to ambiguous data association arising from intermittent speech, close spatial proximity, and complex acoustic conditions. To address these issues, we propose an identity-assisted association that maps unordered direction-of-arrival (DOA) estimates to speaker-consistent source trajectories for reliable speech source tracking. Specifically, speaker identity embeddings are directly integrated into the model input as a complementary cue to spatial features. This enables maintaining identity consistency by combining long-term time-invariant vocal identity characteristics with the short-term continuity of spatial cues. To effectively process these heterogeneous inputs while accommodating their distinct characteristics, we design a unified neural tracker. Within this model, time self-attention modules capture the temporal evolution of each source, while source self-attention modules distinguish between competing source tracks. Experimental results demonstrate the superiority of the proposed neural tracker in mitigating association confusion for speech source tracking.

Figures & tables

Explore similar work

CardsList
  1. STAM-ASR: Speaker-Temporal Anchoring with Memory for Multi-Speaker ASR

    Sep 24, 2026Victor Tolulope Olufemi, Syeda Faiza Ahmed Sara, Shammur Absar ChowdhuryAutomatic Speech RecognitionSpeaker

  2. Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR

    Jun 28, 2026Yichi Wang, Junzhe Chen, Wangjin Zhou +1Target Speaker ExtractionSpeaker

  3. Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

    Jun 19, 2026Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou +4Automatic Speaker VerificationSpeaker