cs.SDSep 28, 2026

Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities

Authors: Xinyuan Qian, Yanghao Zhou, Ziyang Jiang, Yu Chen, Xinjia Zhu, Xueyan Chen, Qiquan Zhang, Zexu Pan, +5 more

Organizations: School of Computer and Communication Engineering, University of Science and Technology Beijing, 100083, China · Department of Computer Science and Technology, Beijing Institute of Technology, China · The Chinese University of Hong Kong, Shenzhen, China · Alibaba Token Foundry, Alibaba Group · Zhejiang Institute of Quality Sciences (Technology Innovation Center of the State Administration for Market Regulation), Hangzhou, 310018, China · School of Computer Software, Tianjin University, Tianjin, 300350, China · University Hospital rechts der Isar, Technical University of Munich, Munich, Germany · School of Artificial Intelligence, the Chinese University of Hong Kong, Shenzhen, 518172, China

Abstract

Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems conditioned on enrollment speech have progressed substantially, enrollment speech as a cue has inherent limitations. Its reliability degrades when the target and interfering speakers have similar voice characteristics, when intra-speaker variability (e.g. changes in emotion or speaking style) creates a mismatch between the enrollment and target speech, or when the enrollment itself is contaminated by noise or competing speakers. This review surveys deep-learning-based TSE from the perspective of auxiliary target cues drawn from multiple modalities. We organize existing methods according to five types of information used to isolate the target speaker: audio enrollment, visual, spatial, textual/semantic, and neural cues. We also trace the evolution from discriminative estimators to variational, diffusion, flow, codec, and foundation-model-based systems and summarize representative datasets and evaluation metrics. We review the benefits and limitations of different cues and discuss challenges involving synchronization, missing or unreliable observations, data scarcity, privacy, computational cost, and real-time operation. Finally, we summarize future directions concerning adaptive cue fusion, instruction-driven extraction, realistic evaluation, and trustworthy deployment. By jointly reviewing cue design, model architecture, training objectives, datasets, and evaluation metrics, this article provides an overview of the current landscape and open problems in multimodal TSE.

Figures & tables

Explore similar work

CardsList
  1. Unified Target-Speaker ASR with Text and Enrollment Speech Cues

    Sep 27, 2026Yuxiang Mei, Yuchen Yan, Dongxing Xu +2Target Speaker ExtractionSpeaker

  2. SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

    Jul 16, 2026Shuai Wang, Zihan Qian, Ke Zhang +9Target Speaker ExtractionSpeaker