cs.SDSep 28, 2026

Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities

Authors: Xinyuan Qian, Yanghao Zhou, Ziyang Jiang, Yu Chen, Xinjia Zhu, Xueyan Chen, Qiquan Zhang, Zexu Pan, +5 more

Organizations: School of Computer and Communication Engineering, University of Science and Technology Beijing, 100083, China · Department of Computer Science and Technology, Beijing Institute of Technology, China · The Chinese University of Hong Kong, Shenzhen, China · Alibaba Token Foundry, Alibaba Group · Zhejiang Institute of Quality Sciences (Technology Innovation Center of the State Administration for Market Regulation), Hangzhou, 310018, China · School of Computer Software, Tianjin University, Tianjin, 300350, China · University Hospital rechts der Isar, Technical University of Munich, Munich, Germany · School of Artificial Intelligence, the Chinese University of Hong Kong, Shenzhen, 518172, China

Abstract

Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems conditioned on enrollment speech have progressed substantially, enrollment speech as a cue has inherent limitations. Its reliability degrades when the target and interfering speakers have similar voice characteristics, when intra-speaker variability (e.g. changes in emotion or speaking style) creates a mismatch between the enrollment and target speech, or when the enrollment itself is contaminated by noise or competing speakers. This review surveys deep-learning-based TSE from the perspective of auxiliary target cues drawn from multiple modalities. We organize existing methods according to five types of information used to isolate the target speaker: audio enrollment, visual, spatial, textual/semantic, and neural cues. We also trace the evolution from discriminative estimators to variational, diffusion, flow, codec, and foundation-model-based systems and summarize representative datasets and evaluation metrics. We review the benefits and limitations of different cues and discuss challenges involving synchronization, missing or unreliable observations, data scarcity, privacy, computational cost, and real-time operation. Finally, we summarize future directions concerning adaptive cue fusion, instruction-driven extraction, realistic evaluation, and trustworthy deployment. By jointly reviewing cue design, model architecture, training objectives, datasets, and evaluation metrics, this article provides an overview of the current landscape and open problems in multimodal TSE.

Figures & tables

Explore similar work

Sep 27, 2026eess.SP

Unified Target-Speaker ASR with Text and Enrollment Speech Cues

Target-speaker automatic speech recognition (TS-ASR) aims to recognize a designated speaker while suppressing interfering speech in multi-talker environments. Conventional TS-ASR typically relies on an enrollment utterance, whereas text-guided methods use known lexical content, such as a wake word, to identify the target speaker from the observed mixture. These two cues provide complementary information but are usually studied separately. We propose a Unified Dual-Cue TS-ASR framework that supports text cues, enrollment speech, or both within a single model. Text cues interact with the mixture representation to extract target-speaker information conditioned on known lexical content, while an independent enrollment utterance provides complementary speaker information. Cross-attention cue-conditioning modules are integrated into shared Conformer blocks, and negative-cue sampling provides cue-validity supervision during dual-cue training. Experiments on 30,000 two-speaker mixtures across five recording/domain conditions and four oracle text-cue lengths show that, with five-character text cues, the concatenated dual-cue method achieves 8.80% CER, compared with 17.32% for text-only and 29.06% for enrollment-only inference. It also outperforms parallel dual-cue fusion (9.49% CER) and yields lower dual-cue CER across all five evaluation subsets. These results demonstrate the benefit of jointly exploiting complementary lexical and speaker information for target-speaker ASR.
Sep 24, 2026cs.SD

Exploring a Single Autoregressive LLM for Unified Target Speech Extraction across Synchronous and Asynchronous Cues

Target speech extraction (TSE) typically trains a separate extractor per cue, and visual-cue systems often need corruption-matched training to remain robust under visual frame corruption. We show that one autoregressive LLM backbone, TSE-Omni, can serve both temporally synchronous cues (lip movements, co-speech gestures) and asynchronous cues (enrollment audio, text). TSE-Omni is driven by next-token prediction: each step predicts target speech semantic tokens from its own past outputs, which we term self-enrollment, forming a continuous target-speech context initialized by the enrollment cue (asynchronous audio or text, or a short visual prefix). This enables audio-visual compensation: the model uses synchronized visuals when intact and its token history when visual frames are missing. Under clean visuals, TSE-Omni matches strong discriminative and generative baselines (SpeechBERTScore 0.81 on VoxCeleb2 and 0.89 on LRS3 zero-shot) with higher DNSMOS. On the same VoxCeleb2 test set, after a 2 s clean visual start, removing the remaining visual frames leaves SpeechBERTScore at 0.81. It remains usable under sparse overlap and multi-speaker interference, and supports streaming inference. Project page: https://alexwxwu.github.io/tseomni-main/.
Jul 16, 2026eess.AS

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.