Organizations: School of Computer and Communication Engineering, University of Science and Technology Beijing, 100083, China · Department of Computer Science and Technology, Beijing Institute of Technology, China · The Chinese University of Hong Kong, Shenzhen, China · Alibaba Token Foundry, Alibaba Group · Zhejiang Institute of Quality Sciences (Technology Innovation Center of the State Administration for Market Regulation), Hangzhou, 310018, China · School of Computer Software, Tianjin University, Tianjin, 300350, China · University Hospital rechts der Isar, Technical University of Munich, Munich, Germany · School of Artificial Intelligence, the Chinese University of Hong Kong, Shenzhen, 518172, China
Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems conditioned on enrollment speech have progressed substantially, enrollment speech as a cue has inherent limitations. Its reliability degrades when the target and interfering speakers have similar voice characteristics, when intra-speaker variability (e.g. changes in emotion or speaking style) creates a mismatch between the enrollment and target speech, or when the enrollment itself is contaminated by noise or competing speakers. This review surveys deep-learning-based TSE from the perspective of auxiliary target cues drawn from multiple modalities. We organize existing methods according to five types of information used to isolate the target speaker: audio enrollment, visual, spatial, textual/semantic, and neural cues. We also trace the evolution from discriminative estimators to variational, diffusion, flow, codec, and foundation-model-based systems and summarize representative datasets and evaluation metrics. We review the benefits and limitations of different cues and discuss challenges involving synchronization, missing or unreliable observations, data scarcity, privacy, computational cost, and real-time operation. Finally, we summarize future directions concerning adaptive cue fusion, instruction-driven extraction, realistic evaluation, and trustworthy deployment. By jointly reviewing cue design, model architecture, training objectives, datasets, and evaluation metrics, this article provides an overview of the current landscape and open problems in multimodal TSE.
Figures & tables
Fig. 1: Illustration of TSE guided by different auxiliary target cues i.e., audio enrollment, visual observations, spatial information, textual or semantic descriptions, neural responses, or combinations of multiple cues.
Fig. 2: Taxonomy of TSE methods according to auxiliary target cues e.g., audio, visual, spatial, semantic and neural information.
Fig. 3: Architecture of TSE networks. Discriminative systems directly estimate a TF representation or waveform of the target. The generative panel illustrates recent pipelines based on diffusion, flow matching, and LLMs. Earlier variational source modeling is reviewed in the text. The panel on the right presents pretraining and task unification as design dimensions that apply across paradigms. These approaches reuse large encoders, generative backbones, or shared models for TSE and related speech processing tasks. The cue encoders shown in each pipeline can be adapted to the auxiliary modalities reviewed in Sec. III .
Resource
Primary role
Cue or modality
Mixture or recording construction
Exact reference
Typical use in TSE research
Source speech and enrollment corpora
WSJ0 [ 169 ]
Clean source corpus
Audio identity
Separate read utterances; mixtures generated downstream
Not a mixture
Target, interference, and enrollment speech
LibriSpeech [ 170 ]
Clean source corpus
Identity; transcript
Separate audiobook utterances; mixtures generated downstream
Not a mixture
Large-scale source and enrollment speech
VoxCeleb1/2 [ 171 , 172 ]
In-the-wild source corpus
Audio-visual identity
Natural video clips; TSE mixtures usually combine independent clips
Source known
Speaker encoders, enrollment, and face-conditioned TSE
VCTK and AISHELL-1 [ 173 , 174 ]
Source corpora
Accent; language
Separate utterances with downstream mixing and augmentation
Not a mixture
Cross-corpus, accent, and language evaluation
Generated mixtures and augmentation resources
TABLE I: Representative source, mixture, augmentation, and recorded-scene resources used in TSE research. “Exact reference” indicates whether the target signal used to form the evaluated mixture is available for waveform supervision or signal-level scoring.
Resource
Primary role
Cue or modality
Mixture or recording construction
Reference status
Typical use in TSE research
Visual and audio-visual resources
GRID [ 187 ]
Controlled audio-visual corpus
Face; lip motion
Fixed-format sentences recorded from individual speakers
Source known
Lip-guided extraction and synchronization
CUAVE [ 188 ]
Controlled audio-visual corpus
Face; lip motion
Individual and multi-speaker recordings with frontal and profile views
Controlled utterances captured from several camera views
Source known
View-robust visual speech representation
LRS2/LRS3 [ 191 , 192 ]
In-the-wild audio-visual corpora
Face; lips; transcript
Television and TED or TEDx clips; mixtures generated downstream
Source known
Large-scale lip-guided extraction and synchronization
TABLE II: Representative cue-bearing multimodal resources used to construct or evaluate visual-, textual-, and neural-conditioned TSE systems. Source clips marked as known usually become exact waveform references only after downstream mixture generation.
Method
Cat.
PESQ ↑
ESTOI ↑
SI-SDR/ SI-SNR ↑
DNSMOS ↑
WER ↓
SIM ↑
Libri2Mix Noisy
TD-SpeakerBeam [ 46 ]
D
1.66
0.70
9.21
3.14
0.21
0.93
X-TF-GridNet [ 155 ]
D
1.72
0.72
9.85
3.42
0.19
0.93
SSL-MHFA [ 37 ]
D
1.76
0.74
10.60
3.22
0.17
0.94
USEF-TSE [ 55 ]
D
1.82
0.72
10.17
3.48
0.17
0.94
AD-FlowTSE [ 168 ]
G
2.15
0.81
12.69
3.48
NR
0.87
TABLE III: Catalogue of independently reported results on Libri2Mix. Values are transcribed from the cited studies and are listed to summarize the range of reported evaluations, not to establish a controlled ranking. Papers differ in training data, mixture recipes, clean or noisy test subsets, sampling rates, model scale, and metric implementations. The SI-SDR/SI-SNR column retains the signal-level measure named by each source. D, G, and H denote discriminative, generative, and hybrid discriminative-generative systems. † marks an adjacent target-sound system and ‡ marks a preprint-only system. NR indicates that the source did not report the metric.
Method
Cat.
DNSMOS ↑
NISQA ↑
SpeechBERT ↑
dWER ↓
WavLM Sim. ↑
WeSpeaker Sim. ↑
SIG
BAK
OVRL
Mixture
N/A
3.383
3.098
2.653
2.453
0.572
0.792
0.847
0.759
SpEx+ [ 58 ]
D
3.472
4.027
3.186
3.349
0.878
0.148
0.973
0.935
WeSep [ 202 ]
D
3.486
3.838
3.118
3.892
0.895
0.123
0.980
0.945
USEF-TSE [ 55 ]
D
3.555
4.051
3.272
4.319
0.935
0.075
0.988
0.968
TSELM-L [ 34 ]
G
3.489
4.041
3.212
3.961
0.793
0.297
0.887
0.627
TABLE IV: Perceptual, linguistic, and speaker-consistency results on Libri2Mix Clean reported under the LauraTSE evaluation protocol [ 35 ] . D, G, and D-G denote discriminative, generative, and discriminative-generative systems. † identifies a unified restoration system rather than a dedicated TSE model, and ‡ marks results from a preprint extension. LauraTSE evaluates dWER with Whisper-base, SpeechBERT with HuBERT-base representations, and speaker similarity with WavLM-base-plus-SV and the WeSpeaker ResNet-221LM. The LauraTSE paper reports that the AnyEnhance values were taken from the original AnyEnhance study, whereas the remaining published baseline outputs were evaluated in its comparison.