cs.SDSep 25, 2026

Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information

Authors: Shuhan Zhang, Wenxuan Wu, Haizhou Li

Abstract

In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-party interference. Observing that anticipating upcoming activity from semantic, acoustic, and facial cues benefits online AV-TSE, we propose an LLM-based target-speaker voice activity projection (TS-VAP) module. Unlike conventional VAP with separated speaker channels, it forecasts the future activity of the target and conversational partner directly from the overlapping mixture, drawing on the linguistic and conversational knowledge of a speech-LLM, and uses this prediction to guide a low-latency separator. We further combine this predictive context with historical and synchronous speaker context. Experiments show that TS-VAP consistently improves streaming extraction across multiple AV-TSE backbones, and that further combining historical, synchronous, and predictive context yields nearly 1 dB gain on real AV conversations. Project page: https://jjjjiaozi.github.io/TS-VAP/.

Explore similar work

CardsList
  1. SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

    Jul 16, 2026Shuai Wang, Zihan Qian, Ke Zhang +9Target Speaker ExtractionSpeech Processing

  2. PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

    Jul 9, 2026Wanyi Ning, Wei Zhou, Yingpeng Li +3Target Speaker ExtractionSpeech Processing

  3. Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

    Jul 11, 2026Shuhai Peng, Jinjiang Liu, Hui Lu +5Target Speaker ExtractionSpeech Processing