cs.SDSep 24, 2026

Exploring a Single Autoregressive LLM for Unified Target Speech Extraction across Synchronous and Asynchronous Cues

Authors: Wenxuan Wu, Shuhan Zhang, Shuai Wang, Haizhou Li

Abstract

Target speech extraction (TSE) typically trains a separate extractor per cue, and visual-cue systems often need corruption-matched training to remain robust under visual frame corruption. We show that one autoregressive LLM backbone, TSE-Omni, can serve both temporally synchronous cues (lip movements, co-speech gestures) and asynchronous cues (enrollment audio, text). TSE-Omni is driven by next-token prediction: each step predicts target speech semantic tokens from its own past outputs, which we term self-enrollment, forming a continuous target-speech context initialized by the enrollment cue (asynchronous audio or text, or a short visual prefix). This enables audio-visual compensation: the model uses synchronized visuals when intact and its token history when visual frames are missing. Under clean visuals, TSE-Omni matches strong discriminative and generative baselines (SpeechBERTScore 0.81 on VoxCeleb2 and 0.89 on LRS3 zero-shot) with higher DNSMOS. On the same VoxCeleb2 test set, after a 2 s clean visual start, removing the remaining visual frames leaves SpeechBERTScore at 0.81. It remains usable under sparse overlap and multi-speaker interference, and supports streaming inference. Project page: https://alexwxwu.github.io/tseomni-main/.

Figures & tables

Explore similar work

CardsList
  1. Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities

    Sep 28, 2026Xinyuan Qian, Yanghao Zhou, Ziyang Jiang +10Target Speaker ExtractionSpeaker

  2. StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model

    Apr 21, 2026Shuhai Peng, Hui Lu, Jinjiang Liu +8Target Speaker ExtractionAutoregressive Text-To-Speech

  3. UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

    Oct 23, 2025Haoyin Yan, Chengwei Liu, Shaofei Xue +4Speech EnhancementAutoregressive Model