cs.SDSep 28, 2026

RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing

Authors: Yu Wang, Haixin Guan, Shuang Wei, Yanhua Long

Organizations: Shanghai Normal University, Shanghai, China · Unisound AI Technology Co., Ltd., Beijing, China

Abstract

Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverages region-wise speaker similarity and silence constraints within a target-aware adaptation framework to jointly optimize the model using synthetic and weakly supervised real-world data, injecting real-world traits while preserving synthetic-learned capabilities. The second stage further adapts the model using only real-world data through our RemixIT-TSE, where quality-filtered teacher pseudo targets, which guarantee reliable student training, provide signal-level supervision via SI-SNR loss. Experiments on the real conversational evaluation set (EVAL-2) of the SLT 2026 REAL-TSE Challenge, the proposed method achieves a 6.53% relative TER reduction, together with relative improvements of 21.84% in speaker similarity, 9.89% in DNSMOS-P808, and 4.10% in target-activity F1 over the source-domain baseline, demonstrating its effectiveness under unseen real-world conditions. Source code at https: //github.com/YuWang-Speech/RemixIT-TSE.

Figures & tables

Explore similar work

Jul 9, 2026cs.SD

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable. We present PS4, a proxy-supervised training framework for TSE in real conversational mixtures, with two main contributions. First, we construct a large-scale corpus of 71,771 training samples derived from four public datasets, covering both Chinese and English scenarios. Each sample contains an overlapping speech mixture, per-speaker enrollment audio, a ground-truth transcript, and frame-level voice activity labels. Second, we propose a proxy-supervised joint training strategy that fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR cross-entropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. Starting from a publicly available pre-trained checkpoint, only the BSRNN separator is updated during fine-tuning. On the REAL-T challenge leaderboard, PS4 ranks 2nd overall, achieving the best speaker similarity and timing F1 among all submitted systems.
Jul 16, 2026eess.AS

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.
Jul 13, 2026cs.SD

The SonicAGI System for the REAL-TSE Challenge

Real-world target speaker extraction (TSE) remains challenging because target speech, interference, and enrollment are recorded under mismatched acoustic conditions with reverberation, noise, and irregular conversational overlap. This paper describes the SonicAGI submission to the REAL-TSE Challenge (IEEE SLT 2026). We take a data-centric approach that combines fully simulated mixtures from clean speech with real meeting overlaps, and use a frozen offline enhancer to provide a denoised mirror of real targets for auxiliary supervision. For the online track, we introduce SwiftNet-Lookahead, which inserts a single bounded-lookahead module before a strictly causal iterative separator and keeps the total system latency at 96 ms. For the offline track, we use a frame-level enrollment cross-attention USEF-TFGridNet with a magnitude-domain fusion stage that trades off perceptual quality and speaker fidelity. In the official evaluation, SwiftNet-Lookahead ranks second in Track1 and USEF-TFGridNet ranks fifth in Track2, both exceeding the challenge baselines. These results suggest that real-data-oriented training and track-specific modeling are effective for conversational TSE.