RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing
Organizations: Shanghai Normal University, Shanghai, China · Unisound AI Technology Co., Ltd., Beijing, China
Abstract
Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverages region-wise speaker similarity and silence constraints within a target-aware adaptation framework to jointly optimize the model using synthetic and weakly supervised real-world data, injecting real-world traits while preserving synthetic-learned capabilities. The second stage further adapts the model using only real-world data through our RemixIT-TSE, where quality-filtered teacher pseudo targets, which guarantee reliable student training, provide signal-level supervision via SI-SNR loss. Experiments on the real conversational evaluation set (EVAL-2) of the SLT 2026 REAL-TSE Challenge, the proposed method achieves a 6.53% relative TER reduction, together with relative improvements of 21.84% in speaker similarity, 9.89% in DNSMOS-P808, and 4.10% in target-activity F1 over the source-domain baseline, demonstrating its effectiveness under unseen real-world conditions. Source code at https: //github.com/YuWang-Speech/RemixIT-TSE.
Figures & tables
| ID | Method | TER | SIM | SIG | BAK | OVRL | P808 | Precision | Recall | F1 |
|---|---|---|---|---|---|---|---|---|---|---|
| 1A | SDP (Baseline) | 0.662 | 0.441 | 2.815 | 2.515 | 2.075 | 2.897 | 0.787 | 0.930 | 0.837 |
| 1B | SDP + SAMoM (synthetic clean) | 0.807 | 0.428 | 1.870 | 2.086 | 1.491 | 2.612 | 0.775 | 0.886 | 0.807 |
| 2B | SDP + SAMoM (real single-spk) | 0.753 | 0.499 | 1.811 | 1.668 | 1.421 | 2.726 | 0.762 | 0.965 | 0.840 |
| 1C | SDP + TAA | 0.661 | 0.459 | 2.884 | 2.822 | 2.221 | 3.103 | 0.772 | 0.949 | 0.848 |
| 2C | SDP + TAA + RTA (RemixIT-TSE) | 0.621 | 0.501 | 2.544 | 2.892 | 2.016 | 2.977 | 0.804 | 0.940 | 0.854 |
| Set | System | TER | SIM | OVRL | P808 | F1 |
|---|---|---|---|---|---|---|
| EVAL-1 | SDP | 0.726 | 0.485 | 2.049 | 2.923 | 0.824 |
| EVAL-1 | RemixIT-TSE | 0.680 | 0.533 | 2.173 | 3.088 | 0.837 |
| EVAL-2 | SDP | 0.763 | 0.335 | 1.850 | 2.710 | 0.804 |
| EVAL-2 | RemixIT-TSE | 0.713 | 0.408 | 2.040 | 2.978 | 0.837 |
| Overall | SDP | 0.748 | 0.395 | 1.929 | 2.790 | 0.812 |
| Overall | RemixIT-TSE | 0.700 | 0.458 | 2.093 | 3.022 | 0.837 |
| Configuration | Data (hrs) | TER | SIM | OVRL | P808 | F1 |
|---|---|---|---|---|---|---|
| RemixIT-TSE | 4.2 | 0.621 | 0.501 | 2.016 | 2.977 | 0.854 |
| w/ residual loss | 4.2 | 0.752 | 0.498 | 1.421 | 2.714 | 0.841 |
| w/o filtering | 51.03 | 0.703 | 0.449 | 1.450 | 2.649 | 0.792 |