The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Authors: Kai Li, Wenze Ren, Junjie Li, Cheng Yu, Peijun Yang, Chien-yu Huang, Haibin Wu, Szu-Wei Fu, +6 more
Organizations: 1Tsinghua University · 2National Taiwan University · 3The Hong Kong Polytechnic University · 4Ohio State University · 5Wuhan University · 6Carnegie Mellon University · 7Meta · 8NVIDIA · 9Nagoya University · 10Academia Sinica · 11The Chinese University of Hong Kong, Shenzhen
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge evaluates two related settings. Track1 comprises two scenarios: real-world mixtures recorded with two speakers speaking simultaneously, without a corresponding clean reference signal, and synthetic remixes obtained by manually mixing the separately recorded speech of two speakers, with a clean reference signal available; Track2 reuses audio but pairs it with a degraded target video and contains additional 3-m far-field recordings. The speakers in the development and test sets are disjoint. Evaluation metrics include clean-waveform fidelity, learned quality estimates, transcription accuracy, and speaker identification. In the remix task on the development set, the baseline model achieved an SI-SDR of −4.069dB and an STOI of 0.388 on Track1, and an SI-SDR of −2.851dB and an STOI of 0.470 on Track2. We release the AV-ConvTasNet checkpoints, the offline evaluator, and the official baseline results on the development and test sets.