cs.SDJun 4, 2026

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems

Authors: Tao ZhongJiajun DengNikita KuzminYinke ZhuTianxiang CaoTristan TsoiZhili TanSimon Lui+1 more

Organizations: The Chinese University of Hong Kong, China · AudioLab Hong Kong, Huawei Leibniz Research Center, China · Nanyang Technological University, Singapore

Abstract

Full-duplex spoken dialogue models allow voice agents to listen and speak concurrently, enabling natural interaction with real-time overlap. However, end-to-end dual-channel models that jointly encode user and agent streams may degrade in realistic acoustic environments: interfering speakers leaking into the user microphone can be encoded as part of the user query, corrupting the LLM's conditioning and causing unstable turn-taking and reduced response quality. We propose Interference-Resilient Adaptive Fusion (IRAF), a lightweight, streaming-compatible module that modulates the contribution of user audio to the LLM frame by frame. IRAF predicts a scalar reliability gate from target-speaker and user audio embeddings and rescales user representations before fusion with agent embeddings. Experiments on MS-MARCO and InstructS2S-200K show consistent gains in response quality and full-duplex interaction under interfering-speaker conditions.

Explore similar work

CardsList
  1. Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

    Jun 9, 2026Atsumoto Ohashi, Neil Zeghidour, Alexandre Défossez +1Full-Duplex Speech Models