cs.SDOct 8, 2026

DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion

Authors: Heeseung Kim

Organizations: Department of Artificial Intelligence University of Seoul Seoul, Republic of Korea

Abstract

Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame. We introduce DiffuPlex, a rolling masked diffusion framework that reduces this sequential computation by predicting multiple future user and assistant frames in a single backbone wake. DiffuPlex consumes only a confident prefix of each predicted future while interaction continues at the original frame rate. As user speech arrives, it checks the corresponding user predictions and, when the interaction diverges, preserves already played assistant content while revising only the unplayed future. We consider two inference policies over the same predictor: DiffuPlex-LISTEN consumes multiple future frames when they predict assistant silence, whereas DiffuPlex-SPEAK can also consume predicted assistant speech. Across full-duplex interaction and spoken-language evaluations, DiffuPlex substantially reduces sequential backbone computation while largely preserving interaction behavior and general capability. DiffuPlex-LISTEN and DiffuPlex-SPEAK achieve 1.46×1.46\times and 1.59×1.59\times deployment-path wall-clock speedups and 1.61×1.61\times and 1.80×1.80\times Core LM speedups, with all measured backbone invocations completing within the 80ms interaction interval. Human evaluation shows that LISTEN preserves speech naturalness and conversational quality, while SPEAK retains conversational quality with some degradation in speech naturalness.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AdaptDuplex: from static to adaptive full-duplex spoken dialogue

    Sep 24, 2026Zhiyang Zhou, Yingxin Shang, Zhou Wang +9Spoken Dialogue SystemsSpeech Language Models

  2. HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models

    Oct 6, 2026Kyudan Jung, Hyunsin Park, Yoonhyung Lee +5RL for Language ModelsSpeech Language Models

  3. BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

    Jun 12, 2026Qingkai Fang, Shoutao Guo, Yang FengSpoken Dialogue SystemsSpeech Language Models