cs.SDSep 24, 2026

AdaptDuplex: from static to adaptive full-duplex spoken dialogue

Authors: Zhiyang Zhou, Yingxin Shang, Zhou Wang, Hongwei Cai, Weixu Wang, Shuran Zhou, Shuofeng Zhao, Wenke Fan, +4 more

Abstract

Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which upgrades Qwen3-Omni with such a mechanism, co-designed across three layers. A compact token-level protocol represents every window as a canonical sequence, trains dual-stream alignment through a bounded text lead over speech, and exposes every behavioral decision as an explicit token for training-free runtime control via logits bias. Adaptive mechanisms dynamically predict among discrete window durations and augment direct response as needed with non-blocking cognitive consolidation and multi-flight external reasoning. A progressive pipeline introduces these behaviors through a three-stage Thinker curriculum, then Talker-only and joint SFT, with GRPO as a further increment. On Full-Duplex-Bench v1 and v1.5, AdaptDuplex outperforms DuplexOmni and MiniCPM-o 4.5 on the majority of comparable turn-taking, overlap-behavior, and timing metrics, with gains in both interaction decisions and response timing. On the human-recorded HumDial-FDBench, it attains the highest Final score (72.9) of the compared duplex models.

Explore similar work

CardsList
  1. SteerDuplex: Steerable Duplex Speech Dialogue Models

    Sep 14, 2026Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai +13Full-Duplex Speech ModelsFull-Duplex Dialogue

  2. DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

    Jul 28, 2026Takyoung Kim, Kang-wook Kim, Sang Hoon Woo +3Turn-TakingDialogue