cs.AISep 28, 2026

DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines

Authors: Haixiao Gao, Yimin Zheng, Linyou Xiao, Zeke Xie

Organizations: Jinan University, China · Hong Kong University of Science and Technology (Guangzhou), China

Abstract

Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model's native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches 2.85×2.85\times the stock runtime's speed at 38.8%38.8\% lower peak memory. On the live duplex path, mean SPEAK time falls from 14%14\% over the one-second cadence to 2%2\% under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-

Figures & tables

Explore similar work

CardsList
  1. AdaptDuplex: from static to adaptive full-duplex spoken dialogue

    Sep 24, 2026Zhiyang Zhou, Yingxin Shang, Zhou Wang +9Full-Duplex Speech ModelsDialogue

  2. BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

    Jun 12, 2026Qingkai Fang, Shoutao Guo, Yang FengFull-Duplex Speech ModelsSpeech Language Models

  3. TASTE2: Text-Aligned Speech Modeling and Deployment toward Full-Duplex Voice Interaction

    Sep 8, 2026Yi-Chang Chen, Chun Wei Chen, Dien-Ruei Wu +7Full-Duplex Speech ModelsSpeech Synthesis