Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model's native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches 2.85× the stock runtime's speed at 38.8% lower peak memory. On the live duplex path, mean SPEAK time falls from 14% over the one-second cadence to 2% under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-
Figures & tables
Figure 1. One second of full-duplex speech synthesis and recurring operations. Conventional runtimes over-provision state at static constants and issue operations individually. DuplexCadence declares native clock rates and retention bounds to derive demand-sized state at a stable address and an exact-shape replay plan without padding. Side-by-side motivation diagram. The conventional runtime over-provisions a fixed state envelope, changes cache addresses, and issues many small GPU kernels. DuplexCadence uses a native-clock contract to derive compact stable state and exact-shape graph replay, producing identical PCM bytes with lower latency and memory.
Figure 2. Stage time breakdown for one full-duplex second. MiniCPM-o 4.5 on A100 (omni-modal duplex, 17 emitted speech tokens; stream-elapsed times). Autoregressive decoding fits within one second, whereas the synthesis tail causes a 16.9% deadline overrun (§ ). Waterfall of stream-elapsed stage times for one SPEAK unit. The LLM backbone and speech-token AR decode already fill one second; the remaining stages push the total 16.9 percent over the 1 s cadence.
Operator
Measured
Bound
γo
Flow DiT step (CFG × 2)
26.6
0.86
31 ×
LLM step ( L=1 , 8B)
∼ 50
10.3
∼ 5 ×
TTS step ( L=1 , small)
∼ 23
≪ 1
> 20 ×
Table 1. Hot operators run far above their roofline bounds. Per-invocation time in ms against the theoretical roofline bound of ( 1 ) (§ ).
Region and implementation constant
Reserved
Declared live
ρr
Flow workspace (steps × frames), duplex
16 × 1,000
5 × 458
7.0 ×
same, standalone decoder
16 × 1,000
10 × 458
3.5 ×
Attention mask envelope (frames)
500
56
8.9 ×
largest bucket
1,000
56
17.9 ×
Stacked carry per frame (fp32)
2 MiB
0.125 MiB
16 ×
Table 2. Persistent state reserved envelopes versus declared live extents. Over-reservation ratios ρr range from 3.5× to 18× across state regions. Derivations are in § .
Figure 3. DuplexCadence system architecture. Adapters declare native clocks, retention rules, and address-stable regions for demand-sized storage and a finite shape catalog. At runtime, wrapped callables enforce fail-closed admission and write back in-place into the fixed-address carry. Three-stage DuplexCadence system diagram. The declaration and capture plane derives a demand-sized region layout and per-callable graph caches. The runtime plane applies a local fail-closed gate to the upsampling encoder, the whole CFM solver loop, and the capturable vocoder core. Audio materialization follows, with an optional side-stream path disabled by default.
Region and declared policy
Wr
sr
Kr
Verdict
measured: catalogs captured in this paper
Moshi temporal KV, fixed-capacity ring
0
–
1
captured
Token2Wav carry, 100 retained frames
100
50
3
captured
Freeze-Omni codec, chunk 40 + pad 10
10
10
2
captured
computed from released configurations
Speech-token KV, windowed mode (steady)
0
–
1
out: Qo
Table 3. Catalog width across declared retention policies. Evaluation of ( 7 ) with Wr and sr read from released configurations. Regions with narrow catalog widths and stable signatures are captured, whereas large widths or variable query lengths fall outside the lossless scope.
Variant
p50
Voc.
Peak
Speedup
Δ Peak
B0
241.7
19.8
4,852
—
—
S
244.2
19.9
2,826
0.99 ×
− 41.8%
R (step)
86.3
4.8
5,194
2.80 ×
+ 7.0%
SR (step)
86.3
4.7
3,109
2.80 ×
− 35.9%
Rc (chunk)
86.6
4.9
7,943
2.79 ×
+ 63.7%
SRc (chunk)
84.9
4.7
2,972
2.85 ×
− 38.8%
Table 4. Factorial evaluation of state and compute rules. Per-chunk p50 latency, vocoder time, and peak memory on A100 ( n=4 matched blocks). Speedup and Δ Peak are relative to B0. All 20/20 optimized cells achieve bit-exact PCM.
Metric
B0
Rc
SRc
deterministic protocol
Synthesis segment (ms)
201.7
72.4
72.2
Peak allocation (MiB)
23,545
26,746
21,198
Token and PCM SHA
—
exact
exact
stock protocol (cadence)
SPEAK mean (ms)
1,141
—
976
Table 5. End-to-end full-duplex evaluation on MiniCPM-o 4.5. A100, n=6 matched blocks. Under the stock protocol (lower block, 48 units), mean SPEAK time drops below the one-second cadence (§ ).
Figure 4. Serving capacity under shared-weight multi-session concurrency. Peak memory allocation in GiB on A100. DuplexCadence consistently reduces peak allocation across session counts, enabling the speak-heavy N=4 workload to complete within ≈ 26 GiB where baseline runs out of memory (Table ). Two-panel line plot of peak GPU allocation versus session count. Mixed and speak-heavy workloads both show a widening gap between the stock baseline and DuplexCadence; speak-heavy N=4 is OOM on A100 for the baseline.
Stack and mode
Latency
Speedup
Exact
LOC
Token2Wav, exact catalog (SRc)
84.9 ms/chunk
2.85 ×
20/20
314
Moshi, vendor graph (width-1 check)
31.6 ms/frame
4.97 ×
yes
0
Moshi, ExactShapeReplay
44.3 ms/frame
3.46 ×
3/3
133
Freeze-Omni, TiCodec generator
5.0 vs. 8.6 ms
1.72 ×
3/3
126
Table 6. Cross-architecture generalization across disjoint decoders. Medians of paired speedup ratios against eager execution; Exact counts gated blocks with bit-identical outputs; LOC denotes executable adapter lines (§ ).
Configuration
p50
Peak
vs. ours
SHA
DuplexCadence SRc
84.5
2,983
—
equal
B0 (eager)
245.3
4,852
+196.8%
ref.
torch.compile , reduce-overhead mode
DiT, dynamic=False
122.3
4,852
+46.2%
differs
DiT, dynamic default
123.1
5,048
+46.5%
differs
DiT, max-autotune a
167.0
5,259
+102.3%
differs
Table 7. Comparison against compiler and vendor-graph controls. NVIDIA A100, n=4 matched blocks. Latency p50 in ms, peak memory in MiB, and gap relative to DuplexCadence SRc (positive means slower). No control achieves speedup, memory reduction, and bitwise exactness together.
Key Laboratory of Intelligent Information Processing Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS) · Key Laboratory of AI Safety, Chinese Academy of Sciences · University of Chinese Academy of Sciences, Beijing, China