A Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations
Authors: Matthew Sun, Vinay Kothapally, Meng Yu, Chao Huang, Hao Zhang, Yixuan Zhang, Steve Yves
Organizations: Tencent Americas, Palo Alto, CA, USA · Tencent Americas, Bellevue, WA, USA · School of Electronic Information, Wuhan University, Wuhan, Hubei, China
Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited control over these events and limited labels for their intent. We present a pipeline for synthesizing intent-labeled, two-channel conversational speech from relational event lists. An LLM authors each event's speaker, text, conversational act, and attachment to an earlier event without predicting absolute timestamps. Events are synthesized independently, aligned with their source text, and placed on a shared clock, so turn-taking landmarks are measured from the rendered signal while silence durations are specified or sampled from turn-taking distributions. The pipeline covers 42 phenomena across eight families in English and Mandarin, derives frame-level system actions from authored intent, and promotes diversity using small, diverse sets of prior examples and batch prompts that request alternatives with self-reported probabilities. Ablations show gains in each targeted diversity dimension. On a four-action label space for taking, holding, releasing, and not holding the conversational floor, a semantic voice-activity detector using only current and past audio reaches start-speaking and start-listening F1 scores of 0.819 and 0.802. When generating its own responses, the full-duplex speech model Moshi takes 0.85 of the reference turns after fine-tuning on the generated corpus, compared with 0.44 before fine-tuning. Its frame-level precision for predicting system-floor occupancy rises from 0.46 to 0.88. With reference context at each step, its frame-level floor F1 rises from 0.893 to 0.962. These results show that controlled synthesis can provide learnable and transferable supervision for full-duplex turn management.
Figures & tables
Structural metric
No hub
Blend
Δ
Barge-ins (total)
0
46
–
Truncations (total)
0
42
–
Backchannels (total)
48
86
+79.2%
Feature-space Vendi ↑
4.29
7.05
+64.2%
NN mean distance ↑
1.37
2.66
+95.0%
Act-seq. ROUGE-L ↓
0.861
0.769
-10.7%
TABLE I: Structural-hub ablation on AF-02. Blend combines feature and act-sequence distances. Arrows indicate the direction of better performance.
Axis
Metric
Seeded
+Hub
Δ
Semantic
RBF-Vendi ↑
9.402
9.738
+3.6%
Lexical
Unique types ↑
1449
1732
+19.5%
Lexical
Pairwise ROUGE-L ↓
0.141
0.150
+6.0%
Lexical
EAD-1 ↑
0.681
0.780
+14.4%
Lexical
EAD-2 ↑
1.060
1.079
+1.7%
Lexical
EAD-3 ↑
1.208
1.241
+2.7%
TABLE II: Semantic-hub ablation on AF-02 with identical topic seeding in both arms.
Axis
Metric
Single
Batch
Δ
Lexical
Unique types ↑
631
874
+38.5%
Lexical
Pairwise ROUGE-L ↓
0.280
0.174
-37.8%
Lexical
EAD-2 ↑
0.724
0.996
+37.6%
Semantic
RBF-Vendi ↑
2.262
3.815
+68.7%
Structural
Feature-space Vendi ↑
3.442
3.776
+9.7%
Structural
NN mean distance ↑
0.704
0.607
-13.8%
TABLE III: Probability-verbalization ablation on TT-01 with one fixed topic and no diversity hubs.
Core metric
Clean
Noisy
Δ
Accuracy
0.9932
0.9929
-0.0003
Speaking F1
0.9953
0.9951
-0.0002
Listening F1
0.9955
0.9953
-0.0002
Start-speaking F1
0.8186
0.8163
-0.0023
Start-listening F1
0.8019
0.7958
-0.0061
TABLE IV: Exact held-out results for the 4-token core semantic-VAD runs.
Merged metric
Clean
Noisy
Δ
Accuracy
0.9613
0.9525
-0.0088
Speaking F1
0.9840
0.9822
-0.0018
Listening F1
0.9860
0.9816
-0.0044
Contested-floor F1
0.9177
0.8901
-0.0276
Hold-backchannel F1
0.9061
0.8568
-0.0493
System-backchannel F1
0.8888
0.8766
-0.0123
TABLE V: Exact held-out results for the 13-token merged semantic-VAD runs.
Mode
Metric
Pretrained
Fine-tuned
Free
Turns taken
0.44
0.85
Free
Floor precision
0.46
0.88
Free
SS F1, ±3 frames
0.066
0.145
Free
SL F1, ±3 frames
0.037
0.122
Teacher
Floor F1
0.893
0.962
Teacher
Exact SL F1
0.60
0.93
TABLE VI: Moshi turn-taking performance before and after 2,000 fine-tuning steps.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 6: Three label granularities for OV-02, a correction barge-in. Each panel shows the same human and system waveforms while increasing the detail of the system-side label taxonomy.
Fig. 7: Three label granularities for OV-03, a cooperative-overlap conversation with backchannels. The finer views distinguish additional floor states that collapse into the four core actions.
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.
Takyoung Kim, Kang-wook Kim, Sang Hoon Woo +3
University of Illinois Urbana-Champaign · Seoul National University · University of California, Berkeley +2
We present DuplexDrama, the first synthesized spoken dialogue dataset that simultaneously covers four dimensions: (i) complete persona and scenario settings; (ii) three full-duplex behaviors (interruption, backchannel, incomplete); (iii) expressive speech with persona-aligned emotion labels; and (iv) script-aware sound events. DuplexDrama is built via a 4-stage pipeline; quality validation on both scripts and synthesized audio confirms its quality. We have produced more than 2,000 hours audio data with a 64-voice timbre pool spanning 13 personas and 5 age buckets; 3.8% of all turns carry at least one full-duplex behavior. This data has been validated through internal full-duplex model training. We will release a curated subset of 6,400 bilingual dialogues (800 h, Chinese ~500 h + English ~300 h) to advance full-duplex spoken dialogue model research. Data samples are available at our demo page and LLM-judge evaluation prompts will be released with the dataset.
Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open, reproducible full-duplex spoken dialogue system for Hindi by adapting Moshi, a state-of-the-art duplex speech architecture, using a custom Hindi tokeniser and training on 26,000 hours of real spontaneous conversations collected from 14,695 speakers with separate speaker channels, enabling direct learning of turn-taking and overlap patterns from natural interactions. To support Hindi text generation, we replace the original English tokeniser and reinitialise text-vocabulary-dependent parameters while retaining the pre-trained audio components. We propose a two-stage training recipe -- large-scale pre-training followed by fine-tuning on 1,000 hours of conversational data. Evaluation through the prompted dialogue continuation paradigm with both automatic metrics and human judgments demonstrates that the resulting model generates natural and meaningful full-duplex conversational behaviour in Hindi. This work serves as a first step toward real-time duplex spoken dialogue systems for Hindi and other Indian languages.