A Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations
Authors: Matthew Sun, Vinay Kothapally, Meng Yu, Chao Huang, Hao Zhang, Yixuan Zhang, Steve Yves
Organizations: Tencent Americas, Palo Alto, CA, USA · Tencent Americas, Bellevue, WA, USA · School of Electronic Information, Wuhan University, Wuhan, Hubei, China
Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited control over these events and limited labels for their intent. We present a pipeline for synthesizing intent-labeled, two-channel conversational speech from relational event lists. An LLM authors each event's speaker, text, conversational act, and attachment to an earlier event without predicting absolute timestamps. Events are synthesized independently, aligned with their source text, and placed on a shared clock, so turn-taking landmarks are measured from the rendered signal while silence durations are specified or sampled from turn-taking distributions. The pipeline covers 42 phenomena across eight families in English and Mandarin, derives frame-level system actions from authored intent, and promotes diversity using small, diverse sets of prior examples and batch prompts that request alternatives with self-reported probabilities. Ablations show gains in each targeted diversity dimension. On a four-action label space for taking, holding, releasing, and not holding the conversational floor, a semantic voice-activity detector using only current and past audio reaches start-speaking and start-listening F1 scores of 0.819 and 0.802. When generating its own responses, the full-duplex speech model Moshi takes 0.85 of the reference turns after fine-tuning on the generated corpus, compared with 0.44 before fine-tuning. Its frame-level precision for predicting system-floor occupancy rises from 0.46 to 0.88. With reference context at each step, its frame-level floor F1 rises from 0.893 to 0.962. These results show that controlled synthesis can provide learnable and transferable supervision for full-duplex turn management.
Figures & tables
Structural metric
No hub
Blend
Δ
Barge-ins (total)
0
46
–
Truncations (total)
0
42
–
Backchannels (total)
48
86
+79.2%
Feature-space Vendi ↑
4.29
7.05
+64.2%
NN mean distance ↑
1.37
2.66
+95.0%
Act-seq. ROUGE-L ↓
0.861
0.769
-10.7%
TABLE I: Structural-hub ablation on AF-02. Blend combines feature and act-sequence distances. Arrows indicate the direction of better performance.
Axis
Metric
Seeded
+Hub
Δ
Semantic
RBF-Vendi ↑
9.402
9.738
+3.6%
Lexical
Unique types ↑
1449
1732
+19.5%
Lexical
Pairwise ROUGE-L ↓
0.141
0.150
+6.0%
Lexical
EAD-1 ↑
0.681
0.780
+14.4%
Lexical
EAD-2 ↑
1.060
1.079
+1.7%
Lexical
EAD-3 ↑
1.208
1.241
+2.7%
TABLE II: Semantic-hub ablation on AF-02 with identical topic seeding in both arms.
Axis
Metric
Single
Batch
Δ
Lexical
Unique types ↑
631
874
+38.5%
Lexical
Pairwise ROUGE-L ↓
0.280
0.174
-37.8%
Lexical
EAD-2 ↑
0.724
0.996
+37.6%
Semantic
RBF-Vendi ↑
2.262
3.815
+68.7%
Structural
Feature-space Vendi ↑
3.442
3.776
+9.7%
Structural
NN mean distance ↑
0.704
0.607
-13.8%
TABLE III: Probability-verbalization ablation on TT-01 with one fixed topic and no diversity hubs.
Core metric
Clean
Noisy
Δ
Accuracy
0.9932
0.9929
-0.0003
Speaking F1
0.9953
0.9951
-0.0002
Listening F1
0.9955
0.9953
-0.0002
Start-speaking F1
0.8186
0.8163
-0.0023
Start-listening F1
0.8019
0.7958
-0.0061
TABLE IV: Exact held-out results for the 4-token core semantic-VAD runs.
Merged metric
Clean
Noisy
Δ
Accuracy
0.9613
0.9525
-0.0088
Speaking F1
0.9840
0.9822
-0.0018
Listening F1
0.9860
0.9816
-0.0044
Contested-floor F1
0.9177
0.8901
-0.0276
Hold-backchannel F1
0.9061
0.8568
-0.0493
System-backchannel F1
0.8888
0.8766
-0.0123
TABLE V: Exact held-out results for the 13-token merged semantic-VAD runs.
Mode
Metric
Pretrained
Fine-tuned
Free
Turns taken
0.44
0.85
Free
Floor precision
0.46
0.88
Free
SS F1, ±3 frames
0.066
0.145
Free
SL F1, ±3 frames
0.037
0.122
Teacher
Floor F1
0.893
0.962
Teacher
Exact SL F1
0.60
0.93
TABLE VI: Moshi turn-taking performance before and after 2,000 fine-tuning steps.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 6: Three label granularities for OV-02, a correction barge-in. Each panel shows the same human and system waveforms while increasing the detail of the system-side label taxonomy.
Fig. 7: Three label granularities for OV-03, a cooperative-overlap conversation with backchannels. The finer views distinguish additional floor states that collapse into the four core actions.