Full-duplex spoken dialogue models listen and speak at the same time, enabling voice agents to have natural, low-latency interactions that turn-based systems cannot offer. However, they are commonly evaluated against single-sided interlocutors: pre-recorded audio that cannot react, or an automated examiner that reacts in real time but only administers a fixed sequence of tests and is never graded. These single-sided frameworks evaluate only half of a two-body problem, where turn-taking, overlap, and interruption are joint products of two coupled speakers. We propose DyaFDB, a framework that evaluates full-duplex models in a dyadic setup: two models converse directly under assigned roles with cooperative or conflicting goals, and both sides are scored offline with an external judge. DyaFDB probes how the two models behave toward each other, such as how they take turns or carry an assigned role under different interests. We instantiate four tasks as 140 scenarios and record 7,560 conversations, covering six self- and cross-play pairings. Throughout the experiments, we observe that how a model behaves continually reshapes its partner. We thus demonstrate that each model must be both the examiner and examinee of the other, and no single fixed interlocutor can play both parts. We will release the scenarios, role prompts, and recording protocols between two full-duplex models, without any pre-recorded audio.
Figures & tables
Figure 1: Evaluation framework for a full-duplex model. (a) Stimulus-response: the model reacts to pre-recorded audio that cannot react back. (b) Automated examiner: the examiner reacts to the model in real time but is bound to asking a set of tests and is never graded. (c) DyaFDB (ours) : two full-duplex models converse on equal standing over the model-to-model bridge. Each model may hold a private role in a shared situation, and both sides are scored offline from the recording.
Benchmark
Interlocutor
Multi-turn
Reactive
Both graded
Partner varied
FDB-v1, -v1.5 ( Lin et al., 2025 ; Lin et al., 2026c )
Recorded audio
×
×
×
×
FDB-v3 ( Lin et al., 2026a )
Recorded human
×
×
×
×
MTR-DuplexBench ( He et al., 2026b )
TTS script
✓
× †
×
×
τ -Voice ( Ray et al., 2026 )
LLM + TTS
✓
✓
×
×
FDB-v2 ( Lin et al., 2026b )
SLM examiner
✓
✓
×
× ‡
DyaFDB (Ours)
Full-duplex model
✓
✓
✓
✓
Table 1: DyaFDB’s position among other benchmarks. Prior benchmarks hold the interlocutor constant, none of them grading both sides of a coupled conversation. Reactive : interlocutor’s outputs depend on what the model actually said. † MTR-DuplexBench replaces the model’s channel with ground-truth speech for all previous turns, so the partner cannot react to it. Partner varied : the partner model is a design factor, so its contribution to a score is also measured. ‡ FDB-v2’s examiner is itself a real-time speech model, but it is a single model that follows a fixed sequence of tests.
Task
ID
Archetype
Roles
Score ∗
Coordination
Role adherence : each model holds an assigned role in a shared situation
T1-1
Engaging partner
symm.
judged adherence
T1-2
Digressing partner
asymm.
judged adherence
Information pooling : each model has partial knowledge that must be pooled for answer
T2
-
symm.
problem solving
Conflict
Contested channel : who holds and dominates the speech channel
T3-1
Competing partner
symm.
channel occupancy
Table 2: Task suite of DyaFDB. Coordination tasks are cooperative, while Conflict tasks require two models to pursue incompatible outcomes. Roles are symmetric when both sides hold the same kind of instruction; in asymmetric types the two roles differ, so the measure is reported from one role’s perspective. See Appendix B for the scenario design and examples.
Figure 2: Turn-taking metrics ( left ) and task-score variance ( right ) across pairings. Above each matrix, the proportion of variance explained is given, η2=model∣partner∣model×partner , and the matrices are ordered from the largest model effect to the largest partner effect. The same decomposition of the task-wise scores is given on the right. See Appendix C for further details.
Role adherence (T1)
Information pooling (T2)
Engaging partner (T1-1)
Digressing partner (T1-2)
Teammate
Model
PP
MCPM
Raon
Avg
PP
MCPM
Raon
Avg
PP
MCPM
Raon
Avg
PersonaPlex
0.44
0.55
0.55
0.51
0.23
0.29
0.35
0.29
0.05
0.07
0.23
0.12
MiniCPM-o
0.64
0.60
0.61
0.62
0.62
0.52
0.62
0.59
0.07
0.14
0.30
0.17
Raon
0.69
0.64
0.72
0.68
0.86
0.87
0.77
0.83
0.23
0.30
0.35
0.29
Table 3: Role adherence (T1) and Information pooling (T2). For role adherence, scores of the model beside the partner under each archetype, engaging or digressing, are reported. For information pooling, scores of the row-column team are reported, where the score is the solved rate; the matrix is symmetric by construction. The diagonal is self-play. PP: PersonaPlex, MCPM: MiniCPM-o.
Figure 3: Timing when models enter or leave the assigned roles (T1-1, engaging partner) . Each curve is the cumulative runs in which the model has entered its role (entry) or drifted off it after entering (deviation) by a given time. Vertical lines mark the Kaplan-Meier (KM) median ( Kaplan and Meier, 1958 ) , the time by which the event has occurred in half of the runs, while n.r. indicates that a KM median has not reached within 60 s. (a) and (b) show the three self-play pairings, entry and deviation, and (c)–(e) show the entry in each cross-play pairing.
Competing partner (T3-1)
Avoiding partner (T3-2)
Model (Presser)
PP
MCPM
Raon
Avg
PP
MCPM
Raon
Avg
PersonaPlex
(0.50)
0.48
0.38
0.43
0.52
0.56
0.45
0.51
MiniCPM-o
0.52
(0.50)
0.25
0.39
0.45
0.56
0.21
0.41
Raon
0.62
0.75
(0.50)
0.69
0.62
0.75
0.52
0.63
Avg (avoider) ↑
0.53
0.62
0.39
Table 4: Contested channel (T3). Channel occupancy of the row (pressing) model against the column partner, where the competing self-plays are pinned at 0.50. The bottom row averages each avoider column, and higher means that model as avoider yielded the floor.
Preference (T4-1)
Attack (T4-2)
Model
PP
MCPM
Raon
Avg
PP
MCPM
Raon
Avg
PersonaPlex
0.24
0.12
0.11
0.16
0.46
0.30
0.63
0.46
MiniCPM-o
0.27
0.12
0.05
0.14
0.36
0.30
0.41
0.36
Raon
0.45
0.20
0.16
0.27
0.31
0.28
0.39
0.33
Avg (defender) ↓
0.38
0.29
0.48
Table 5: Contested goal (T4). Score of the row model against the column partner in both archetypes. Preference : the rate at which the row model’s option became the pair’s choice. Attack : the rate at which the row model extracted the secret from the column defender. The bottom row averages each defender column, and lower means that model defended its secret better.
Figure 4: Scores split into acting and succeeding. x-axis is the rate of the scored action (decided or asked) and y-axis is the score when the model acts, so the gray curves connect equal products. Small dots indicate per-partner runs. Open circles in T4-2 add the leaks that occurred without a demand.
Figure 5: Scores across conversation windows of 30, 60, and 120 seconds. The top row is the action step (entered, delivered, or asked) and the bottom row is the score, with 95% confidence intervals.
Figure 6: Balance of the contested games. The left end of each track is the side that gives something up (permission, secret, money) and the right end the side that obtains it. The marker is the win rate of the side on the right and the small dots are the per-model win rates in that role. For T3-2 and T4-2 every valid run counts and a run where nothing happens is a win for the avoider or defender, while for T4-3 and T4-4 the balance is conditional on a deal, since a no-deal is a loss for both sides.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Measure
Task
Model
Partner
Model × Partner
Residual
TORpause
0.02 ± 0.01
0.74 ± 0.03
0.05 ± 0.01
0.07 ± 0.01
0.12 ± 0.02
Backchannel
0.01 ± 0.00
0.49 ± 0.02
0.12 ± 0.01
0.32 ± 0.02
0.07 ± 0.01
Barge-in success
0.01 ± 0.01
0.29 ± 0.03
0.49 ± 0.04
0.05 ± 0.02
0.17 ± 0.03
TORend
0.01 ± 0.00
0.19 ± 0.01
0.65 ± 0.01
0.12 ± 0.01
0.04 ± 0.00
Appendix
Table 6: Variance decomposition of the turn-taking metrics. The fraction of the variance explained by each term ( η2 ), with its bootstrap standard deviation, is shown. The residual holds the interactions that involve the task.
Task ID
Measure
Question to the judge
Labels
T1
role adherence
did the model behave as its own assigned role
yes / partly / no
T1
role timing
first in-role and first off-role utterance
seconds or none
T2
decided, solved (final choice)
which candidate is left standing as the pair’s choice
candidate or none
T3-2
asked
did the presser voice its request
yes / no
T3-2
agreement
did the avoider grant the request
yes / no
T4-1
decided, won (landed option)
which option did the pair land on
option or none
Appendix
Table 7: Judge rubric. From a choice label, decided is whether the label is not none, and solved or won is whether it matches the answer or the model’s own option.
Task ID
Measure
n
Opus-5 vs. Sonnet-5 (%)
T1-1
role adherence (per seat)
2,160
86.6 ( ρ=0.75 )
T1-2
role adherence (per seat)
1,080
93.3 ( ρ=0.86 )
T2
decided (per run)
1,080
92.4
T2
final choice (per decided run)
541
95.2
T3-2
asked (per run)
1,080
86.2
T3-2
agreement (per run)
1,080
89.9
Appendix
Table 8: Agreement with Sonnet-5 and human evaluators. Top: every recording is judged again by Sonnet-5 with the same prompts and rubric. Bottom: T1 samples are judged by majority voting of human experts. EM (exact match) is the percentage of the n scored units, runs or seats, on which the two sides assign the same binary score. ρ is Spearman’s rank correlation over the three-level label (no < partly < yes), given only for the role adherence.
Model
T1-1
T2
T3-2
T4-2
PersonaPlex ( Roy et al., 2026 )
0.51
0.12
0.09
0.46
MiniCPM-o 4.5 ( Cui et al., 2026 )
0.62
0.17
0.19
0.36
Raon-SpeechChat ( Kim et al., 2026 )
0.68
0.29
0.17
0.33
Qwen2.5-7B-Instruct ( Qwen, 2025 )
0.82
0.55
0.55
0.53
Qwen3-8B ( Yang et al., 2025 )
0.88
0.55
0.50
0.68
Appendix
Table 9: Task performance of the text-only model pairing. Score of each model on the four tasks, where speech rows average each model over its three pairings and text rows come from one pairing. T2 is scored per pair, so the two text rows share one value. Here, the T3-2 score is the presser’s agreement rate, whether the avoider granted its request, not the channel occupancy.
Ref. Task
Behavior
PersonaPlex
MiniCPM-o 4.5
Raon-SpeechChat
T1
Does the partner’s stance matter?
yes
yes
yes
Enters its role?
slow
fast
fast
Stays in it?
yes
yes
no
Holds the role
0.51
0.62
0.68
T2
Discloses what it holds
0.37
0.34
0.47
Solves once decided
0.22
0.40
0.40
Appendix
Table 10: Per-model behavioral profiles. Each answer summarizes the model over its three pairings, averaging the self-play and the two cross-plays.
Claim (T4-3)
Negotiation (T4-4)
Model
PP
MCPM
Raon
Avg
PP
MCPM
Raon
Avg
PersonaPlex
0.38 (n=68)
0.33 (n=49)
0.26 (n=54)
0.32
0.74 (n=39)
0.54 (n=31)
0.48 (n=63)
0.59
MiniCPM-o
0.76 (n=49)
0.54 (n=41)
0.52 (n=21)
0.61
0.71 (n=34)
0.66 (n=8)
0.44 (n=31)
0.60
Raon
0.89 (n=65)
0.85 (n=40)
0.56 (n=36)
0.77
0.66 (n=78)
0.79 (n=32)
0.47 (n=70)
0.64
Avg (payer / buyer) ↓
0.68
0.57
0.45
0.70
0.66
0.46
Appendix
Table 11: Claim and negotiation tasks. Score of the row model, the side that receives the money (claimant, seller), against the column model, the side that pays (payer, buyer), among decided runs (n) . The score is the share of decided runs that landed on the claimant’s line, or the seller’s share of the gap between the two openings.
Figure 8: Turn-taking metrics and variance decomposition over four models. Figure 2 recomputed with Lychee-FD as a fourth model and partner, with the task-score panel taken from the 4×4 matrices of Tables 12 – 14 .
Role adherence (T1)
Information pooling (T2)
Engaging partner
Digressing partner
Teammate
Model
PP
MCPM
Raon
LFD
Avg
PP
MCPM
Raon
LFD
Avg
PP
MCPM
Raon
LFD
Avg
PersonaPlex
0.44
0.55
0.55
0.51
0.51
0.23
0.29
0.35
0.32
0.30
0.05
0.07
0.23
0.19
0.13
MiniCPM-o
0.64
0.60
0.61
0.56
0.60
0.62
0.52
0.62
0.61
0.59
0.07
0.14
0.30
0.20
0.18
Raon
0.69
0.64
0.72
0.62
0.67
0.86
0.87
0.77
0.78
0.82
0.23
0.30
0.35
0.28
0.29
Lychee-FD
0.46
0.53
0.37
0.40
0.44
0.45
0.33
0.33
0.33
0.36
0.19
0.20
0.28
0.15
0.20
Appendix
Table 12: Role adherence (T1) and information pooling (T2) with Lychee-FD. Table 3 extended by a fourth model, appended to the last row and column. Avg spans the four pairings. PP: PersonaPlex, MCPM: MiniCPM-o, LFD: Lychee-FD.
Competing partner
Avoiding partner
Model (Presser)
PP
MCPM
Raon
LFD
Avg
PP
MCPM
Raon
LFD
Avg
PersonaPlex
(0.50)
0.48
0.38
0.36
0.41
0.52
0.56
0.45
0.44
0.49
MiniCPM-o
0.52
(0.50)
0.25
0.28
0.35
0.45
0.56
0.21
0.35
0.39
Raon
0.62
0.75
(0.50)
0.49
0.62
0.62
0.75
0.52
0.54
0.61
Lychee-FD
0.64
0.72
0.51
(0.50)
0.62
0.63
0.72
0.51
0.53
0.60
Avg (avoider) ↑
0.56
0.65
0.42
0.46
Appendix
Table 13: Contested channel (T3) with Lychee-FD. Table 4 extended by a fourth model, appended to the last row and column. Avg spans the four pairings, and the bottom row averages each avoider column, where higher means that model as avoider yielded the floor.
Preference
Attack (column: defender)
Model
PP
MCPM
Raon
LFD
Avg
PP
MCPM
Raon
LFD
Avg
PersonaPlex
0.24
0.12
0.11
0.27
0.18
0.46
0.30
0.63
0.14
0.38
MiniCPM-o
0.27
0.12
0.05
0.44
0.22
0.36
0.30
0.41
0.20
0.32
Raon
0.45
0.20
0.16
0.48
0.32
0.31
0.28
0.39
0.13
0.28
Lychee-FD
0.23
0.05
0.05
0.11
0.11
0.35
0.28
0.51
0.19
0.33
Avg (defender) ↓
0.37
0.29
0.48
0.16
Appendix
Table 14: Contested goal (T4) with Lychee-FD. Table 5 extended by a fourth model, appended to the last row and column. Avg spans the four pairings, and the bottom row averages each defender column, where lower means that model defended its secret better.
Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. In addition, existing benchmarks provide limited coverage of languages and dialogue domains. We propose M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs. M3-DuplexBench supports English and Japanese and covers both casual conversation and multi-turn question answering. In addition, we evaluate models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior. Experiments with recent FDSDSs reveal model-specific turn-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context.
Full-duplex spoken dialogue models (SDMs) can listen and speak simultaneously, enabling interaction dynamics closer to human conversation than turn-based systems. Inspired by neural coupling in human communication, we study how such models coordinate their internal representations during interaction. We simulate full-duplex dialogues between two instances of the pretrained \textit{Moshi} model under controlled conditions, manipulating channel noise and decoding bias. Synchronization is measured using Centered Kernel Alignment (CKA) across temporal lags, while anticipatory turn-taking cues are probed from delayed internal activations using causal LSTM models, from both speaker and listener perspectives. We find strong representational synchronization under no noise conditions, peaking near zero lag and degrading with noise, and we show that internal states encode anticipatory information that supports turn-taking prediction ahead of time.
Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.