Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation. A duplex model supports continuous listening and speaking, but complex reasoning and tool use may exceed its capabilities. A coding agent can plan and execute extended tasks, but its sequential interface is a poor fit for live conversation. Combining them requires a harness that coordinates task acceptance, progress, cancellation, replacement, and result delivery while keeping the conversation responsive. Existing harnesses often rely on coupled heuristics, making them difficult to improve systematically from evidence. We present DuplexAgent, a full-duplex collaboration system whose harness expresses this workflow as six editable modules, and Duplex-Harness-RSI, a closed loop that revises them from interaction traces. A simulator automatically generates timed test conversations, runs the system, and produces failure traces that identify the collaboration modules requiring repair. Reasoning LLMs and coding agents in the delegation pool also serve the improvement loop: the Exam Planner selects the next tests from observed weaknesses and the repair archive, and the Harness Editor proposes targeted module changes. The capabilities that serve the user thus also improve the system's coordination. Experiments on intelligence, agentic, and duplex benchmarks show that DuplexAgent combines continuous interaction with difficult reasoning and complex task execution, achieving stronger spoken-knowledge and executable-tool scores than the compared delegated systems while maintaining strong interruption response. A harness ablation further shows that this modular, verifiable loop outperforms the initial harness and repeated editing that lacks its diagnosis and repair archive.
Figures & tables
Figure 1: Benchmark results for DuplexAgent (blue), compared with other collaboration systems (gray). Higher is better on every panel. Intelligence reports MMSU accuracy, agentic tool use reports BFCL parallel-multiple and FDB-v3 Pass@1, and duplex reports FDB-v1 interrupt TOR. DuplexAgent is stronger on all three capabilities. All collaboration systems use the same delegation LLM. See Tables 4 and 5 .
Figure 2: Schematic of DuplexAgent and Duplex-Harness-RSI.
Figure 3: A full-duplex collaboration session. The user requests Snake, then cancels it and asks for Tetris while work is running. The obsolete Snake result is discarded. The current Tetris result returns through the interaction model after the user pauses.
Module
Collaboration decision
Editable surface
Policy
Answer, delegate, or task control
Instructions, delegation schema, routing examples
Intake
Operation expressed, and which live task
Parsing, control cues, task relations, resolution rules
Hold
What the user hears while a task is pending
Receipts, progress timing, control replies
Worker
Delegate, context, and result contract
Routing, task instructions, result schemas, executable overrides
Deliver
When and how a current result returns
Speaking window, waiting, success or failure rendering
Memory
Context shared by conversation and background work
Table 1: Six editable modules of the DuplexAgent harness.
Figure 4: Test-conversation construction. The Exam Planner selects coverage over request types, interaction patterns, and input conditions. Content generation and deterministic temporal composition produce a corpus with explicit expectations. Speech rendering and acoustic augmentation turn it into timed audio input.
Failure family
Default attribution
Trace evidence
Perception
Outside harness
Requested turn not recovered well enough for the harness to act
Routing
Policy
Missed or extra delegation, or a reply to non-directed speech
Interpretation
Intake
Malformed request, or control applied to the wrong task
Continuity
Hold
Receipts conflict with speech, handoffs repeat, or pending work lacks progress
Execution
Worker
Execution fails, content is wrong, or the result contract is violated
Delivery
Deliver
Missing, stale, duplicated, late, unspoken, spoken over the user, or failure shown as success
Table 2: Default failure attribution, refined using intermediate decisions and returned payloads.
Component
Experimental setting
Content bank
Content split before temporal composition into development, protection, and sealed sets (60/20/20).
Exam Planner
DeepSeek-v4-flash; bounded weights over request types, interaction patterns, and input conditions, with request-coverage floors.
Adaptive exam
Fixed core share 40%; fresh-item budget of 30 per round, with weighted sampling of the remaining cases.
Harness Editor
Codex using gpt-5.6-sol ; three candidates and up to two refinement attempts per round.
Development gate
More paired wins than losses, preserved collaboration invariants, and cause regressions within a fixed tolerance.
Protection gate
A separate protection set; at least as many paired wins as losses and preserved invariants.
Table 3: Configuration used in the RSI experiments.
BFCL-v3
FDB-v3
Method
Delegate
Average
Simp.
Mult.
Para.
P-M
Irre.
F1
Argu.
Pass@1
Moshi
✗
20.0
0.0
0.0
0.0
0.0
100.0
0.0
0.0
0.0
MoshiRAG
✓
20.0
0.0
0.0
0.0
0.0
100.0
0.0
0.0
0.0
MiniCPM-o 4.5
✗
20.0
0.0
0.0
0.0
0.0
100.0
0.0
0.0
0.0
Gander
✓
17.7
0.0
0.0
0.0
0.0
88.3
0.0
0.0
0.0
Realtime-Omni-9B
✗
68.8
74.0
57.0
71.0
53.0
89.2
7.0
4.2
2.0
Table 4: Executable tool use on BFCL and FDB-v3. DuplexAgent scores higher than the standalone interaction model and Qwen-Audio-Agent on both attachments. Scores are percentages. ∗ denotes VoiceChat and † denotes Qwen-Realtime.
Figure 5: DuplexAgent scores above both its standalone interaction model and Qwen-Audio-Agent on the same attachment.
Intelligence
FDB-v1
Method
Delegate
OBQA
MMSU
Pause Synthetic ↓
Pause Candor ↓
Turn TOR ↑
Interrupt TOR ↑
Moshi
✗
23.8
24.2
1.000
0.991
1.000
0.94
MoshiRAG
✓
30.8
28.1
1.000
0.981
1.000
0.83
MiniCPM-o 4.5
✗
69.6
55.0
0.132
0.222
0.898
0.85
Gander
✓
82.4
62.7
0.000
0.019
0.390
0.72
Realtime-Omni-9B
✗
84.1
64.5
0.000
0.019
0.864
0.68
Table 5: Spoken knowledge and duplex timing. DuplexAgent raises spoken-knowledge accuracy on both attachments, while duplex timing stays near the attached interaction model. Marks follow Table 4 .
Figure 6: VoiceChat harness ablation on the fixed core. The left panel is the collaboration score Q , and the right panel is the rate of each diagnostic failure. Duplex-Harness-RSI scores higher than one-shot or sequential editing.
Figure 7: Duplex-Harness-RSI on Qwen-Realtime. Retained versions raise the collaboration score on the fixed core from the initial harness g0 to g8 .
Module
Rule before repair
Rule after repair
Recorded local evidence
Policy †
Conversational and familiar-fact turns are routed too broadly.
State the direct-answer categories explicitly, while still delegating requests that need additional capability.
Extra calls reduced by 50%.
Intake ∗
A late cancellation can be treated as work that has already completed.
Resolve the cancellation against live task state, and apply the same control handling to replacement preambles.
Ignored cancellations reduced by 81%. Replacements reduced by 65%, then by 83.3%.
Hold †
Pending work does not receive a timely progress update.
Report progress earlier, and keep receipts brief and compatible with the adapter.
Missing progress updates reduced by 75%.
Worker ∗
Conversions skip the required calculation, or reformulate output that was already verified.
Preserve the requested units and precision, calculate when the request requires it, and retain verified result content.
Wrong answers reduced by 66.7%.
Deliver ∗
Completion can start speech during the user’s turn, and an inability to execute is reported as success.
Speak only once the user is silent, and distinguish unsuccessful execution from a successful answer.
Overlaps reduced by 100% on both attachments. Concealed failures reduced by 100%.
Memory ∗
Delegation is correct, but the delegate receives an incomplete question or missing context.
Assemble the complete user turn, and retain the context that a delegated request needs.
Missing-context failures reduced by 60%.
Table 6: Representative harness repairs. Each change reduces the targeted failures on the local exam used to test it. Marks follow Table 4 .
Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The contribution may be a missing word, a correction or a clarification. We introduce Duplex Cue, an evaluation of this \emph{in-turn adaptation} in full-duplex voice agents. Duplex Cue separates listener intent (backchannel, collaboration, or interruption) from speaker behavior: continuing unchanged, adapting within the turn, or yielding. Adaptation includes acknowledgment as well as content revision. In a single-model case study using 300 human-confirmed cues from unscripted English conversations, we compare recorded human responses with PersonaPlex continuations generated while replaying the listener's audio. We retain 208 pairs with the ongoing speaker active at cue onset and a scorable response in each condition. On the 66 collaborative pairs, recorded speakers adapt in 68.2% of cases, compared with 34.8% for PersonaPlex. The model otherwise continues unchanged (42.4%) or yields (22.7%). These findings show why evaluating natural voice interaction requires measuring how an agent responds to a listener's contribution as well as whether it keeps speaking.
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM (system 1) with a powerful asynchronous slow-thinking LLM agent (system 2). Beyond handling real-time interaction, system 1 learns when to answer directly and when to delegate, remaining responsive during backend execution and seamlessly integrating returned information into the ongoing dialogue without exposing tool traces or losing conversational context. Evaluations on single-turn spoken question answering (QA) and multi-turn conversations demonstrate that adaptive delegation substantially improves accuracy on knowledge-intensive and multi-hop reasoning questions, while knowledge-boundary-aware training avoids unnecessary system 2 invocations. On a customized version of τ-Voice, SALMONN-duo further demonstrates its ability to complete environment-grounded, policy-constrained tasks through multi-turn interactions in realistic business scenarios. Finally, cost-aware reinforcement learning further enhances the trade-off between task performance and backend usage across the QA and conversation tasks, while improving task success and response safety on τ-Voice with an acceptable increase in the delegation rate.
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.