Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation. A duplex model supports continuous listening and speaking, but complex reasoning and tool use may exceed its capabilities. A coding agent can plan and execute extended tasks, but its sequential interface is a poor fit for live conversation. Combining them requires a harness that coordinates task acceptance, progress, cancellation, replacement, and result delivery while keeping the conversation responsive. Existing harnesses often rely on coupled heuristics, making them difficult to improve systematically from evidence. We present DuplexAgent, a full-duplex collaboration system whose harness expresses this workflow as six editable modules, and Duplex-Harness-RSI, a closed loop that revises them from interaction traces. A simulator automatically generates timed test conversations, runs the system, and produces failure traces that identify the collaboration modules requiring repair. Reasoning LLMs and coding agents in the delegation pool also serve the improvement loop: the Exam Planner selects the next tests from observed weaknesses and the repair archive, and the Harness Editor proposes targeted module changes. The capabilities that serve the user thus also improve the system's coordination. Experiments on intelligence, agentic, and duplex benchmarks show that DuplexAgent combines continuous interaction with difficult reasoning and complex task execution, achieving stronger spoken-knowledge and executable-tool scores than the compared delegated systems while maintaining strong interruption response. A harness ablation further shows that this modular, verifiable loop outperforms the initial harness and repeated editing that lacks its diagnosis and repair archive.
Figures & tables
Figure 1: Benchmark results for DuplexAgent (blue), compared with other collaboration systems (gray). Higher is better on every panel. Intelligence reports MMSU accuracy, agentic tool use reports BFCL parallel-multiple and FDB-v3 Pass@1, and duplex reports FDB-v1 interrupt TOR. DuplexAgent is stronger on all three capabilities. All collaboration systems use the same delegation LLM. See Tables 4 and 5 .
Figure 2: Schematic of DuplexAgent and Duplex-Harness-RSI.
Figure 3: A full-duplex collaboration session. The user requests Snake, then cancels it and asks for Tetris while work is running. The obsolete Snake result is discarded. The current Tetris result returns through the interaction model after the user pauses.
Module
Collaboration decision
Editable surface
Policy
Answer, delegate, or task control
Instructions, delegation schema, routing examples
Intake
Operation expressed, and which live task
Parsing, control cues, task relations, resolution rules
Hold
What the user hears while a task is pending
Receipts, progress timing, control replies
Worker
Delegate, context, and result contract
Routing, task instructions, result schemas, executable overrides
Deliver
When and how a current result returns
Speaking window, waiting, success or failure rendering
Memory
Context shared by conversation and background work
Table 1: Six editable modules of the DuplexAgent harness.
Figure 4: Test-conversation construction. The Exam Planner selects coverage over request types, interaction patterns, and input conditions. Content generation and deterministic temporal composition produce a corpus with explicit expectations. Speech rendering and acoustic augmentation turn it into timed audio input.
Failure family
Default attribution
Trace evidence
Perception
Outside harness
Requested turn not recovered well enough for the harness to act
Routing
Policy
Missed or extra delegation, or a reply to non-directed speech
Interpretation
Intake
Malformed request, or control applied to the wrong task
Continuity
Hold
Receipts conflict with speech, handoffs repeat, or pending work lacks progress
Execution
Worker
Execution fails, content is wrong, or the result contract is violated
Delivery
Deliver
Missing, stale, duplicated, late, unspoken, spoken over the user, or failure shown as success
Table 2: Default failure attribution, refined using intermediate decisions and returned payloads.
Component
Experimental setting
Content bank
Content split before temporal composition into development, protection, and sealed sets (60/20/20).
Exam Planner
DeepSeek-v4-flash; bounded weights over request types, interaction patterns, and input conditions, with request-coverage floors.
Adaptive exam
Fixed core share 40%; fresh-item budget of 30 per round, with weighted sampling of the remaining cases.
Harness Editor
Codex using gpt-5.6-sol ; three candidates and up to two refinement attempts per round.
Development gate
More paired wins than losses, preserved collaboration invariants, and cause regressions within a fixed tolerance.
Protection gate
A separate protection set; at least as many paired wins as losses and preserved invariants.
Table 3: Configuration used in the RSI experiments.
BFCL-v3
FDB-v3
Method
Delegate
Average
Simp.
Mult.
Para.
P-M
Irre.
F1
Argu.
Pass@1
Moshi
✗
20.0
0.0
0.0
0.0
0.0
100.0
0.0
0.0
0.0
MoshiRAG
✓
20.0
0.0
0.0
0.0
0.0
100.0
0.0
0.0
0.0
MiniCPM-o 4.5
✗
20.0
0.0
0.0
0.0
0.0
100.0
0.0
0.0
0.0
Gander
✓
17.7
0.0
0.0
0.0
0.0
88.3
0.0
0.0
0.0
Realtime-Omni-9B
✗
68.8
74.0
57.0
71.0
53.0
89.2
7.0
4.2
2.0
Table 4: Executable tool use on BFCL and FDB-v3. DuplexAgent scores higher than the standalone interaction model and Qwen-Audio-Agent on both attachments. Scores are percentages. ∗ denotes VoiceChat and † denotes Qwen-Realtime.
Figure 5: DuplexAgent scores above both its standalone interaction model and Qwen-Audio-Agent on the same attachment.
Intelligence
FDB-v1
Method
Delegate
OBQA
MMSU
Pause Synthetic ↓
Pause Candor ↓
Turn TOR ↑
Interrupt TOR ↑
Moshi
✗
23.8
24.2
1.000
0.991
1.000
0.94
MoshiRAG
✓
30.8
28.1
1.000
0.981
1.000
0.83
MiniCPM-o 4.5
✗
69.6
55.0
0.132
0.222
0.898
0.85
Gander
✓
82.4
62.7
0.000
0.019
0.390
0.72
Realtime-Omni-9B
✗
84.1
64.5
0.000
0.019
0.864
0.68
Table 5: Spoken knowledge and duplex timing. DuplexAgent raises spoken-knowledge accuracy on both attachments, while duplex timing stays near the attached interaction model. Marks follow Table 4 .
Figure 6: VoiceChat harness ablation on the fixed core. The left panel is the collaboration score Q , and the right panel is the rate of each diagnostic failure. Duplex-Harness-RSI scores higher than one-shot or sequential editing.
Figure 7: Duplex-Harness-RSI on Qwen-Realtime. Retained versions raise the collaboration score on the fixed core from the initial harness g0 to g8 .
Module
Rule before repair
Rule after repair
Recorded local evidence
Policy †
Conversational and familiar-fact turns are routed too broadly.
State the direct-answer categories explicitly, while still delegating requests that need additional capability.
Extra calls reduced by 50%.
Intake ∗
A late cancellation can be treated as work that has already completed.
Resolve the cancellation against live task state, and apply the same control handling to replacement preambles.
Ignored cancellations reduced by 81%. Replacements reduced by 65%, then by 83.3%.
Hold †
Pending work does not receive a timely progress update.
Report progress earlier, and keep receipts brief and compatible with the adapter.
Missing progress updates reduced by 75%.
Worker ∗
Conversions skip the required calculation, or reformulate output that was already verified.
Preserve the requested units and precision, calculate when the request requires it, and retain verified result content.
Wrong answers reduced by 66.7%.
Deliver ∗
Completion can start speech during the user’s turn, and an inability to execute is reported as success.
Speak only once the user is silent, and distinguish unsuccessful execution from a successful answer.
Overlaps reduced by 100% on both attachments. Concealed failures reduced by 100%.
Memory ∗
Delegation is correct, but the delegate receives an incomplete question or missing context.
Assemble the complete user turn, and retain the context that a delegated request needs.
Missing-context failures reduced by 60%.
Table 6: Representative harness repairs. Each change reduces the targeted failures on the local exam used to test it. Marks follow Table 4 .