APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction
Organizations: University of Maryland College Park, USA
Abstract
Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completion, corporate negotiation, coordination, consulting, and interviewing. Each workflow executes in a stateful Voice Workbench environment with task-specific knowledge, typed tools, gold-annotated final work artifact, authorization constraints, and a user simulation policy backed by validated, pre-compiled speech realizations. We evaluate both artifact field accuracy and end-to-end workflow success, which requires the correct terminal state, valid process, completed actions, and a valid final artifact. Across five frontier real-time voice agents-GPT-Live-1, Gemini-3.8-Live, Grok-Voice-Think-2.0, Step-Audio3, and GPT-realtime-2.1, none exceeds 25% Pass@1, and the best Reliable@3 is only 10.8%. Moreover, stateful coordination is the dominant failure point across systems, while success decreases further on workflows requiring greater knowledge retrieval and mid-speech corrections. Overall, APEX-Voice is the first benchmark for evaluating whether voice agents can translate conversational competence into dependable professional work.
Figures & tables
| Benchmark | Full-duplex voice | User simulation | Knowledge grounding | Stateful tools | Professional work | Work-artifact output | Correction / authorization | Verifiable outcome |
|---|---|---|---|---|---|---|---|---|
| Full-duplex and voice-agent benchmarks | ||||||||
| Full-Duplex-Bench v1/1.5 ( Lin et al., 2025 ; Lin et al., 2026d ) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Full-Duplex-Bench v2 ( Lin et al., 2026c ) | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ |
| FD-Bench ( Peng et al., 2025 ) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| MTR-DuplexBench ( He et al., 2026 ) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Full-Duplex-Bench v3 ( Lin et al., 2026b ) | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ |
| Model | Pass@1 | Pass@3 | Reliable@3 | Tool-use Efficiency ( ) |
|---|---|---|---|---|
| Cascaded (Whisper-LV3–GPT-5.6–Chatterbox-TurboTTS) | 1.2 | 3.2 | 0.0 | 0.25 |
| Step-Audio3 | 2.5 | 5.0 | 0.8 | 0.38 |
| Gemini-3.8-Live | 13.6 | 28.3 | 1.7 | 1.09 |
| GPT-live-1 | 8.9 | 15.8 | 2.5 | 0.94 |
| Grok-Voice-Think-2.0 | 23.1 | 41.7 | 5.0 | 1.51 |
| GPT-realtime-2.1 | 23.6 | 36.7 | 10.8 | 0.98 |
| Model | Target State (TS) | Process Validity (PV) | Action Completion (AC) | Artifact Validity (AV) | Artifact Field Accuracy (AV) |
|---|---|---|---|---|---|
| GPT-realtime-2.1 | 81.1 | 89.2 | 61.4 | 35.3 | 91.4 |
| Grok-Voice-Think-2.0 | 99.4 | 85.3 | 85.8 | 27.8 | 88.5 |
| Gemini-3.8-Live | 84.4 | 74.2 | 49.7 | 24.4 | 84.5 |
| GPT-live-1 | 88.6 | 91.9 | 65.8 | 11.9 | 71.7 |
| Step-Audio3 | 34.4 | 56.4 | 56.9 | 4.2 | 72.3 |
| Floor Control | Correction Uptake | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | Speech Overlap (%) | Barge-in Yield (%) | Stop Latency p50 (ms) | Stop Latency p95 (ms) | AFA: Corrected Fields (%) | AFA: Other Fields (%) | Uptake Gap (pts) | Correction-Linked AV Fails (%) |
| GPT-realtime-2.1 | 2.53 | 99.8 | 156 | 294 | 74.0 | 95.1 | 21.1 | 74 |
| Grok-Voice-Think-2.0 | 1.31 | 100.0 | 28 | 64 | 72.7 | 92.6 | 20.0 | 71 |
| Gemini-3.8-Live | 1.82 | 100.0 | 20 | 47 | 64.3 | 88.2 | 23.8 | 77 |
| Step-Audio3 | 4.55 | 100.0 | 59 | 105 | 38.4 | 75.2 | 36.8 | 89 |
| GPT-live-1 | 34.79 | 95.6 | 912 | 1424 | 49.4 | 78.6 | 29.2 | 82 |
| Text controls | Voice reference | |||||
| Metric | Claude Opus-5.5 | GPT- 6-sol | GPT- 5.5 | Gemini 3.8-Flash | Kimi K3 | GPT- realtime-2.1 |
| Pass@1 (%) | 58.7 | 55.0 | 62.0 | 54.3 | 55.8 | 23.3 |
| AFA (%) | 97 | 96 | 95 | 98 | 97 | 91 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Taxonomy dimension | Realized labels |
|---|---|
| Work archetype | form-fill, interview, intake, troubleshoot, negotiate, coordinate, discovery, advise, facilitate, inspect |
| Industry / setting | Software/SaaS, horizontal enterprise, manufacturing/field operations, professional services, workplace/HR, healthcare, insurance |
| Work artifact | Structured form, case record, CRM record, evidence matrix, memo/report, plan/checklist, schedule, ticket, timeline, negotiation record, work order |
| Economic role / function | Recruiting, HR operations, sales, customer success, technical support, insurance operations, finance operations, procurement, project management, operations, field service, consulting, compliance, executive assistance |
| Duplex phenomenon | Backchannel, user barge-in, mid-speech correction, cancellation/revocation, intent switch, clarification, overlapping speech |
| Delegation pattern | delegate, complete, revise, follow-through, approve |
| Dimension | Operational interpretation |
|---|---|
| Work archetype | Underlying professional operation, such as interviewing, troubleshooting, coordination, negotiation, or form completion. |
| Industry / setting | Organizational context that determines task terminology, policies, resources, and constraints. |
| Work artifact | Persistent and verifiable output produced by the workflow, such as a form, case record, schedule, memo, or negotiation record. |
| Economic role / function | Professional function carrying out the workflow, such as recruiting, sales, technical support, procurement, or project management. |
| Duplex phenomenon | Task-critical real-time conversational event, including barge-ins, overlap, corrections, cancellations, clarifications, and backchannels. |
| Delegation pattern | How the workflow evolves: inferring procedure (delegate), discovering missing information (complete), propagating corrections (revise), adapting to environment changes (follow-through), or obtaining authorization (approve). |
| Grader | Rule | Example (gold agent score) |
|---|---|---|
| exact | String equality after whitespace trim. | E4471 E4471 |
| casefold_exact | Case-insensitive equality after trim. | HDHP hdhp |
| normalized_phone | Compare digits only (strip formatting). | 4155550134 (415) 555-0134 |
| normalized_date | Canonicalize to YYYY-MM-DD (ISO, US slash, or month-name forms). | 1988-07-09 July 9, 1988 |
| normalized_address | Lowercase, collapse non-alphanumerics to spaces, compare. | 12 oak st 12 Oak St. |
| enum | Exact match against a controlled value. | submitted submitted |
| Human validation metric | Result | 95% CI |
|---|---|---|
| Semantic fidelity (1–5) | 4.5 | [4.34, 4.61] |
| Speech naturalness (1–5) | 4.2 | [4.14, 4.33] |
| Unauthorized-fact leakage | 4.1% | [4.09%, 4.16%] |
| Ordinal agreement ( ) | 0.79 | — |
| Leakage agreement ( ) | 0.88 | — |
| Model | Synthetic Pass@1 | Human Pass@1 | Pass@1 |
|---|---|---|---|
| GPT-realtime-2.1 | 24.1 | 20.8 | |
| Grok-Voice-Think-2.0 | 24.1 | 16.4 | |
| Gemini-3.8-Live | 20.9 | 16.3 | |
| Step-Audio3 | 12.1 | 7.5 | |
| GPT-live-1 | 7.5 | 3.1 |
| Setting | GPT-rt | Grok | Gemini | Step3 | GPT-live1 | |
|---|---|---|---|---|---|---|
| Software / SaaS | 43 | 27% (94%) | 26% (90%) | 14% (86%) | 2% (76%) | 12% (72%) |
| Horizontal enterprise | 30 | 23% (91%) | 22% (90%) | 12% (85%) | 1% (75%) | 16% (78%) |
| Manufacturing / field ops | 25 | 9% (88%) | 20% (88%) | 7% (83%) | 1% (66%) | 1% (66%) |
| Professional services | 10 | 30% (89%) | 20% (87%) | 10% (79%) | 10% (76%) | 3% (63%) |
| Workplace HR | 8 | 33% (91%) | 21% (86%) | 17% (84%) | 4% (70%) | 12% (73%) |
| Healthcare | 3 | 56% (92%) | 22% (87%) | 33% (82%) | 0% (61%) | 0% (75%) |
| Artifact | GPT-rt | Grok | Gemini | Step3 | GPT-live1 | |
|---|---|---|---|---|---|---|
| Negotiation record | 20 | 10% (88%) | 15% (86%) | 3% (79%) | 0% (63%) | 3% (75%) |
| Case record | 10 | 27% (91%) | 20% (87%) | 20% (88%) | 0% (55%) | 3% (71%) |
| CRM record | 10 | 23% (96%) | 23% (91%) | 13% (86%) | 0% (88%) | 7% (65%) |
| Evidence matrix | 10 | 47% (95%) | 27% (89%) | 27% (86%) | 10% (69%) | 7% (64%) |
| Memo / report | 10 | 43% (96%) | 27% (83%) | 13% (90%) | 3% (74%) | 17% (79%) |
| Plan / checklist | 10 | 47% (95%) | 43% (91%) | 27% (87%) | 7% (82%) | 23% (69%) |
| Autonomy level | GPT-rt | Grok | Gemini | Step3 | GPT-live1 | |
|---|---|---|---|---|---|---|
| Prepare only | 51 | 34% (94%) | 29% (89%) | 20% (87%) | 5% (75%) | 9% (70%) |
| Draft + confirm | 29 | 14% (91%) | 18% (89%) | 8% (86%) | 2% (77%) | 14% (74%) |
| Low-risk execute | 16 | 21% (89%) | 17% (89%) | 10% (86%) | 0% (67%) | 6% (73%) |
| Approval-gated commit | 24 | 15% (88%) | 17% (87%) | 1% (78%) | 0% (65%) | 8% (70%) |
| Slice | GPT-rt | Grok | Gemini | Step3 | GPT-live1 | |
|---|---|---|---|---|---|---|
| No external knowledge | 44 | 39% (93%) | 31% (90%) | 26% (86%) | 5% (77%) | 15% (71%) |
| Supplied documents | 6 | 6% (88%) | 6% (85%) | 0% (76%) | 0% (73%) | 0% (67%) |
| Small knowledge search | 51 | 12% (90%) | 19% (88%) | 3% (84%) | 1% (68%) | 6% (72%) |
| Multi-document policy | 19 | 25% (93%) | 18% (86%) | 11% (85%) | 2% (72%) | 11% (75%) |
| Light tools | 44 | 39% (93%) | 31% (90%) | 26% (86%) | 5% (77%) | 15% (71%) |
| Moderate tools | 76 | 15% (90%) | 18% (88%) | 4% (83%) | 1% (69%) | 7% (72%) |
| Profile | GPT-rt | Grok | Gemini | Step3 | GPT-live1 | |
|---|---|---|---|---|---|---|
| Cooperative | 24 | 28% (93%) | 33% (92%) | 12% (84%) | 1% (71%) | 10% (72%) |
| Correction-prone | 24 | 19% (89%) | 18% (88%) | 11% (86%) | 0% (76%) | 8% (72%) |
| Ambiguous / underspecified | 12 | 31% (93%) | 17% (87%) | 14% (84%) | 0% (69%) | 14% (75%) |
| Distracted / time-pressured | 12 | 8% (90%) | 11% (85%) | 19% (90%) | 3% (77%) | 8% (71%) |
| Domain expert | 12 | 17% (91%) | 28% (87%) | 11% (82%) | 3% (74%) | 11% (61%) |
| Low-tech expertise | 12 | 31% (95%) | 31% (89%) | 11% (84%) | 11% (59%) | 14% (77%) |
| Slice | GPT-rt | Grok | Gemini | Step3 | GPT-live1 | |
|---|---|---|---|---|---|---|
| Not APPROVE-gated | 91 | 26% (92%) | 25% (89%) | 15% (86%) | 3% (75%) | 10% (71%) |
| APPROVE-gated | 29 | 15% (88%) | 14% (88%) | 2% (80%) | 0% (64%) | 9% (72%) |
| 9 required fields | 24 | 31% (93%) | 26% (88%) | 12% (85%) | 4% (68%) | 14% (73%) |
| 10 required fields | 74 | 17% (91%) | 22% (89%) | 9% (84%) | 1% (76%) | 10% (74%) |
| 11–13 required fields | 22 | 38% (92%) | 21% (89%) | 21% (85%) | 5% (66%) | 3% (64%) |
| Routine | 63 | 21% (92%) | 25% (90%) | 12% (86%) | 3% (78%) | 13% (73%) |
| Model | Reliable@3 (%) | Median TTR (s) |
|---|---|---|
| GPT-realtime-2.1 | 10.8 | 222 |
| Grok-Voice-Think-2.0 | 5.0 | 206 |
| Gemini-3.8-Live | 1.7 | 223 |
| Step-Audio3 | 0.8 | 233 |
| GPT-live-1 | 2.5 | 207 |
| Model | TS | PV | AC | AV | |
|---|---|---|---|---|---|
| GPT-realtime-2.1 | 35 | 14 | 48 | 80 | 120 |
| Grok-Voice-Think-2.0 | 0 | 19 | 17 | 91 | 120 |
| Gemini-3.8-Live | 20 | 28 | 58 | 91 | 120 |
| Step-Audio3 | 75 | 54 | 55 | 115 | 120 |
| GPT-live-1 | 15 | 10 | 36 | 105 | 120 |
| Archetype | GPT-rt | Grok | Gemini | Step3 | GPT-live1 |
|---|---|---|---|---|---|
| ADVISE | 43% | 27% | 13% | 3% | 17% |
| COORDINATE | 20% | 22% | 12% | 3% | 23% |
| DISCOVERY | 23% | 23% | 13% | 0% | 7% |
| FACILITATE | 47% | 43% | 27% | 7% | 23% |
| FORM_FILL | 7% | 13% | 0% | 0% | 3% |
| INSPECT | 13% | 30% | 3% | 0% | 0% |
| Model | Repetition | ASR WER | Tone stability |
|---|---|---|---|
| GPT-realtime-2.1 | 8.6% | 0.09 | 0.94 |
| Grok-Voice-Think-2.0 | 5.6% | 0.04 | 0.94 |
| Gemini-3.8-Live | 9.2% | 0.05 | 0.93 |
| Step-Audio3 | 14.3% | 0.10 | 0.89 |
| GPT-live-1 | 0.9% | 1.15 | 0.90 |
| Field | Grader | Gold | Agent wrote | Score |
| employee_id | exact | E4471 | E4471 | 1 |
| legal_name | semantic | Morgan Reyes | Morgan Ray | 0 |
| date_of_birth | normalized_date | 1988-07-09 | July 9, 1988 | 1 |
| medical_plan | semantic | HDHP | HDHP | 1 |
| dental_plan | semantic | Standard | Standard | 1 |
| vision_plan | semantic | Vision Basic | Basic | 1 |