cs.AISep 28, 2026

APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction

Authors: Puneet Mathur, Dinesh Manocha

Organizations: University of Maryland College Park, USA

Abstract

Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completion, corporate negotiation, coordination, consulting, and interviewing. Each workflow executes in a stateful Voice Workbench environment with task-specific knowledge, typed tools, gold-annotated final work artifact, authorization constraints, and a user simulation policy backed by validated, pre-compiled speech realizations. We evaluate both artifact field accuracy and end-to-end workflow success, which requires the correct terminal state, valid process, completed actions, and a valid final artifact. Across five frontier real-time voice agents-GPT-Live-1, Gemini-3.8-Live, Grok-Voice-Think-2.0, Step-Audio3, and GPT-realtime-2.1, none exceeds 25% Pass@1, and the best Reliable@3 is only 10.8%. Moreover, stateful coordination is the dominant failure point across systems, while success decreases further on workflows requiring greater knowledge retrieval and mid-speech corrections. Overall, APEX-Voice is the first benchmark for evaluating whether voice agents can translate conversational competence into dependable professional work.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Talk2Agent: Benchmarking Voice Interfaces for Text Agents

    Sep 30, 2026Terumi Chiba, Guangzhi Sun, Zheqi Yuan +1TalkAutomatic Speech Recognition Evaluation

  2. VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation

    Sep 29, 2026Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5Voice AgentsDialogue Benchmarks

  3. VAmoS Bench: Voice Agent Simulation Bench

    Jul 29, 2026Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5Voice AgentsDialogue Benchmarks