Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information required for successful execution has been preserved. We introduce Talk2Agent, a benchmark for evaluating how effectively voice interfaces convey human-spoken instructions to LLM-based computer-use agents. Talk2Agent builds human-spoken versions of tasks from WildClawBench and OSWorld and evaluates a range of voice interfaces, including dedicated ASR models, audio-capable LLMs, contextual biasing, and LLM-based ontology repair. Because repeatedly executing long-horizon computer-use tasks is costly and stochastic, we further propose an execution-free, task-conditioned evaluation framework that projects the original task grader onto prompt-addressable intentions and measures how much task-relevant information is retained after the voice interface. On WildClawBench, Talk2Agent's execution-free native projection provides a practical, execution-grounded measure of voice-interface quality, correlating with downstream task completion and improving Pearson correlation by 0.246 over WER/CER on 32 hours of real human speech.
Figures & tables
Source
Coverage
Hours
Median [IQR] (s)
WCB
60×10=600
24.84
113.66 [55.01–212.31]
OSWorld
351×5=1,755
6.98
11.54 [7.00–18.28]
Table 1: Talk2Agent speech collections. Coverage is tasks × recordings per task = total recordings.
Figure 1: Talk2Agent dataset construction and evaluation. Written agent tasks are colloquialized to spoken forms and recorded by human speakers. Voice-interface outputs are evaluated either by a fixed downstream text agent and the source task grader or by execution-free projection of that grader onto retained task intentions.
Input / interface
Biasing
Ontology Repair
Execution (%)
Original text
–
–
40.75±1.67
Spoken-form text
–
–
36.39±3.56
Parakeet
✗
✗
24.20±2.76
Parakeet
✗
✓
25.32±1.13
Parakeet
✓
✗
25.28±3.39
Parakeet
✓
✓
25.03±2.95
Table 2: Results on Talk2Agent WCB split. Mean ± std. deviation reported across three round-level interface means; each round equally averages the 60 tasks. Voice rows use TTS audio.
Agent / input
Biasing
Ontology Repair
Mean
SD
Gemini agent
Text
–
–
61.76
0.37
GPT-4o
✗
✗
58.49
0.64
GPT-4o
✗
✓
59.28
1.09
Whisper
✗
✗
56.07
1.34
Whisper
✗
✓
56.77
1.50
Table 3: OSWorld execution from the frozen summary. Values are means and SDs over five speaker recordings per task
Metric
Task-interface
Type-interface
Interface
1− WER/CER
0.220
0.176
0.514
1− SemDist
0.218
0.178
0.606
1− AER
0.286
0.297
0.916
Atomic Rubric
0.259
0.243
0.787
1− Semantic WER
0.058
0.143
0.834
Flat retention
0.349
0.441
0.895
Table 4: PCC against WCB execution over 500 task–interface pairs, 50 task-type–interface aggregates, and ten interface means. Flat equally weights intentions; other metrics are in Sec. 5 .
Cohort
Biasing
Ont. fix
Shapley
Native
Execution
Targeted
✗
✗
39.93
14.00
40.60
( n=15 )
✗
✓
69.33
34.00
52.45
✓
✗
41.30
14.67
33.97
✓
✓
59.25
33.67
50.06
Full
✗
✗
33.39
7.73
22.64
( n=50 )
✗
✓
54.84
23.24
22.68
Table 5: Parakeet diagnosis on WCB. Targeted split comprises the 15 non-Safety tasks in the preselected 16-task workspace-intensive probe. Full comprises all 50 non-Safety WCB tasks. Bold indicates improvements from ontology repair under matched biasing. All values are percentages.
Measure
Raw
+Ont.
Δ
Whisper
54.93±1.18
62.89±2.05
+7.96±1.52
Table 6: OSWorld execution-free transfer, mean ± SD across five paired Whisper takes. Improvements are percentage points.
Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be converted into controlled audio-based tool calling evaluations without re-annotating the tool schema and gold labels. Our dataset-agnostic framework uses text-to-speech, speaker variation, and environmental noise to create paired text-audio instances while preserving the original dataset annotations. Based on extensive evaluation of 7 omni-modal models on audio-converted versions of Confetti and When2Call, our framework demonstrates that the performance is strongly model- and task-dependent: Gemini-3.1-Flash-Live obtains the highest Confetti score (70.4), whereas GPT-Realtime-1.5 performs best on When2Call (71.9). On Confetti, the text-to-voice gap ranges from 1.8 points for Qwen3-Omni to 4.8 points for GPT-Realtime-1.5. A targeted analysis of failure cases demonstrates that degradations most often reflect misunderstandings of argument values in the speech. Considering real-world deployment scenarios, we further report text-only results, an ambiguity-based reformulation stress test, and a reference-free LLM-as-judge protocol validated against human preferences. Notably, we find that open-source Qwen3 judges with at least 8B parameters exceed 80% agreement with proprietary judges, supporting privacy-preserving evaluation. Overall, our framework provides a verifiable and reproducible first-stage diagnostic that complements purpose-built audio corpora.
Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning complexity to evaluate tool-calling performance. We introduce Audio2Tool, a large-scale dataset comprising approximately 30,000 queries designed to assess tool-calling capabilities of SpeechLMs across three primary domains: Smart Car, Smart Home, and Wearables. Our benchmark features a multi-tier complexity hierarchy, ranging from simple direct commands to complex multi-intent and needle-in-a-haystack extraction to isolate distinct failure modes. To ensure realism, we employ zero-shot voice cloning text-to-speech synthesis and diverse noise profiles to simulate in-the-wild conditions. Evaluations of state-of-the-art SpeechLMs and ASR-LLM pipelines show strong performance on simple commands but significant degradation under compositional and acoustic challenges. Code and dataset are publicly available on the project page: https://audio2tool.github.io/.
Production voice agents span cascaded, speech-to-speech, and hybrid architectures. Voice-agent benchmarks typically measure component quality and conversational properties such as word error rate, latency, naturalness, and turn-taking. Fewer measure whether the agent handled a phone call correctly on its own. Contact centers refer to this as ``containment'': the share of phone calls the automated system resolves without handing off to a human. On some phone calls the right outcome is refusal or a redirect. To address this gap, we introduce VAmoS Bench, the Voice Agent Simulation Bench. It measures complete voice-agent systems end to end in a stateful customer-support task. The agent is Riley, a credit-card support representative for a fictional bank who can freeze, cancel, replace, or activate a card. Each of 100 scenarios supplies a simulated caller with a private goal and a seeded PostgreSQL backend. The platform uses each scenario to populate and activate an isolated simulation in which the caller reaches Riley over audio; roughly one-third apply adversarial pressure. The agent can use five tools that execute real SQL against the backend. Each scenario also defines binary assertions. A grader evaluates them against the complete trace of what the caller and agent said and what the agent did, including tool invocations, arguments, and returned rows. This catches an agent that claims to have changed a card without updating the database, as well as one that makes the right database change while disclosing protected information. This first benchmark version focuses on financial services. Its evaluation protocol supports an evolving leaderboard: additional voice agents can be evaluated on the same version, while later versions can expand the tasks and scenarios.