Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information required for successful execution has been preserved. We introduce Talk2Agent, a benchmark for evaluating how effectively voice interfaces convey human-spoken instructions to LLM-based computer-use agents. Talk2Agent builds human-spoken versions of tasks from WildClawBench and OSWorld and evaluates a range of voice interfaces, including dedicated ASR models, audio-capable LLMs, contextual biasing, and LLM-based ontology repair. Because repeatedly executing long-horizon computer-use tasks is costly and stochastic, we further propose an execution-free, task-conditioned evaluation framework that projects the original task grader onto prompt-addressable intentions and measures how much task-relevant information is retained after the voice interface. On WildClawBench, Talk2Agent's execution-free native projection provides a practical, execution-grounded measure of voice-interface quality, correlating with downstream task completion and improving Pearson correlation by 0.246 over WER/CER on 32 hours of real human speech.
Figures & tables
Source
Coverage
Hours
Median [IQR] (s)
WCB
60×10=600
24.84
113.66 [55.01–212.31]
OSWorld
351×5=1,755
6.98
11.54 [7.00–18.28]
Table 1: Talk2Agent speech collections. Coverage is tasks × recordings per task = total recordings.
Figure 1: Talk2Agent dataset construction and evaluation. Written agent tasks are colloquialized to spoken forms and recorded by human speakers. Voice-interface outputs are evaluated either by a fixed downstream text agent and the source task grader or by execution-free projection of that grader onto retained task intentions.
Input / interface
Biasing
Ontology Repair
Execution (%)
Original text
–
–
40.75±1.67
Spoken-form text
–
–
36.39±3.56
Parakeet
✗
✗
24.20±2.76
Parakeet
✗
✓
25.32±1.13
Parakeet
✓
✗
25.28±3.39
Parakeet
✓
✓
25.03±2.95
Table 2: Results on Talk2Agent WCB split. Mean ± std. deviation reported across three round-level interface means; each round equally averages the 60 tasks. Voice rows use TTS audio.
Agent / input
Biasing
Ontology Repair
Mean
SD
Gemini agent
Text
–
–
61.76
0.37
GPT-4o
✗
✗
58.49
0.64
GPT-4o
✗
✓
59.28
1.09
Whisper
✗
✗
56.07
1.34
Whisper
✗
✓
56.77
1.50
Table 3: OSWorld execution from the frozen summary. Values are means and SDs over five speaker recordings per task
Metric
Task-interface
Type-interface
Interface
1− WER/CER
0.220
0.176
0.514
1− SemDist
0.218
0.178
0.606
1− AER
0.286
0.297
0.916
Atomic Rubric
0.259
0.243
0.787
1− Semantic WER
0.058
0.143
0.834
Flat retention
0.349
0.441
0.895
Table 4: PCC against WCB execution over 500 task–interface pairs, 50 task-type–interface aggregates, and ten interface means. Flat equally weights intentions; other metrics are in Sec. 5 .
Cohort
Biasing
Ont. fix
Shapley
Native
Execution
Targeted
✗
✗
39.93
14.00
40.60
( n=15 )
✗
✓
69.33
34.00
52.45
✓
✗
41.30
14.67
33.97
✓
✓
59.25
33.67
50.06
Full
✗
✗
33.39
7.73
22.64
( n=50 )
✗
✓
54.84
23.24
22.68
Table 5: Parakeet diagnosis on WCB. Targeted split comprises the 15 non-Safety tasks in the preselected 16-task workspace-intensive probe. Full comprises all 50 non-Safety WCB tasks. Bold indicates improvements from ontology repair under matched biasing. All values are percentages.
Measure
Raw
+Ont.
Δ
Whisper
54.93±1.18
62.89±2.05
+7.96±1.52
Table 6: OSWorld execution-free transfer, mean ± SD across five paired Whisper takes. Improvements are percentage points.