cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
Organizations: Carnegie Mellon University
Abstract
Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces standardized infrastructure and task sets, with a focus on evaluating the speed and efficiency of CUAs. cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks. Across four different CUA benchmarks, we evaluate how reasoning effort, agent harnesses, and environment latency affect performance, speed, and cost. We find no single model family is optimal for all three; none of the open-weight models are on the frontier, and also, unintuitively, for some models increasing the reasoning effort can speed up task completion, while faster environment input-output can slow down overall task completion time. We also demonstrate that we can effectively reduce the evaluation task set of most CUA benchmarks without degrading overall statistical power, allowing for more efficient benchmarking and comparison. We believe cua-speedrun will enable structured progress towards fast, efficient CUAs, unlocking new real-world use cases and applications. All code, infrastructure, and analysis are available at https://cuaspeedrun.com.
Figures & tables
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
| Infrastructure | Click | Escape | Ctrl+U | Type 100 characters |
|---|---|---|---|---|
| Gym-Anything | 2758.70 | 2865.00 | 2858.29 | 3521.72 |
| OSWorld | 2515.41 | 2707.74 | 2708.41 | 2756.47 |
| Fast I/O (ours) | 27.58 | 2.43 | 2.91 | 14.98 |
| Component (ms) | Click | Escape | Ctrl+U | Type 100 characters |
|---|---|---|---|---|
| Gym-Anything | ||||
| Fixed waits | 2100.2 | 2100.2 | 2100.2 | 2100.2 |
| Action request | 153.5 | 153.7 | 154.0 | 805.5 |
| Screenshot capture | 258.9 | 354.0 | 354.3 | 354.1 |
| Screenshot transfer | 238.1 | 212.4 | 220.6 | 211.3 |
| Cleanup and other | 53.1 | 53.1 | 53.1 | 53.1 |
| Component (ms) | Click | Escape | Ctrl+U | Type 100 characters |
|---|---|---|---|---|
| Input execution | 27.30 | 1.85 | 2.55 | 18.39 |
| Frame capture | 1.03 | 0.77 | 0.77 | 0.80 |
| Foreground image preparation | 0.27 | 0.02 | 0.02 | 0.18 |
| Other environment processing | 0.93 | 0.94 | 1.00 | 0.90 |
| Transport | 4.67 | 4.99 | 4.76 | 4.62 |
| Screenshot-file delivery | 0.43 | 0.45 | 0.43 | 0.44 |
| Run | Score (%) | Time (s) | Agent (s) | Steps | Env. (s) | Waits (s) | Env. waits (s) |
|---|---|---|---|---|---|---|---|
| 1 | 91.62 | 103.96 | 98.07 | 8.68 | 5.891 | 5.553 | 0.338 |
| 2 | 91.62 | 93.98 | 89.02 | 8.04 | 4.955 | 4.628 | 0.327 |
| 3 | 91.62 | 96.48 | 91.67 | 7.96 | 4.812 | 4.488 | 0.324 |
| 4 | 87.62 | 100.22 | 95.28 | 8.20 | 4.934 | 4.604 | 0.330 |
| 5 | 89.62 | 100.46 | 95.68 | 8.12 | 4.785 | 4.424 | 0.361 |
| Mean | 90.42 | 99.02 | 93.94 | 8.20 | 5.075 | 4.739 | 0.336 |
| ID | Configuration | Batches | Actions | Responses | Tokens |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra / xhigh | 7.14 | 19.90 | – | 2,060 |
| 2 | Gemini 3.8 Flash / low | 16.60 | 20.02 | 18.52 | 1,564 |
| 3 | Claude Opus 5 / high | 19.92 | 26.92 | 20.62 | 3,735 |
| 4 | Gemini 3.8 Flash / medium | 27.16 | 34.70 | 29.10 | 5,365 |
| 5 | GPT-6 Astra / medium | 6.30 | 19.26 | – | 1,366 |
| 6 | GPT-6 Astra / high | 6.52 | 19.76 | – | 1,593 |
| ID | Configuration | Batches | Actions | Responses | Tokens |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra / xhigh | 52.81 | 194.10 | – | 18,332 |
| 2 | GPT-6 Astra / high | 51.29 | 149.02 | – | 14,146 |
| 3 | GPT-6 Astra / low | 42.58 | 151.40 | – | 10,190 |
| 4 | GPT-6 Astra / medium | 45.46 | 151.08 | – | 11,382 |
| 5 | Gemini 3.8 Flash / high | 216.54 | 331.88 | 218.48 | 72,149 |
| 6 | Gemini 3.8 Flash / medium | 167.79 | 252.65 | 169.65 | 49,532 |
| ID | Configuration | Score (%) | Time (s) | Cost ($) | Tokens |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra / xhigh | 91.6 | 126.8 | 0.708 | 2,060 |
| 2 | Gemini 3.8 Flash / low | 91.6 | 127.1 | 0.108 | 1,564 |
| 3 | Claude Opus 5 / high | 91.6 | 136.0 | 0.661 | 3,735 |
| 4 | Gemini 3.8 Flash / medium | 91.6 | 250.1 | 0.221 | 5,365 |
| 5 | GPT-6 Astra / medium | 89.6 | 90.2 | 0.600 | 1,366 |
| 6 | GPT-6 Astra / high | 89.6 | 100.8 | 0.636 | 1,593 |
| ID | Configuration | Score (%) | Time (s) | Cost ($) | Tokens |
|---|---|---|---|---|---|
| 27 | GPT-5.6 Luna / low / Codex | 75.6 | 144.8 | 0.025 | 3,191 |
| 28 | MiniMax M3 / thinking_off | 75.6 | 253.8 | – | – |
| 29 | GPT-5.6 Luna / high / API | 81.6 | 160.5 | 0.029 | 3,358 |
| 30 | Kimi K3 / low / batched | 73.6 | 149.5 | 0.174 | 1,865 |
| 31 | Muse Spark 1.3 / high | 73.6 | 506.9 | 0.148 | 10,538 |
| 32 | Kimi K3 / low / single | 71.6 | 176.6 | 0.280 | 3,171 |
| ID | Configuration | Score (%) | Time (s) | Cost ($) | Tokens |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra / xhigh | 76.9 | 1237.3 | 8.175 | 18,332 |
| 2 | GPT-6 Astra / high | 75.0 | 914.0 | 7.391 | 14,146 |
| 3 | GPT-6 Astra / low | 68.2 | 904.2 | 6.336 | 10,190 |
| 4 | GPT-6 Astra / medium | 67.2 | 828.1 | 6.519 | 11,382 |
| 5 | Gemini 3.8 Flash / high | 60.2 | 3058.8 | 5.019 | 72,149 |
| 6 | Gemini 3.8 Flash / medium | 59.9 | 2051.8 | 2.915 | 49,532 |
| MyPCBench (38 tasks) | CUA-World (26 tasks) | |||||
|---|---|---|---|---|---|---|
| Effort | Score (%) | Time (s) | Cost ($) | Score (%) | Time (s) | Cost ($) |
| Low | 89.9 | 514 | 4.57 | 90.3 | 807 | 9.29 |
| Medium | 88.3 | 495 | 4.33 | 86.7 | 810 | 8.42 |
| High | 88.6 | 568 | 4.81 | 94.2 | 1,108 | 11.11 |
| Xhigh | 93.6 | 619 | 4.81 | 93.0 | 1,515 | 15.06 |
| Run | Score (%) | Time (s) | Agent (s) | Steps | Env. (s) | Waits (s) | Env. waits (s) |
|---|---|---|---|---|---|---|---|
| 1 | 87.62 | 86.11 | 72.90 | 5.90 | 13.209 | 1.328 | 11.881 |
| 2 | 91.62 | 89.56 | 75.31 | 5.82 | 14.248 | 1.082 | 13.166 |
| 3 | 91.62 | 90.13 | 76.52 | 5.88 | 13.606 | 1.040 | 12.566 |
| 4 | 91.62 | 89.81 | 75.78 | 5.78 | 14.033 | 1.152 | 12.881 |
| 5 | 91.62 | 92.10 | 78.19 | 5.94 | 13.914 | 1.012 | 12.902 |
| Mean | 90.82 | 89.54 | 75.74 | 5.86 | 13.802 | 1.123 | 12.679 |
| I/O | Pass@ (%) | Best-of- partial score (%) | |
|---|---|---|---|
| Standard I/O | 1 | 87.20 | 90.82 |
| 2 | 88.00 | 91.62 | |
| 3 | 88.00 | 91.62 | |
| 4 | 88.00 | 91.62 | |
| 5 | 88.00 | 91.62 | |
| Fast I/O | 1 | 86.80 | 90.42 |
| Benchmark | Full set | Selected subset |
|---|---|---|
| OSWorld ( Xie et al., 2024 ) | 295 | 50 |
| OSWorld2 ( Yuan et al., 2026 ) | 63 | 52 |
| CUA-World ( Aggarwal et al., 2026 ) | 143 | 26 |
| MyPCBench ( Jang et al., 2026a ) | 184 | 38 |
| Benchmark | Partial | Exact | Partial MAE | Exact MAE | |
|---|---|---|---|---|---|
| OSWorld | 49 | 0.967 | 0.967 | 4.36 | 4.57 |
| 50 | 0.983 | 0.975 | 2.16 | 2.17 | |
| 51 | 1.000 | 0.992 | 4.21 | 3.95 | |
| OSWorld2 | 51 | 0.964 | 0.955 | 1.70 | 2.28 |
| 52 | 0.964 | 0.955 | 1.33 | 2.03 | |
| 53 | 0.964 | 0.982 | 0.74 | 1.04 |
| Method | Partial MAE | Exact MAE | Mean MAE |
|---|---|---|---|
| Uniform random | 5.88 | 6.01 | 5.95 |
| Difficulty-stratified random | 5.05 | 5.11 | 5.08 |
| IRT-inspired selection | 7.94 | 8.21 | 8.08 |
| Energy-based selection | 4.50 | 4.34 | 4.42 |
| Application | Task UUID |
|---|---|
| chrome | 06fe7178-4491-4589-810f-2e2bc9502122 |
| chrome | 2ae9ba84-3a0d-4d4c-8338-3a1478dc5fe3 |
| chrome | 3720f614-37fd-4d04-8a6b-76f54f8c222d |
| chrome | 44ee5668-ecd5-4366-a6ce-c1c9b8d4e938 |
| chrome | af630914-714e-4a24-a7bb-f9af687d3b91 |
| gimp | 045bf3ff-9077-4b86-b483-a1040a949cff |
| Environment | Task name |
|---|---|
| ardour_env | broadcast_podcast_stem_delivery |
| dhis2_env | rmncah_scorecard_dashboard |
| docker_desktop_env | diagnose_broken_microservices_stack |
| gpredict_env | poes_downlink_schedule_setup |
| gvsig_desktop_env | vulnerability_map_remote_communities |
| jstock_env | quarterly_portfolio_rebalance |
| Reviewer or outcome | Included | Excluded |
|---|---|---|
| Reviewer 1 | 297 | 72 |
| Reviewer 2 | 302 | 67 |
| Reviewer 3 | 295 | 74 |
| Unanimous decision | 295 | 67 |
| Disagreement | 7 | |
| Requested input | Received input | Cause |
|---|---|---|
| Type < | > | Incorrect X11 key/modifier mapping |
| Type accented text | Accented characters omitted | Characters absent from the key map |
| Type a literal shell variable | Variable expanded into a path | Text interpreted by the command shell |
| Keypad Enter or Menu | No corresponding key event | Missing key-name translation |
| Toggle Caps Lock, then type | Lowercase text | Caps Lock press silently omitted |
| Harness | Model response | Executed behavior |
|---|---|---|
| Gemini | Scroll down by five wheel clicks | Scroll up by 600 ticks |
| Gemini | Press F5 or Page Down | Type the key name as text |
| Qwen3.5 | Middle-click the target | No action |
| Qwen3.5 | Ctrl-click the target | Release Ctrl before clicking |
| Qwen3.5 | Move, then scroll or drag | Execute only the first tool call |
| Action family | Operations |
|---|---|
| Mouse clicks | Left, right, middle, double, and triple click at a coordinate |
| Pointer motion | Move, drag along coordinates, hold or release a mouse button |
| Scrolling | Signed vertical scroll at the current pointer position |
| Keyboard | Type text, press a key or chord, hold and release modifiers |
| Waiting | Wait for an agent-specified duration |
| Observation | Request a screenshot without changing the desktop |
| Model | Effort | Interaction rule | Score | Time | Batches | Actions | Tokens |
|---|---|---|---|---|---|---|---|
| Astra | xhigh | Batched | 91.62 | 126.82 | 7.14 | 19.90 | 2,060 |
| Astra | xhigh | One action + image | 91.62 | 182.73 | 16.50 | 16.50 | 2,559 |
| Kimi K3 | low | Single tool | 71.62 | 176.61 | 8.12 | 55.10 | 3,171 |
| Kimi K3 | low | Batched tools | 73.62 | 149.50 | 5.74 | 30.00 | 1,865 |
| Kimi K3 | high | Single tool | 81.62 | 305.07 | 8.92 | 113.60 | 7,567 |
| Kimi K3 | high | Batched tools | 79.62 | 358.13 | 8.76 | 60.20 | 7,125 |
| Model | Tasks | Effort | Time | Batches | Actions | Tokens | Tok/s |
|---|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | 5 | low | 1291.0 | 116.40 | 171.60 | 29,596 | 47.57 |
| medium | 1363.9 | 112.60 | 158.40 | 34,446 | 43.10 | ||
| high | 3024.5 | 190.60 | 246.60 | 73,578 | 42.18 | ||
| GPT-6 Astra | 13 | low | 791.8 | 47.08 | 146.69 | 7,878 | 15.18 |
| medium | 752.6 | 47.15 | 138.85 | 8,376 | 17.07 | ||
| high | 817.3 | 53.15 | 171.23 | 11,001 | 21.31 |