Computer-Use Agent Benchmarks

Latest papers 91

All topics
CardsList
  1. JobBench: Aligning Agent Work With Human Will

    May 25, 2026Yuetai Li, Yichen Feng, Zhangchen Xu +21Computer-Use Agent BenchmarksLLM Agent Evaluation

  2. Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

    May 25, 2026Yusong Lin, Xinyuan Liang, Haiyang Wang +8Computer-Use Agent BenchmarksAI Agent Benchmarks

  3. AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

    May 25, 2026Jingwei Sun, Jianing Zhu, Yuanyi Li +3Computer-Use Agent BenchmarksComputer-Use Agents

  4. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11LLM Safety BenchmarksComputer-Use Agent Benchmarks

  5. TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

    May 21, 2026Zhaoyang Chu, Jiarui Hu, Xingyu Jiang +8Computer-Use Agent BenchmarksTerminal Agents

  6. AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

    May 19, 2026Parsa Mazaheri, Kasra MazaheriAgent Failure AnalysisComputer-Use Agent Benchmarks

  7. OpenComputer: Verifiable Software Worlds for Computer-Use Agents

    May 19, 2026Jinbiao Wei, Qianran Ma, Yilun Zhao +4Computer-Use Agent BenchmarksComputer-Use Agents

  8. Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks

    May 18, 2026Yubin Qu, Ying Zhang, Yanjun Zhang +4Computer-Use Agent BenchmarksCoding Agents

  9. ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

    May 13, 2026Yuxiang Lai, Peng Xia, Haonian Ji +8Computer-Use Agent BenchmarksLLM Agent Evaluation

  10. Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

    May 12, 2026Miaosen Zhang, Xiaohan Zhao, Zhihong Tan +14Computer-Use Agent BenchmarksComputer-Use Agents

  11. No More, No Less: Task Alignment in Terminal Agents

    May 12, 2026Sina Mavali, David Pape, Jonathan Evertz +5Computer-Use Agent BenchmarksAI Alignment

  12. WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

    May 11, 2026Shuangrui Ding, Xuanlang Dai, Long Xing +14Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  13. An Executable Benchmarking Suite for Tool-Using Agents

    May 10, 2026Zhiqing Zhong, Zhijing Ye, Jiamin Wang +1Computer-Use Agent BenchmarksTool-Use Evaluation

  14. MMTB: Evaluating Terminal Agents on Multimedia-File Tasks

    May 8, 2026Chiyeong Heo, Jaechang Kim, Junhyuk Kwon +4Computer-Use Agent BenchmarksAudio-Visual Understanding

  15. Computer Use at the Edge of the Statistical Precipice

    May 7, 2026Pierluca D'Oro, Sneha Silwal, William Wong +6Computer-Use Agent BenchmarksComputer-Use Agents

  16. MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents

    May 7, 2026Ashwani Anand, Ivi Chatzi, Ritam Raha +1Computer-Use Agent BenchmarksAI Agent Reliability

  17. Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

    May 5, 2026Zirui Tang, Xuanhe Zhou, Yumou Liu +19Computer-Use Agent BenchmarksAI Agent Evaluation

  18. AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

    May 1, 2026Ranit Karmakar, Jayita ChatterjeeComputer-Use Agent BenchmarksLLM Agent Evaluation

  19. WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments

    Apr 30, 2026Jinchao Li, Yunxin Li, Chenrui Zhao +3Computer-Use Agent BenchmarksComputer-Use Agents

  20. Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

    Apr 27, 2026Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried +1Computer-Use Agent BenchmarksWeb Agent Benchmarks

  21. AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

    Apr 3, 2026Yunhao Feng, Yifan Ding, Yingshui Tan +6LLM Safety BenchmarksComputer-Use Agent Benchmarks

  22. LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

    Mar 20, 2026Xiang Long, Li Du, Yilong Xu +11Computer-Use Agent BenchmarksTool-Augmented Language Model Agents