Web Agent Benchmarks

Latest papers 33

All topics
CardsList
  1. Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

    Oct 1, 2026Chengguang Gan, Zimeng He, Yoshihiro Tsujii +3AI Agent AuditingWeb Agent Benchmarks

  2. WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents

    Sep 28, 2026Anton Emelyanov, Maria Tikhonova, Zaven Martirosian +2Web Agent BenchmarksGUI Agents

  3. CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

    Aug 9, 2026Zejun Xu, Taiyi Chen, Jin Li +13Computer-Use AgentsWeb Agent Benchmarks

  4. Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents

    Aug 6, 2026Jiaming Wei, Zekun Wu, Adriano Koshiyama +1Web Agent BenchmarksWeb Agents

  5. The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory

    Jul 12, 2026Yixiong Chen, Xinyi Bai, Alan YuilleWeb Agent BenchmarksLong-Horizon Agent Evaluation

  6. MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

    Jul 11, 2026Chengguang Gan, Hanjun Wei, Yunhao Liang +3Computer-Use AgentsWeb Agent Benchmarks

  7. WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

    Jul 7, 2026Wei Dong, Tianyu Fu, Zhe Yu +9LLM-as-a-JudgeWeb Agent Benchmarks

  8. When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration

    Jun 16, 2026Aagam Sogani, Botao Rui, Swetha Vaidyanathan +3Agent Failure AnalysisWeb Agent Benchmarks

  9. LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings

    Jun 16, 2026Yi Zhao, Zhen Yang, Mengpan Chen +5VLM EvaluationWeb Application Generation

  10. MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

    Jun 15, 2026Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh +1Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  11. Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents

    Jun 12, 2026Sina Hajimiri, Masih Aminbeidokhti, Jose Dolz +4Web Agent BenchmarksLLM Agent Evaluation

  12. Running the Gauntlet: Hard Agentic Tasks

    Jun 12, 2026Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +203D Spatial ReasoningWeb Agent Benchmarks

  13. SentinelBench: A Benchmark for Long-Running Monitoring Agents

    Jun 3, 2026Matheus Kunzler Maldaner, Adam Fourney, Amanda Swearngin +5Computer-Use Agent BenchmarksWeb Agent Benchmarks

  14. K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

    Jun 1, 2026Nahyun Lee, Dongkeun Yoon, Guijin Son +12Computer-Use Agent BenchmarksLLM Evaluation

  15. GTA: Generating Long-Horizon Tasks for Web Agents at Scale

    May 28, 2026Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey +4Computer-Use Agent BenchmarksWeb Agent Benchmarks

  16. VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

    May 22, 2026JunJia Guo, Yuhang Yao, Jiawei +2Web Application GenerationWeb Agent Benchmarks

  17. ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents

    May 15, 2026Chinmay Savadikar, Mingyu Zhao, Yuanzheng Zhu +5Benchmark DesignWeb Agent Benchmarks

  18. An Executable Benchmarking Suite for Tool-Using Agents

    May 10, 2026Zhiqing Zhong, Zhijing Ye, Jiamin Wang +1Computer-Use Agent BenchmarksTool-Use Evaluation

  19. InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?

    Apr 30, 2026Qiyao Wang, Haoran Hu, Longze Chen +4Web Application GenerationWeb Agent Benchmarks

  20. Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

    Apr 27, 2026Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried +1Computer-Use Agent BenchmarksWeb Agent Benchmarks

  21. Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces

    Apr 26, 2026Zijing Shi, Meng Fang, Ling ChenWeb Agent BenchmarksAI Agent Safety

  22. MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation

    Apr 16, 2026Yan Li, Zezi Zeng, Yifan Yang +12Web Application GenerationWeb Agent Benchmarks

  23. Do Web Agents Investigate Before They Decide?

    Feb 5, 2026Syed Nazmus Sakib, Nafiul Haque, Tapodhir Karmakar Taton +2Web Agent BenchmarksHallucination in Language Models

  24. It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

    Dec 29, 2025Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki +7Web Agent BenchmarksIndirect Prompt Injection

  25. WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks

    Jul 1, 2025Zihao Sun, Zijing Shi, Ling ChenWeb Agent BenchmarksMultimodal Agents