AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

    May 13, 2026Yuxiang Lai, Peng Xia, Haonian Ji +8Computer-Use Agent BenchmarksLLM Agent Evaluation

  2. Reinforcement Learning for Tool-Calling Agents in Fast Healthcare Interoperability Resources (FHIR)

    May 13, 2026Marius S. Knorr, Robert Müller, Jan P. Bremer +1HealthcareRL for Language Model Reasoning

  3. PBT-Bench: Benchmarking AI Agents on Property-Based Testing

    May 13, 2026Lucas Jing, Xinqi Wang, Liao Zhang +1Automated Software TestingSoftware Engineering Agents

  4. Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction

    May 13, 2026Darius A. Faroughy, Sofia Palacios Schweitzer, Ian Pang +2AI Agents for Scientific DiscoveryAI Agent Evaluation

  5. Turning Intent into Specifications: A Benchmark and an Interactive User-Assistant Agent

    May 13, 2026Hao Wang, Ligong Han, Kai Xu +1Requirements EngineeringSoftware Engineering Agents

  6. Evaluating Memory Condensation Strategies for Coding Agents in Data-Driven Scientific Discovery

    May 13, 2026Renuka Chintalapati, Sid Raskar, Anurag Acharya +3AI Coding AgentsAI Agent Benchmarks

  7. Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

    May 13, 2026Qinchuan Cheng, Zhantao Gong, Pengzhan Sun +3Belief-Space PlanningLong-Horizon Agent Evaluation

  8. Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue

    May 13, 2026Vardhan Dongre, Dilek Hakkani-TürMulti-Agent CoordinationAI Agent Benchmarks

  9. Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse

    May 12, 2026Ling-Qi Zhang, Kristin BransonAI Coding AgentsAI Agent Benchmarks

  10. No More, No Less: Task Alignment in Terminal Agents

    May 12, 2026Sina Mavali, David Pape, Jonathan Evertz +5Computer-Use Agent BenchmarksAI Alignment

  11. Rollout Cards: A Reproducibility Standard for Agent Research

    May 12, 2026Charlie Masters, Ziyuan Liu, Stefano V. AlbrechtAI Agent EvaluationAgent Evaluation

  12. When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents

    May 12, 2026Xiaolin Zhou, Aojie Yuan, Zheng Luo +12Domain RandomizationAI Agent Benchmarks

  13. Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations

    May 12, 2026Junjue Wang, Weihao Xuan, Heli Qi +7LLM Agent EvaluationGeospatial Reasoning

  14. PRISM: : Planning and Reasoning with Intent in Simulated Embodied Environments

    May 12, 2026Yunn Kang Lim, Pengzhan Sun, Ziyi Bai +4Long-Horizon Agent EvaluationAI Agent Benchmarks

  15. An Empirical Study of Automating Agent Evaluation

    May 12, 2026Kang Zhou, Sangmin Woo, Haibo Ding +15Agent EvaluationAI Agent Benchmarks

  16. gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy

    May 11, 2026Tousif Islam, Digvijay Wadekar, Zihan ZhouGravitational-Wave AstronomyAI Coding Agents

  17. ABRA: Agent Benchmark for Radiology Applications

    May 11, 2026Bulat Maksudov, Vladislav Kurenkov, Kathleen M. Curran +1Radiology Report GenerationHealthcare

  18. Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?

    May 11, 2026Maximilian Triebel, Marco Menner, Dominik HelfensteinVLM EvaluationPhysical Reasoning

  19. WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

    May 11, 2026Shuangrui Ding, Xuanlang Dai, Long Xing +14Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  20. AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

    May 11, 2026Edward De Brouwer, Carl Edwards, Alexander Wu +9LLM EvaluationAI Agent Benchmarks

  21. MaD Physics: Evaluating information seeking under constraints in physical environments

    May 11, 2026Moksh Jain, Mehdi Bennani, Johannes Bausch +4Scientific ReasoningAI Agent Benchmarks

  22. ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

    May 11, 2026Yuanyang Li, Xue Yang, Longyue Wang +2LLM Agent EvaluationAI Agent Benchmarks