AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

    May 8, 2026Bohan Hou, Jiuning Gu, Jiayan Guo +5Multimodal IRMultimodal Search Agents

  2. Business Utility of Large Language Models as Exploratory Data Analysis Agents

    May 8, 2026Rafał Łabędzki, Patryk Miziuła, Hubert Rutkowski +5LLM Agent EvaluationAI Agent Benchmarks

  3. Tools as Continuous Flow for Evolving Agentic Reasoning

    May 8, 2026Tairan Huang, Siyu Shang, Qiang Chen +2Tool-Use PlanningAI Agent Benchmarks

  4. EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams

    May 8, 2026Dongchuan Ran, Linyu Ou, Xueheng Li +5Egocentric Video UnderstandingAI Agent Benchmarks

  5. Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning

    May 8, 2026Yuyang Wu, Yue Huang, Shuaike Shen +8AI Agents for Scientific DiscoveryLLM Agent Evaluation

  6. EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

    May 8, 2026Yi Liu, TingFeng Hui, Wei Zhang +4AI Agent EvaluationAI Agent Benchmarks

  7. SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

    May 8, 2026Jackson Clark, Yiming Su, Saad Mohammad Rafid Pial +5Software EngineeringAI Agent Benchmarks

  8. TeamBench: Evaluating Agent Coordination under Enforced Role Separation

    May 8, 2026Yubin Kim, Chanwoo Park, Taehan Kim +9Multi-Agent CoordinationLLM Agent Evaluation

  9. SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair

    May 7, 2026Ion George Dinu, Marian Cristian Mihăescu, Traian RebedeaAutomated Program RepairLLM Agent Evaluation

  10. Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

    May 7, 2026Roger Creus Castanyer, Pablo Samuel Castro, Glen BersethAI Agent EvaluationAI Agent Benchmarks

  11. Coordination Matters: Evaluation of Cooperative Multi-Agent Reinforcement Learning

    May 7, 2026Maria Ana Cardei, Matthew Landers, Afsaneh DoryabMulti-Agent CoordinationAI Agent Benchmarks

  12. STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?

    May 7, 2026Hanxiang Chao, Yihan Bai, Rui Sheng +2Agent MemoryLLM Agent Evaluation

  13. Stop Comparing LLM Agents Without Disclosing the Harness

    May 7, 2026Yunbei Zhang, Janet Wang, Yingqiang Ge +3LLM Agent EvaluationLong-Horizon Agent Evaluation

  14. More Than Can Be Said: A Benchmark and Framework for Pre-Question Scientific Ideation

    May 7, 2026Jie Yu, Song QiuAI for ScienceScientific Hypothesis Generation

  15. MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents

    May 7, 2026Ashwani Anand, Ivi Chatzi, Ritam Raha +1Computer-Use Agent BenchmarksAI Agent Reliability

  16. BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

    May 7, 2026Jinge Wu, Hongjian Zhou, Mingde Zeng +8AI Agents for Scientific DiscoveryDeep Research Agents

  17. SODE: Analyzing Social Dynamics in LLM Agents

    May 6, 2026Inseo Jung, Yoonseok Oh, Kyungryul Back +2Multi-Agent LLM SystemsLLM Agent Evaluation

  18. Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

    May 5, 2026Zirui Tang, Xuanhe Zhou, Yumou Liu +19Computer-Use Agent BenchmarksAI Agent Evaluation

  19. SOTOPIA-TOM: Evaluating Privacy and Information Management in Multi-Agent Interaction with Theory of Mind

    May 4, 2026Yashwanth YS, Ruichen Wang, Shihua Zeng +4Multi-Agent CoordinationTheory of Mind

  20. Towards Understanding Specification Gaming in Reasoning Models

    May 4, 2026Kei Nishimura-Gasparian, Robert McCarthy, David LindnerReward HackingRL for Language Model Reasoning

  21. PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

    May 4, 2026Ruoqi Liu, Imran Q. Mohiuddin, Austin J. Schoeffler +10HealthcareLong-Horizon Agent Evaluation

  22. NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

    May 3, 2026Xiao JiaAI Agent ReliabilityLLM Agent Evaluation

  23. LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation

    May 2, 2026Dong Xu, Jialun Cao, Guozhao Mo +9LLM Agent EvaluationFormal Verification

  24. ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue

    May 2, 2026Daoxuan Zhang, Ping Chen, Jianyi Zhou +1Aerial RoboticsRobotics Simulation