AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. VISTA: A Generative Egocentric Video Framework for Daily Assistance

    May 11, 2026Yu-Hsiang Liu, Yu-Chien Tang, An-Zi YenSynthetic Data GenerationAI Agent Benchmarks

  2. Agentic Performance at the Edge: Insights from Benchmarking

    May 11, 2026Shiqiang Wang, Herbert WoisetschlägerLLM Agent EvaluationAI Agent Evaluation

  3. Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

    May 11, 2026Haonan Dong, Qiguan Feng, Kehan Jiang +3AI AlignmentAI Agent Evaluation

  4. SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

    May 11, 2026Zonglin Yang, Xingtong Liu, Xinyan XuAI Agents for Scientific DiscoveryAI Agent Evaluation

  5. Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse

    May 11, 2026Kuan Zhang, Dongchen Liu, Qiyue Zhao +12Game-Playing AgentsAI Agent Benchmarks

  6. EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents

    May 11, 2026Gurusha Juneja, Dylan Lu, Saaket Agashe +7Multi-Agent CoordinationTheory of Mind

  7. An Executable Benchmarking Suite for Tool-Using Agents

    May 10, 2026Zhiqing Zhong, Zhijing Ye, Jiamin Wang +1Computer-Use Agent BenchmarksTool-Use Evaluation

  8. Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents

    May 10, 2026Josefa Lia Stoisser, Marc Boubnovski Martell, Sidsel Boldsen +2LLM Agent EvaluationAI Agent Benchmarks

  9. DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

    May 10, 2026Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi +7VLM EvaluationTool Use in VLMs

  10. The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory

    May 10, 2026Luoxi Tang, Rupali Rajendra Vaje, Yuqiao Meng +5AI Agent BenchmarksLLM Agent Reliability

  11. MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions

    May 10, 2026Yaoning Yu, Ye Yu, Haojing Luo +1Human Behavior SimulationLLM Agent Evaluation

  12. LLM Agents Already Know When to Call Tools -- Even Without Reasoning

    May 10, 2026Chung-En Sun, Linbo Liu, Ge Yan +2Tool-Augmented Language Model AgentsAI Agent Benchmarks

  13. MDGYM: Benchmarking AI Agents on Molecular Simulations

    May 9, 2026Vinay Kumar, Satyendra Rajput, Mausam +1AI Agents for Scientific DiscoveryAI Agent Evaluation

  14. OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

    May 9, 2026Xiaozhe Li, Jixuan Chen, Xinyu Fang +4Large Language Model-Guided OptimizationLLM Agent Self-Improvement

  15. MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

    May 9, 2026Bohan Lyu, Yucheng Yang, Siqiao Huang +25AI Agents for Scientific DiscoveryAutomated Algorithm Discovery

  16. AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

    May 9, 2026Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia +10AI Agent ReliabilityMulti-Agent Collaboration

  17. Log analysis is necessary for credible evaluation of AI agents

    May 8, 2026Peter Kirgis, Sayash Kapoor, Stephan Rabanser +8Agent Failure AnalysisAI Agent Evaluation

  18. Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge

    May 8, 2026Dhaval Patel, Chathurangi Shyalika, Suryanarayana Reddy Yarrabothula +4Competitive AnalysisMulti-Agent Orchestration

  19. Learning CLI Agents with Structured Action Credit under Selective Observation

    May 8, 2026Haoyang Su, Ying WenComputer-Use AgentsCredit Assignment in RL

  20. AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents

    May 8, 2026Zhengkang Guo, Yiyang Li, Lin Qiu +7LLM Agent EvaluationLong-Horizon Agent Evaluation

  21. DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

    May 8, 2026Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang +2LLM Agent EvaluationAI Agent Benchmarks

  22. MMTB: Evaluating Terminal Agents on Multimedia-File Tasks

    May 8, 2026Chiyeong Heo, Jaechang Kim, Junhyuk Kwon +4Computer-Use Agent BenchmarksAudio-Visual Understanding