AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition

    Oct 10, 2025Yushuo Zheng, Tongrui Ye, Zicheng Zhang +3Multimodal Large Language ModelsAI Agent Benchmarks

  2. HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

    Sep 10, 2025Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes +1LLM Agent EvaluationHuman-AI Interaction

  3. MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

    Aug 14, 2025Shilong Li, Xingyuan Bu, Wenjie Wang +22Multimodal IRMultimodal Reasoning

  4. Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents

    Aug 3, 2025Yuhan Guo, Cong Guo, Aiwen Sun +12Cognitive ModelingMultimodal Reasoning

  5. StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

    Jul 10, 2025Weihao Tan, Changjiu Jiang, Yu Duan +5Game-Playing AgentsAgent Evaluation

  6. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

    Jul 7, 2025Yuanzhe Hu, Yu Wang, Julian McAuleyLLM Agent MemoryLLM Agent Evaluation

  7. AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

    Jun 4, 2025Akshat Naik, Emma Gouné, Patrick Quinn +4AI AlignmentLLM Agent Evaluation

  8. AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance

    Jun 4, 2025Dhaval Patel, Shuxin Lin, James Rayfield +7LLM Agent OrchestrationAI Agent Evaluation

  9. SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments

    May 29, 2025Zixiang Xu, Yanbo Wang, Yue Huang +13LLM EvaluationAI Agent Benchmarks

  10. Towards LLM Agents for Earth Observation

    Apr 16, 2025Chia Hsiang Kao, Wenting Zhao, Cheryl Lam +10Code GenerationLLM Agent Evaluation

  11. GraphChase: A Platform and Benchmark for Urban Network Security Games

    Jan 29, 2025Shuxin Zhuang, Shuxin Li, Tianji Yang +4Benchmark DesignAI Agent Benchmarks

  12. How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?

    Dec 12, 2024Qiaoyu Zheng, Chaoyi Wu, Weike Zhao +5HealthcareLLM Agent Evaluation

  13. DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

    Date pendingRuizhe Li, Mingxuan Du, Benfeng Xu +3Deep Research AgentsLLM Agent Evaluation

  14. WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics

    Date pendingMadhav Kanda, Sharad Agarwal, Rodrigo Fonseca +2LLM Agent EvaluationAI Agent Benchmarks

  15. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

    Date pendingTara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz +10Voice Agent EvaluationSpoken Dialogue Systems

  16. GRACE-DS: a Guarded Reward-guided Agent Correction Environment in Data Science

    Date pendingAleksandr Tsymbalov, Danis Zaripov, Artem Epifanov +1LLM Agent EvaluationAI Agent Benchmarks

  17. OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

    Date pendingKaicheng Zhang, Wen Ge, Lei Jiang +5Quantitative FinanceLLM Agent Evaluation

  18. PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation

    Date pendingHyungseok Song, Junseok Park, Won-Seok Choi +4Electronic Design AutomationAI Agent Benchmarks

  19. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

    Date pendingWilliam CabanInter-Rater ReliabilityAI Agent Evaluation

  20. SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

    Date pendingYuqiao Tan, Shizhu He, Jun Zhao +1AI Agent EvaluationSparse Autoencoders

  21. Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

    Date pendingMinghao Guo, Meng Cao, Sui Zhao +10Deep Research AgentsLong-Horizon Agent Evaluation