AI Agent Benchmarks

Momentum

209 papers in the last four weeks, up 61% on the four weeks before. 1.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,159

All topics
CardsList
  1. MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

    Aug 14, 2025Shilong Li, Xingyuan Bu, Wenjie Wang +22Multimodal IRMultimodal Reasoning

  2. Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents

    Aug 3, 2025Yuhan Guo, Cong Guo, Aiwen Sun +12Cognitive ModelingMultimodal Reasoning

  3. StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

    Jul 10, 2025Weihao Tan, Changjiu Jiang, Yu Duan +5Game-Playing AgentsAgent Evaluation

  4. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

    Jul 7, 2025Yuanzhe Hu, Yu Wang, Julian McAuleyLLM Agent MemoryLLM Agent Evaluation

  5. AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

    Jun 4, 2025Akshat Naik, Emma Gouné, Patrick Quinn +4AI AlignmentLLM Agent Evaluation

  6. AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance

    Jun 4, 2025Dhaval Patel, Shuxin Lin, James Rayfield +7LLM Agent OrchestrationAI Agent Evaluation

  7. SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments

    May 29, 2025Zixiang Xu, Yanbo Wang, Yue Huang +13LLM EvaluationAI Agent Benchmarks

  8. Towards LLM Agents for Earth Observation

    Apr 16, 2025Chia Hsiang Kao, Wenting Zhao, Cheryl Lam +10Code GenerationLLM Agent Evaluation

  9. GraphChase: A Platform and Benchmark for Urban Network Security Games

    Jan 29, 2025Shuxin Zhuang, Shuxin Li, Tianji Yang +4Benchmark DesignAI Agent Benchmarks

  10. How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?

    Dec 12, 2024Qiaoyu Zheng, Chaoyi Wu, Weike Zhao +5HealthcareLLM Agent Evaluation

  11. DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

    Date pendingRuizhe Li, Mingxuan Du, Benfeng Xu +3Deep Research AgentsLLM Agent Evaluation

  12. WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics

    Date pendingMadhav Kanda, Sharad Agarwal, Rodrigo Fonseca +2LLM Agent EvaluationAI Agent Benchmarks

  13. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

    Date pendingTara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz +10Voice Agent EvaluationSpoken Dialogue Systems

  14. GRACE-DS: a Guarded Reward-guided Agent Correction Environment in Data Science

    Date pendingAleksandr Tsymbalov, Danis Zaripov, Artem Epifanov +1LLM Agent EvaluationAI Agent Benchmarks

  15. OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

    Date pendingKaicheng Zhang, Wen Ge, Lei Jiang +5Quantitative FinanceLLM Agent Evaluation

  16. PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation

    Date pendingHyungseok Song, Junseok Park, Won-Seok Choi +4Electronic Design AutomationAI Agent Benchmarks

  17. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

    Date pendingWilliam CabanInter-Rater ReliabilityAI Agent Evaluation