AI Agent Benchmarks

Momentum

209 papers in the last four weeks, up 61% on the four weeks before. 1.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,159

All topics
CardsList
  1. ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences

    Feb 11, 2026Bang Nguyen, Dominik Soós, Qian Ma +8LLM Agent EvaluationAI Agent Benchmarks

  2. GameDevBench: Evaluating Agentic Capabilities Through Game Development

    Feb 11, 2026Wayne Chi, Yixiong Fang, Arnav Yayavaram +8Software Engineering AgentsAgent Evaluation

  3. Evaluating Memory Structure in LLM Agents

    Feb 11, 2026Alina Shutova, Alexandra Olenina, Ivan Vinogradov +1LLM Agent MemoryLLM Agent Evaluation

  4. ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response

    Feb 3, 2026Xiaomeng Zhu, Fengming Zhu, Weijie Zhou +8AI Agent BenchmarksProactive Assistance

  5. Declarative by Design, Assistable Only by Convention: Benchmarking Multi-Agent Frameworks for AI-Assistability

    Feb 3, 2026Shafiuddin Rehan Ahmed, Sourabh DeshpandeAI Agent BenchmarksCode Generation Evaluation

  6. AgentRx: Diagnosing AI Agent Failures from Execution Trajectories

    Feb 2, 2026Shraddha Barke, Arnav Goyal, Alind Khare +3Agent Failure AnalysisAI Agent Reliability

  7. Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

    Jan 30, 2026Yuhao Zhan, Tianyu Fan, Linxuan Huang +2Deep Research AgentsHallucination Detection

  8. EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents

    Jan 23, 2026Xinze Li, Ziyue Zhu, Siyuan Liu +4Episodic MemoryLong-Horizon Agent Evaluation

  9. Toward Efficient Agents: Memory, Tool learning, and Planning

    Jan 20, 2026Xiaofang Yang, Lijun Li, Heng Zhou +12LLM Agent MemoryAI Agent Benchmarks

  10. LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning

    Jan 20, 2026Ye Tian, Zihao Wang, Onat Gungor +2HealthcareLong-Horizon Agent Evaluation

  11. DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems

    Jan 20, 2026Maojun Sun, Yifei Xie, Yue Wu +5LLM Agent EvaluationAgent Evaluation

  12. AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks

    Jan 16, 2026Weiyi Wang, Xinchi Chen, Jingjing Gong +2LLM Agent EvaluationAI Agent Benchmarks

  13. IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents

    Jan 10, 2026Yingchaojie Feng, Qiang Huang, Xiaoya Xie +4Deep Research AgentsLLM Agent Evaluation

  14. DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation

    Dec 19, 2025Janghoon Han, Heegyu Kim, Changho Lee +6Deep Research AgentsLLM Agent Evaluation

  15. PPTArena: A Benchmark for PowerPoint Editing

    Dec 2, 2025Michael Ofengenden, Yunze Man, Ziqi Pang +2LLM Agent EvaluationAI Agent Benchmarks

  16. CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows

    Nov 25, 2025Chenyue Li, Hyeonjae Kim, Wen Deng +4AI for ScienceMulti-Agent Orchestration

  17. Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

    Nov 24, 2025Dayong Liu, Chao Xu, Weihong Chen +5AI Agent EvaluationMultimodal Large Language Models

  18. IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

    Nov 21, 2025Yifan Li, Lichi Li, Anh Dao +15Visual Spatial ReasoningEmbodied Navigation

  19. SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

    Nov 20, 2025Juntao Cheng, Wanyue Zhang, Zhiwei Yu +7Long-Horizon Agent EvaluationEgocentric Video Understanding

  20. CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

    Nov 4, 2025Jiayu Liu, Cheng Qian, Zhaochen Su +4LLM PlanningLLM Agent Evaluation

  21. When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

    Oct 22, 2025Tanya Shourya, Yingfan Wang, Zhaoyi Joey Hou +3Tool-Use EvaluationAI Agent Benchmarks

  22. Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

    Oct 21, 2025Howard Yen, Yoonsang Lee, Ashwin Paranjape +5Agentic SearchLong-Horizon Agent Evaluation

  23. HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning

    Oct 16, 2025Chance Jiajie Li, Zhenze Mo, Yuhan Tang +11Cognitive ModelingHuman Behavior Simulation

  24. Compositional Machine Design as Program Synthesis with LLMs

    Oct 16, 2025Wenqian Zhang, Yangyi Huang, Weiyang Liu +1LLM-Based Program SynthesisProgram Synthesis

  25. LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition

    Oct 10, 2025Yushuo Zheng, Tongrui Ye, Zicheng Zhang +3Multimodal Large Language ModelsAI Agent Benchmarks

  26. HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

    Sep 10, 2025Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes +1LLM Agent EvaluationHuman-AI Interaction