AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

    May 26, 2026Anmol Agarwal, Natalie Neamtu, Pranjal Aggarwal +6LLM Agent EvaluationFormal Verification

  2. JobBench: Aligning Agent Work With Human Will

    May 25, 2026Yuetai Li, Yichen Feng, Zhangchen Xu +21Computer-Use Agent BenchmarksLLM Agent Evaluation

  3. Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

    May 25, 2026Jianing Zhu, Yeonju Ro, John Robertson +5AI Agent ReliabilityAI Agent Benchmarks

  4. Sentinel: Embodied Cooperative Spatial Reasoning and Planning

    May 25, 2026Xiangye Lin, Hongxin Zhang, Ruxi Deng +2AI Agent BenchmarksMulti-Agent Planning

  5. MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

    May 25, 2026Dingbang Wu, Rui Hao, Haiyang Wang +8Mobile GUI AutomationGUI Agents

  6. DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

    May 25, 2026Matt L. Wiemann, Lindsay M. Smith, Peter Melchior +4Physical ReasoningScientific Discovery

  7. Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

    May 25, 2026Yusong Lin, Xinyuan Liang, Haiyang Wang +8Computer-Use Agent BenchmarksAI Agent Benchmarks

  8. Automated Benchmark Auditing for AI Agents and Large Language Models

    May 25, 2026Junlin Wang, Federico Bianchi, Shang Zhu +4LLM EvaluationBenchmark Auditing

  9. CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

    May 25, 2026Junlin Yang, Dylan Zhang, Xiangchen Song +7Structural Causal ModelsLLM Agent Evaluation

  10. Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

    May 25, 2026Haoyi Hu, Qirong Lyu, Xianghan Kong +7AI Agent EvaluationAI Agent Benchmarks

  11. AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

    May 25, 2026Jingwei Sun, Jianing Zhu, Yuanyi Li +3Computer-Use Agent BenchmarksComputer-Use Agents

  12. Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents

    May 25, 2026Hao-Hsuan ChenAI Risk ManagementAI Agent Evaluation

  13. CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

    May 25, 2026Bowen Wang, Dunjie Lu, Junli Wang +11Computer-Use AgentsSynthetic Data Generation

  14. GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

    May 24, 2026Xiang Cheng, Yulan Hu, Lulu Zheng +4AI-Assisted Decision MakingLLM Planning

  15. AION: Next-Generation Tasks and Practical Harness for Time Series

    May 24, 2026Tianxiang Zhan, Xiaobao Song, Tong Guan +2Time Series ReasoningTime Series Forecasting

  16. Memory-Induced Tool-Drift in LLM Agents

    May 24, 2026Mahavir Dabas, Jihyun Jeong, Ming Jin +1AI Agent BenchmarksLLM Tool Use

  17. The Open Source Economic Index of AI Adoption and Capability

    May 23, 2026Seamus Somerstep, Aritra Guha, Divesh Srivastava +1AI Agent EvaluationAI Agent Benchmarks

  18. ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

    May 22, 2026Xianzhong Ding, Yangyang Yu, Changwei Liu +1Tool-Augmented Language Model AgentsAI Coding Agents

  19. AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

    May 22, 2026Darek Kleczek, Fuheng Zhao, Alexander W. Lee +4AI Agent EvaluationAI Agent Benchmarks

  20. SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

    May 22, 2026Yingtie Lei, Zhongwei Wan, Jiankun Zhang +13LLM Agent EvaluationLLM Agent Skill Learning

  21. EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions

    May 22, 2026Haiyang Shen, Xuanzhong Chen, Wendong Xu +3AI Coding AgentsSoftware Engineering Benchmarks

  22. OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

    May 22, 2026Jiahao Ying, Boxian Ai, Wei Tang +2LLM Agent EvaluationAI Agent Benchmarks

  23. GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models

    May 22, 2026Vartan Shadarevian, Kia Ghods, Alex Kenich +1LLM EvaluationImperfect-Information Games

  24. MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

    May 21, 2026Thomson Yen, Julian Poeltl, Harshith Srinivas Gear +9LLM Agent EvaluationAI Agent Benchmarks