AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists

    Oct 4, 2026Xiaxun Xie, Qingqing Long, Meng Xiao +4AI Agents for Scientific DiscoveryScientific Hypothesis Generation

  2. MetaKernelBench: Measuring GPU Kernel Knowledge Transfer Beyond Code

    Oct 4, 2026Xueyi Chen, Shiyu Liu, Xin Jin +5GPU Kernel OptimizationAI Agent Benchmarks

  3. EMBER-Bench: Benchmarking Cross-Event Causal Memory in Long-Horizon Embodied Tasks

    Oct 4, 2026Aoyang Cai, Boning Zhao, Shaoxuan Xie +5Next-Action PredictionEmbodied QA

  4. Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos

    Oct 3, 2026Jinzhou Tang, Zijun Zhang, Jing Yang +10Interactive World ModelsCoding Agents

  5. ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

    Oct 1, 2026Sohyeon Kim, Yoonho Lee, Bo Liu +11AI Agents for Scientific DiscoveryAgentic Search

  6. Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

    Oct 1, 2026Hao Wang, Ting HuangSmall Language ModelsAI Agent Benchmarks

  7. Agents Are Systems, Not Models: Rethinking Agentic Evaluation

    Oct 1, 2026Luis Wiedmann, Leander Girrbach, Cordelia Schmid +1AI Agent ReliabilityAI Agent Evaluation

  8. Revision-Aware Independent Agent Graphs for Dynamic Reasoning

    Oct 1, 2026Yan Luo, Selim-Antoine Lali, Jeremy Moebel +3Temporal ReasoningAI Agent Benchmarks

  9. Finding the Right Fit: Model-Harness Interactions across Agent Tasks

    Oct 1, 2026Yixuan Li, Yiyun Zhou, Yao Long Teng +6Tool-Augmented Language Model AgentsLLM Agent Evaluation

  10. Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena

    Oct 1, 2026Haojian Huang, Pukun Zhao, Zexi Li +7VLM EvaluationVisual Reasoning

  11. ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

    Sep 30, 2026Xisen Jin, Jingheng Li, Zhenglun Chen +2Continual Learning for LLM AgentsLong-Horizon Agent Evaluation

  12. Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

    Sep 30, 2026Michael Hardy, Ruhana Azam, Anka Reuel +2AI Agent EvaluationAgent Evaluation

  13. Incident-Arena: Getting agents to the last nine of reliability

    Sep 30, 2026Andre Fu, Malik Drabla, Leon Liu +5Agent ReliabilityAI Coding Agents

  14. CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

    Sep 30, 2026Chenmu Zhang, Levi Felix, Jun-Jie Zhang +6AI Agents for Scientific DiscoveryAI Agent Evaluation

  15. Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

    Sep 30, 2026Sahan Paliskara, Nattaput Namchittai, Andrew LampinenMulti-Agent LLM SystemsMulti-Agent Coordination

  16. EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

    Sep 30, 2026Jiayi Geng, Zhengxuan Wu, Kevin S. Chen +12Scientific DiscoveryLong-Horizon Agent Evaluation

  17. From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning

    Sep 30, 2026Minye Shao, Chaohui Yu, Yixuan Wu +3AI Agent BenchmarksMedical VLMs

  18. A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

    Sep 30, 2026Seonho Lee, Wonryeol Jeong, Alberto Cereser +4AI Coding AgentsAI Agent Benchmarks

  19. EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning

    Sep 30, 2026Yitong Qiao, Yancheng Jin, Lei Liu +4Evidence-Grounded ReasoningAI Agent Benchmarks

  20. RealWorldShop: Benchmarking and Improving Conversational Shopping Agents in Real-World E-commerce

    Sep 30, 2026Xinwei Yang, Kelong Mao, Yudong Guo +4LLM Agent EvaluationAI Agent Benchmarks