AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

    Sep 23, 2026Jingjie Ning, Xueqi Li, Yibo Kong +1AI Agent EvaluationActive Learning

  2. MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

    Sep 23, 2026Yongjun Jeong, Hanbum Ko, Ye Rin Kim +6Molecular OptimizationLLM Agent Evaluation

  3. CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

    Sep 23, 2026Yuxuan Li, Will Epperson, Wesley Deng +1Computer-Use AgentsAgent Harness Optimization

  4. SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

    Sep 22, 2026Jennifer Williams, Dave Farris, Jeff Farris +1Software EngineeringLLM Inference Serving

  5. The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

    Sep 22, 2026Wenbo Pan, Zhichao Liu, Shujie Liu +6Software Engineering BenchmarksLong-Horizon Agent Evaluation

  6. Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development

    Sep 21, 2026Haocheng Xia, Eugene Wu, Yongjoo ParkLLM Agent EvaluationAI Agent Benchmarks

  7. The AI Neuroscientist: An Interactive Agentic Interface for Neuroimaging Analysis

    Sep 21, 2026Aakash Patel, Panos Ketonis, Shreya Saxena +2AI Agents for Scientific DiscoveryLLM Agent Evaluation

  8. Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

    Sep 21, 2026Weihang Ding, Junfei ZhanLLM Agent EvaluationAI Agent Benchmarks

  9. FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

    Sep 21, 2026Wenqing Wang, Haitao Xiang, Xinyi Zhao +8Financial QAAI Agent Benchmarks

  10. MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents

    Sep 21, 2026Demetris Paschalides, Moysis Symeonides, George Pallis +1AI Agent BenchmarksTool-Using Agents

  11. XYEval: Agents say yes to bad advice

    Sep 20, 2026Zhengxuan Wu, Yuxuan Li, Oyvind Tafjord +1LLM SycophancyAI Agent Evaluation

  12. BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

    Sep 20, 2026Peng Kuang, Yuchun Fan, Jiangnan Li +7Multilingual Language Model EvaluationLLM Agent Evaluation

  13. F2^{2}DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

    Sep 17, 2026Bojian Xiong, Wentao Ding, Yujing Lu +11Reward ModelingLLM Evaluation

  14. SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership

    Sep 17, 2026Run Peng, Zinnia Nie, Jing Ding +7Human Behavior SimulationLong-Horizon Agent Evaluation

  15. Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

    Sep 16, 2026Xinshuai Guo, Junjie Wu, Dolly Deng +4AI Agent EvaluationAI Agent Benchmarks

  16. Clueing up LLMs with Tool-Augmented Deductive Reasoning

    Sep 16, 2026Rebecca Ansell, Autumn Toney-WailsMulti-Agent LLM SystemsAI Agent Benchmarks

  17. Locating Hidden Failures Makes Long-Horizon Agents More Reliable

    Sep 15, 2026Salman Rahman, Yubin Kim, Mihir Parmar +15Agent Failure AnalysisAI Agent Reliability

  18. GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

    Sep 15, 2026Sikun Wang, Yixi Zhou, Lei Fan +1LLM Reasoning with GraphsLLM Agent Evaluation

  19. ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures

    Sep 15, 2026Muhammad Ashar Ishfaq, Glaucia MeloTheory of MindLLM Agent Evaluation

  20. V-ICAL Bench: Evaluating Video In-Context Learning for Multimodal Agents in Interactive Environments

    Sep 14, 2026Ziqian Fan, Shibo Xu, Junjie Li +8Multimodal ICLAI Agent Benchmarks

  21. BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

    Sep 14, 2026Yolo Y. Tang, Daiki Shimada, Jiayue Meng +14AI Agent BenchmarksTemporal Video Understanding

  22. When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

    Sep 14, 2026Kaiyuan Liu, Qiuyang Mang, Bo Peng +6LLM Agent EvaluationTest-Time Scaling