AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

    Jul 18, 2026Mihir Shriniwas AryaCompositional ReasoningLLM Agent Memory

  2. CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents

    Jul 18, 2026Peilong Zhou, Zhirong Chen, Cangyuan Li +4Benchmark DesignElectronic Design Automation

  3. Interactive Task Alignment as a POMDP

    Jul 17, 2026Andy Dai, Zexue He, Zhenyu Zhang +2Pragmatic Reasoning in Language ModelsAI Agent Benchmarks

  4. Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents

    Jul 17, 2026Zedong Yu, Qianxing Li, Zhi Gao +8Computer-Use Agent BenchmarksGUI Agents

  5. PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

    Jul 16, 2026Bowen Jiang, Yuan Yuan, Zhuoqun Hao +12Personal AI AgentsLLM Personalization

  6. BrainPilot: Automating Brain Discovery with Agentic Research

    Jul 16, 2026Haoxuan Li, Tianci Gao, Jianhe Li +13AI for ScienceLLM Agent Orchestration

  7. OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

    Jul 16, 2026Chengyu Shen, Yujie Fu, Gangtao Xin +13Computer-Use Agent BenchmarksAI Agent Evaluation

  8. MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers

    Jul 16, 2026Huanxi Liu, Kun Hu, Jiaqi Liao +6Tool-Use EvaluationAI Agent Benchmarks

  9. DeepStress: Stress-Testing Deep Search Agents

    Jul 15, 2026Ismael Rousseau, Geraldine Damnati, Frederic BechetAI Agent ReliabilityAI Agent Evaluation

  10. EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent

    Jul 15, 2026Junlong Li, Junxi Li, Yuxiang Yang +3Egocentric Video QAVideo QA

  11. AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

    Jul 15, 2026Kai Chen, Zichen Ding, Jiaye Ge +20LLM Agent EvaluationAI Agent Benchmarks

  12. DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

    Jul 15, 2026Huatao Li, Xinwei Geng, Yuheng Wang +9Computer-Use Agent BenchmarksLLM Agent Evaluation

  13. Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals

    Jul 14, 2026Mohotarema Rashid, Lingzi Hong, Junhua Ding +1Multi-Agent LLM SystemsLLM Agent Evaluation

  14. RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls

    Jul 13, 2026Brenda Lelis, Rodrigo Cabral-CarvalhoMulti-Agent CoordinationAI Agent Benchmarks

  15. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

    Jul 13, 2026Chunzheng Zhu, Lei Tian, Bohan Tan +15HealthcareLLM Agent Self-Improvement

  16. The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

    Jul 13, 2026Chenglin Yu, Hongquan Gui, Ying Yu +3LLM Agent EvaluationAI Agent Benchmarks

  17. BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

    Jul 13, 2026Yuzhe Guo, Mengzhou Wu, Yuan Cao +4Tool-Augmented Language Model AgentsLanguage Model Generation Evaluation

  18. Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems

    Jul 12, 2026Yongchang Fu, Xinjie Huang, Chengjun Dai +3Large Language Model-Guided OptimizationLLM Agent Evaluation

  19. The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory

    Jul 12, 2026Yixiong Chen, Xinyi Bai, Alan YuilleWeb Agent BenchmarksLong-Horizon Agent Evaluation

  20. UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp

    Jul 12, 2026Xiyu Wei, Qingwei Zong, Zhuocheng Yu +1Multimodal QAAI Agent Benchmarks

  21. RideGym: A Standardized Interface for Real-World Large-Scale Ride-Sharing System

    Jul 11, 2026Zijian Zhao, Yulong Hu, Sen LiVehicle RoutingAI Agent Benchmarks

  22. Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

    Jul 11, 2026Dongping Liu, Aoyu Zhang, Luyao ZhangAI Agent EvaluationAI Agent Benchmarks

  23. Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

    Jul 10, 2026Jiale Liu, Huajun Xi, Shaokun Zhang +6Agent Failure AnalysisLLM Agent Evaluation

  24. DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

    Jul 10, 2026Jerzy Kamiński, Ilya Galyukshev, Artem Kuznetsov +4Computer-Use Agent BenchmarksLLM Agent Evaluation

  25. Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

    Jul 9, 2026Zongxia Li, Zhongzhi Li, Yucheng Shi +10Long-Horizon Agent EvaluationAI Agent Benchmarks