AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

    Aug 7, 2026Joshua Zuniga, Srinivasan Subramanian, Ramya Madhuri Narapureddy +1Agent Failure AnalysisAI Agent Benchmarks

  2. FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

    Aug 6, 2026Bo Deng, Kang Zhou, Lifan Guo +6Long-Horizon Agent EvaluationAI Agent Benchmarks

  3. Unified Agent: Managing Interactions across Devices

    Aug 6, 2026Xinshuang Liu, Runfa Blark Li, Shaoxiu Wei +2AI Agent BenchmarksState Tracking

  4. StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

    Aug 6, 2026Xichen Zhang, Guankai Li, Yinghao Zhu +6Memory-Augmented VLMsStreaming Video Understanding

  5. SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

    Aug 6, 2026Zhi Han, Chenxi Zeng, Liuhaichen Yang +3LLM-as-a-JudgeLong-Horizon Agent Evaluation

  6. RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough

    Aug 5, 2026Anchen Sun, Kaiqi YangMulti-Agent LLM SystemsAI Agent Benchmarks

  7. Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems

    Aug 4, 2026J. de Curtò, I. de ZarzàCyber-Physical SystemsLLM Planning

  8. ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

    Aug 4, 2026Tianyi Guan, Yiding Wang, Haotong Yang +5Continual Learning for LLM AgentsLLM Agent Evaluation

  9. GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

    Aug 4, 2026Leijun Zhou, Zhihao Liu, Xiang Qu +9LLM Agent EvaluationAI Agent Benchmarks

  10. Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

    Aug 4, 2026William Bolton, Philip TorrAI Agents for Scientific DiscoveryBenchmark Design

  11. Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

    Aug 4, 2026Zejun Liu, Jian Wu, Ru Peng +4AI for ScienceBenchmark Design

  12. MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

    Aug 4, 2026Qiming Li, Shujie Hu, Haohan Liu +3Code GenerationAI Agent Benchmarks

  13. EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

    Aug 4, 2026Unggi Lee, Sookbun Lee, Yeil Jeong +3Intelligent Tutoring SystemsAI in Education

  14. ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision

    Aug 3, 2026Wei-Jung Huang, Bonan ShenPairwise ComparisonLLM Agent Evaluation

  15. Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce

    Aug 3, 2026Shicheng Fan, Mingdai Yang, Duohao Wang +9AI Agent EvaluationAgent Evaluation

  16. PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

    Aug 3, 2026Abdulrahman AlRabah, Xiaocheng Yang, Dilek Hakkani-Tür +1AI Agent ReliabilityAI-Assisted Decision Making

  17. ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

    Aug 3, 2026Vernon Toh, Navonil Majumder, Zhengyuan Liu +2LLM Agent EvaluationAI Agent Benchmarks

  18. From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents

    Aug 3, 2026Jiajia Song, Bobo Li, Haiwen Yi +6LLM PersonalizationLLM Agent Evaluation