AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles

    Sep 28, 2026Bingo Zhang, Haochuan Lu, Zongjie Li +3GUI AgentsLong-Horizon Agent Evaluation

  2. VehicleArena: A Realistic Urban Environment for Multi-Agent Driving

    Sep 28, 2026Jie Yang, Jiajun Chen, Jiazheng Zhou +4Autonomous Driving BenchmarksAutonomous Driving

  3. FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management

    Sep 28, 2026Peiyu ZangLong-Horizon Agent EvaluationAI Agent Benchmarks

  4. AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs

    Sep 28, 2026Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin +6LLM Inference EfficiencyLLM Inference

  5. The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions

    Sep 28, 2026Xiaoting Lyu, Xinbo Ma, Yufei Han +5AI Agent BenchmarksScientific Reasoning in Language Models

  6. MASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent Systems

    Sep 28, 2026Yapeng Li, Songze Li, Shuang Yu +4Multi-Agent LLM SystemsMulti-Agent Collaboration

  7. PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems

    Sep 28, 2026Xijing Wang, Yinsheng Yao, Jinru Ding +5LLM Agent EvaluationAI Agent Benchmarks

  8. AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

    Sep 28, 2026Chanhee Park, Jeongho Yoon, Sungbin Han +2Multi-Hop QATool-Augmented Language Model Agents

  9. BIABench: Evaluating AI agents on real-world bioimage analysis tasks

    Sep 28, 2026Zixuan Pan, Davide Panzeri, Lukas Johanns +5AI Agent ReliabilityAI Agent Evaluation

  10. SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals

    Sep 28, 2026Jingyun Jia, Antoine Remond-Tiedrez, Aaron Alvarez +3LLM EvaluationAI Agent Benchmarks

  11. Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures

    Sep 28, 2026Dong Xu, Zhangfan Yang, Jiantao Wu +5LLM Agent EvaluationAI Agent Benchmarks

  12. RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

    Sep 28, 2026Haitong Ma, Chenxiao Gao, Rushi Qiang +2AI Agent BenchmarksRobot Learning

  13. PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models

    Sep 28, 2026Shane K. A. Dalumura Hettige, Jonas OppenlaenderCreativity AssessmentAgentic Image Generation

  14. When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

    Sep 27, 2026Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur +2Long-Horizon Agent EvaluationAI Agent Benchmarks

  15. Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models

    Sep 27, 2026Eduardo Ariño de la Rubia, Szilard PafkaAI Coding AgentsLLM Agent Evaluation

  16. Auditing Agent Actions through Query-Conditioned Attribution

    Sep 27, 2026Yifan Liu, Praveen Venkateswaran, Abdulhamid Adebayo +1Gradient-Based AttributionAI Agent Benchmarks

  17. LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs

    Sep 27, 2026Haochen Luo, Yifan Li, Binh Minh An +4Quantitative FinanceLLM Agent Evaluation

  18. DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

    Sep 27, 2026Nan Huang, Mario Tapia-Pacheco, Kun Zhou +4AI Agent ReliabilityScientific Hypothesis Generation

  19. TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

    Sep 27, 2026Dehai Min, Daoan Zhang, Yiming Zeng +13Benchmark ConstructionLLM Agent Evaluation

  20. ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

    Sep 24, 2026Ming Zhang, Zhenghao Xiang, Peizhong Gao +17AI Agent EvaluationAI Agent Benchmarks

  21. RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

    Sep 23, 2026Mithil Salunkhe, Haochen Ding, Samridhi Verma +1LLM Agent EvaluationAI Agent Benchmarks

  22. TRACER: Trajectory-Aligned Learning for Multi-Turn User Simulation

    Sep 23, 2026Geng Chen, Ruotong Pan, Zhirui Yang +9AI Agent BenchmarksUser Simulation

  23. SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

    Sep 23, 2026Zhilong Ge, Yuting Shao, Yutao Yang +6Tool-Augmented Language Model AgentsLLM Agent Skill Learning