AI Agent Benchmarks

Momentum

96 papers in the last four weeks, up 30% on the four weeks before. 1.2% of all new papers.

Jul 6Week of Sep 21

Latest papers 1,013

All topics
CardsList
  1. MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

    Jul 1, 2026Zhishang Xiang, Zerui Chen, Yunbo Tang +5Agent MemoryLLM Sycophancy

  2. LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

    Jul 1, 2026Ruotong Zhao, Zhiyu Chen, Xurui Liu +7Human Preference EvaluationLLM-as-a-Judge

  3. Evaluating Agentic Harness Systems for Autonomous Computational Pathology

    Jul 1, 2026Jie Lin, Zongyi Chen, Qiaoling Zheng +6Computational PathologyLLM Agent Evaluation

  4. EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

    Jul 1, 2026Jihyeok Jung, Jeewu Lee, Sanghyeop Kim +2Egocentric Video UnderstandingAI Agent Benchmarks

  5. PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

    Jul 1, 2026Ke Zhang, Sahchit Chundur, Mohammad Javad Qomi +1Computer-Use Agent BenchmarksAI for Science

  6. QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

    Jun 30, 2026Sergio Hernández-Gutiérrez, Matteo Merler, Ilze Amanda Auzina +3Long-Horizon Agent EvaluationAI Agent Benchmarks

  7. DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

    Jun 30, 2026Meng Chen, Anya Ji, Tsung-Han Wu +4Computer-Use AgentsAI in Education

  8. MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

    Jun 30, 2026Qingyun Liu, Jiwen Zhang, Jingyi Hu +2Multi-Agent CollaborationAI Agent Benchmarks

  9. Xiaomi-GUI-0 Technical Report

    Jun 30, 2026Wanxia Cao, Chengzhen Duan, Pei Fu +29Mobile GUI AutomationRL for GUI Agents

  10. ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

    Jun 30, 2026Kaiwen Xiong, Haonian Ji, Shi Qiu +4LLM Agent OrchestrationLLM Agent Evaluation

  11. AhaBench: Do Agents Turn Experience into Reusable Insights? A Long-Horizon Benchmark for Continual Learning

    Jun 30, 2026Zerui Cheng, Jiawei Xu, Huacan Chai +3Continual Learning for LLM AgentsLifelong Learning Agents

  12. MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning

    Jun 30, 2026Sheng Zhang, Qinglin Li, Yuechao Zang +3Aerial RoboticsLarge Language Model-Based Robot Planning

  13. What Drives Interactive Improvement from Feedback?

    Jun 29, 2026Bartłomiej Cupiał, Jan Łojek, Mikołaj Garstecki +3LLM Self-RefinementLLM Agent Evaluation

  14. MirrorCode: AI can rebuild entire programs from behavior alone

    Jun 29, 2026Tom Adamczewski, David Owen, David Rein +4Software Engineering AgentsAI Agent Benchmarks

  15. Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

    Jun 29, 2026Yangda Peng, Yunjia Qi, Haotian Xia +8LLM-as-a-JudgeAI Agent Benchmarks

  16. CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

    Jun 29, 2026Bo Qu, Mingguang ChenQuantitative FinanceLLM Agent Evaluation

  17. ECHO: Learning Epistemically Adaptive Language Agents with Turn-Level Credit

    Jun 29, 2026Abhijnan Nath, Nikhil KrishnaswamyBelief-Space PlanningCredit Assignment in RL

  18. Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

    Jun 28, 2026Sunqi Fan, Qingle Liu, Runqi Yin +2GUI AgentsVideo QA

  19. Hierarchical Experimentalist Agents

    Jun 28, 2026Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka +2LLM Agent Skill LearningAI Agent Benchmarks

  20. A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

    Jun 28, 2026Yuanhong Cai, Xiaohui Nie, Kanglin Yin +8Software EngineeringLLM Agent Evaluation

  21. Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

    Jun 27, 2026Ananto Nayan Bala, Faisal Muhammad ShahMulti-Label ClassificationAI Agent Benchmarks

  22. SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

    Jun 27, 2026My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak +4LLM Agent EvaluationAI Agent Evaluation

  23. GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

    Jun 26, 2026Amit Parekh, Sabrina McCallum, Kareem Al-Hasan +3Game-Playing AgentsMulti-Agent Collaboration