AI Agent Benchmarks

Momentum

146 papers in the last four weeks, up 161% on the four weeks before. 1.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio

    Jun 15, 2026Anzhe Xie, Weihang Su, Yujia Zhou +2AI-Assisted Scientific ResearchLLM Agent Evaluation

  2. LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

    Jun 15, 2026Anqi Zou, Han Deng, Chengyu Zhang +9Computer-Use Agent BenchmarksComputer-Use Agents

  3. MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

    Jun 15, 2026Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh +1Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  4. AgentFairBench: Do LLM Agents Discriminate When They Act?

    Jun 15, 2026Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh +1Algorithmic FairnessLLM Agent Evaluation

  5. VisualClaw: A Real-Time, Personalized Agent for the Physical World

    Jun 15, 2026Haoqin Tu, Jianwen Chen, Zijun Wang +14Continual Learning for LLM AgentsAI Agent Benchmarks

  6. UXBench: Measuring the Actionability of LLM-Generated UX Critiques

    Jun 15, 2026Wenjie Wang, Yue Huang, Zipeng Ling +11LLM-as-a-JudgeLLM Agent Evaluation

  7. ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

    Jun 13, 2026Rahul Suresh Babu, Laxmipriya Ganesh IyerAI Agent ReliabilityTool-Use Evaluation

  8. CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?

    Jun 13, 2026Yuxin Zhang, Ju Fan, Meihao Fan +2AI Coding AgentsSoftware Engineering Benchmarks

  9. Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

    Jun 13, 2026Shijun Wan, Xuehai Wu, Jiwen Zhang +2LLM Agent EvaluationAI Agent Benchmarks

  10. Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning

    Jun 13, 2026Sanhorn Chen, Xiaoyang Chen, Boyu Liu +1Irregular Time-Series ModelingLLM Evaluation

  11. Running the Gauntlet: Hard Agentic Tasks

    Jun 12, 2026Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +203D Spatial ReasoningWeb Agent Benchmarks

  12. OdysSim: Building Foundation Models for Human Behavior Simulation

    Jun 12, 2026Xuhui Zhou, Weiwei Sun, Weihua Du +6Human Behavior SimulationAI Agent Benchmarks

  13. LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents

    Jun 12, 2026Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle +2Educational AssessmentGenerative AI in Education

  14. Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents

    Jun 12, 2026Brendan King, Jeffrey FlaniganCoding AgentsLLM Agent Evaluation

  15. SANA: What Matters for QA Agents over Massive Data Lakes?

    Jun 11, 2026Austin Senna Wijaya, Jiaxiang Liu, Haonan Wang +1LLM Agent EvaluationQuestion Answering

  16. Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs

    Jun 11, 2026Pratham Singla, Shivank Garg, Vihan SinghGame-Playing AgentsStrategic Reasoning Risks in Language Models

  17. EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

    Jun 11, 2026Jundong Xu, Qingchuan Li, Jiaying Wu +11Benchmark DesignLong-Horizon Agent Evaluation

  18. AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

    Jun 11, 2026Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26Software Engineering AgentsAI Agent Evaluation

  19. EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

    Jun 11, 2026Harihara Muralidharan, Reema Baskar, Soo Hee Lee +2AI Agents for Scientific DiscoveryAI Agent Evaluation

  20. RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue

    Jun 11, 2026Sara Candussio, Emanuele Ballarin, Lorenzo Bonin +2AI Agent ReliabilityLanguage Model Safety Evaluation

  21. ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm

    Jun 11, 2026Jiaxin Ai, Tao Hu, Xuemeng Yang +11Computer-Use AgentsProgram Synthesis

  22. TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

    Jun 11, 2026Dat Tien Nguyen, Thao Nguyen, Fadillah Adamsyah Maani +5AI Agents for Scientific DiscoveryScientific Reasoning

  23. EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge

    Jun 11, 2026Yunhan Wang, Jiaan Wang, Lianzhe Huang +2Benchmark ContaminationWeb Search Agents

  24. The Illusion of Multi-Agent Advantage

    Jun 11, 2026Prathyusha Jwalapuram, Hehai Lin, Chuyuan Li +7AI Agent EvaluationAgent Evaluation