AI Agent Benchmarks

Momentum

96 papers in the last four weeks, up 30% on the four weeks before. 1.2% of all new papers.

Jul 6Week of Sep 21

Latest papers 1,013

All topics
CardsList
  1. GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

    May 28, 2026Xiaoyi Chen, Yifei Gao, Yang Xu +3GUI AgentsAI Agent Benchmarks

  2. WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

    May 28, 2026Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu +14Multimodal MemoryLLM Agent Memory

  3. STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments

    May 28, 2026Junyang Wang, Haiyang Xu, Xi Zhang +4GUI AgentsAI Agent Benchmarks

  4. OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

    May 28, 2026Yibing Liu, Yangze Liu, Xiaolong Yin +4Agent Failure AnalysisAnomaly Localization

  5. BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

    May 28, 2026Jiahao Huang, Fei Cheng, Junfeng Jiang +2LLM Agent EvaluationAI Agent Benchmarks

  6. Personal Visual Memory from Explicit and Implicit Evidence

    May 27, 2026Viet Nguyen, Thao Nguyen, Vishal M. Patel +1Multimodal MemoryAI Agent Benchmarks

  7. LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

    May 27, 2026HuiMing Fan, Xiao Wang, Zheng Chu +5Computer-Use Agent BenchmarksLLM Agent Evaluation

  8. VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

    May 27, 2026Yuting Xu, Jiayi Tian, Jian Liang +4Multimodal ReasoningAI Agent Benchmarks

  9. Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

    May 27, 2026Jihyeong Park, Ingeol Baek, Jeonghyun Park +1LLM Agent EvaluationAI Agent Benchmarks

  10. From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets

    May 27, 2026Taojie Zhu, Wentao Zhao, Rui Sun +7Benchmark ContaminationLLM Agent Evaluation

  11. OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents

    May 27, 2026Chenyu Zhou, Xinyun Lu, Jiangyue Zhao +3LLM Agent EvaluationAI Agent Benchmarks

  12. Verifiable Benchmarking of Long-Horizon Spatial Biology

    May 27, 2026Ian Diks, Harihara Muralidharan, Tim Proctor +1Spatial TranscriptomicsBenchmark Design

  13. MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents

    May 27, 2026Zihan Li, Xingyu Fan, Feifei Li +1Conversational MemoryAI Agent Benchmarks

  14. AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

    May 27, 2026Kou Shi, Ziao Zhang, Shiting Huang +7Tool-Use EvaluationAI Agent Benchmarks

  15. DisasterBench: Benchmarking LLM Planning under Typed Tool Interface Constraints

    May 27, 2026Zhitong Chen, Kai Yin, Weifeng Zhang +7Tool-Use PlanningAI Agent Benchmarks

  16. UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

    May 27, 2026Pengyu Zhu, Lijun Li, Yaxing Lyu +8LLM Agent EvaluationAI Agent Benchmarks

  17. Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

    May 27, 2026Jaechang Kim, Sunung Mun, Seungjoon Lee +2Explainable Artificial IntelligenceAI Agent Benchmarks

  18. AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

    May 26, 2026Yifan Sui, Xin Huang, Hongbing Li +14Mobile GUI AutomationComputer-Use Agent Benchmarks

  19. DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents

    May 26, 2026Shijie Cao, Yuan Yuan, Jing LiuLLM Agent EvaluationCombinatorial Optimization

  20. Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

    May 26, 2026Yipeng Ouyang, Xin Huang, Bingjie Liu +3Software Engineering AgentsLLM Agent Evaluation

  21. ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents

    May 26, 2026Xing Fu, Yulin Hu, Mengtong Ji +5Emotional Support ConversationAI Agent Benchmarks

  22. VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

    May 26, 2026Yuxin Chen, Yi Zhang, Zhengzhou Cai +11LLM Agent MemoryAI Agent Evaluation

  23. IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

    May 26, 2026Jinzhao Li, Yinuo Chen, Wenxuan Song +5Streaming Video UnderstandingAI Agent Benchmarks

  24. ChartAct: A Benchmark for Dynamic Chart Understanding

    May 26, 2026Muye Huang, Lin Wu, Lingling Zhang +5Data VisualizationChart QA