AI Agent Benchmarks

Momentum

209 papers in the last four weeks, up 61% on the four weeks before. 1.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,159

All topics
CardsList
  1. From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

    Apr 21, 2026Md Nayem Uddin, Kumar Shubham, Eduardo Blanco +2Long-Horizon Agent EvaluationAI Agent Benchmarks

  2. SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

    Apr 21, 2026Josue Torres-Fonseca, Naihao Deng, Yinpei Dai +5LLM Safety BenchmarksLarge Language Model-Based Robot Planning

  3. Time Series Augmented Generation for Financial Applications

    Apr 21, 2026Anton Kolonin, Alexey Glushchenko, Evgeny Bochkov +1LLM Agent EvaluationAI Agent Benchmarks

  4. Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

    Apr 21, 2026Bobo Li, Rui Wu, Zibo Ji +5AI Agent BenchmarksMulti-Agent Reasoning

  5. AutomationBench

    Apr 21, 2026Daniel Shepard, Robin SalimansLLM Agent EvaluationBusiness Process Automation

  6. ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

    Apr 20, 2026Xirui Li, Ming Li, Ion Stoica +2Agent EvaluationAI Agent Benchmarks

  7. MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline

    Apr 20, 2026Jiyao Liu, Jianghan Shen, Sida Song +19HealthcareAI Agent Benchmarks

  8. ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship

    Apr 20, 2026Zhaopei Huang, Yanfeng Jia, Jiayi Zhao +3AI CompanionsAI Agent Benchmarks

  9. Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

    Apr 20, 2026Guanting Dong, Junting Lu, Junjie Huang +17Lifelong Learning AgentsAI Agent Benchmarks

  10. AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

    Apr 20, 2026Wentao Shi, Yu Wang, Yuyang Zhao +8LLM-as-a-JudgeLLM Agent Evaluation

  11. ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

    Apr 20, 2026Yindong Zhang, Wenmian Yang, Yiquan Zhang +1LLM Agent OrchestrationAI Agent Benchmarks

  12. Latent Preference Modeling for Multi-Session Personalized Tool Calling

    Apr 20, 2026Yejin Yoon, Minseo Kim, Taeuk KimLLM PersonalizationLLM Agent Memory

  13. Scaling Human-AI Coding Collaboration Requires a Governable Consensus Layer

    Apr 20, 2026Tianfu Wang, Zhezheng Hao, Yin Wu +5AI Agent BenchmarksAI Agent Governance

  14. Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots

    Apr 20, 2026Shiquan Zhang, Tianyi Zhang, Le Fang +3Mobile GUI AutomationLLM Agent Evaluation

  15. SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

    Apr 19, 2026Ziao Zhang, Kou Shi, Shiting Huang +13Lifelong Learning AgentsLong-Horizon Agent Evaluation

  16. HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents

    Apr 19, 2026Chao Jin, Wenkui Yang, Hao Sun +6GUI AgentsHallucination in Language Models

  17. SeekerGym: A Benchmark for Reliable Information Seeking

    Apr 18, 2026Remy Kim, Minseung Lee, Shuo Li +1AI Agent ReliabilityAgentic Search

  18. Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents

    Apr 18, 2026Sukai Huang, Chenyuan Zhang, Fucai Ke +4AI Agent EvaluationAI Agent Benchmarks

  19. Agentic Frameworks for Reasoning Tasks: An Empirical Study

    Apr 17, 2026Zeeshan Rasheed, Abdul Malik Sami, Muhammad Waseem +3LLM Agent OrchestrationAI Agent Evaluation

  20. Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent Workflows

    Apr 17, 2026Luay Gharzeddine, Samer SaabMulti-Agent CoordinationMulti-Agent Orchestration

  21. SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems

    Apr 17, 2026Hikaru Shindo, Hanzhao Lin, Lukas Helff +2LLM Agent EvaluationAI Agent Benchmarks

  22. GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

    Apr 17, 2026Jize Wang, Xuanxuan Liu, Yining Li +7Tool-Use EvaluationLong-Horizon Agent Evaluation