AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors

    Apr 23, 2026Chenghao Yang, Yuning Zhang, Zhoufutu Wen +4Tool-Augmented Language Model AgentsLLM Agent Evaluation

  2. How VLAs (Really) Work In Open-World Environments

    Apr 23, 2026Amir Rasouli, Yangzheng Wu, Zhiyuan Li +4VLM RobustnessLong-Horizon Robotic Manipulation

  3. From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

    Apr 21, 2026Md Nayem Uddin, Kumar Shubham, Eduardo Blanco +2Long-Horizon Agent EvaluationAI Agent Benchmarks

  4. SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

    Apr 21, 2026Josue Torres-Fonseca, Naihao Deng, Yinpei Dai +5LLM Safety BenchmarksLarge Language Model-Based Robot Planning

  5. Time Series Augmented Generation for Financial Applications

    Apr 21, 2026Anton Kolonin, Alexey Glushchenko, Evgeny Bochkov +1LLM Agent EvaluationAI Agent Benchmarks

  6. Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

    Apr 21, 2026Bobo Li, Rui Wu, Zibo Ji +5AI Agent BenchmarksMulti-Agent Reasoning

  7. AutomationBench

    Apr 21, 2026Daniel Shepard, Robin SalimansLLM Agent EvaluationBusiness Process Automation

  8. ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

    Apr 20, 2026Xirui Li, Ming Li, Ion Stoica +2Agent EvaluationAI Agent Benchmarks

  9. MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline

    Apr 20, 2026Jiyao Liu, Jianghan Shen, Sida Song +19HealthcareAI Agent Benchmarks

  10. ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship

    Apr 20, 2026Zhaopei Huang, Yanfeng Jia, Jiayi Zhao +3AI CompanionsAI Agent Benchmarks

  11. Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

    Apr 20, 2026Guanting Dong, Junting Lu, Junjie Huang +17Lifelong Learning AgentsAI Agent Benchmarks

  12. AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

    Apr 20, 2026Wentao Shi, Yu Wang, Yuyang Zhao +8LLM-as-a-JudgeLLM Agent Evaluation

  13. ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

    Apr 20, 2026Yindong Zhang, Wenmian Yang, Yiquan Zhang +1LLM Agent OrchestrationAI Agent Benchmarks

  14. Latent Preference Modeling for Multi-Session Personalized Tool Calling

    Apr 20, 2026Yejin Yoon, Minseo Kim, Taeuk KimLLM PersonalizationLLM Agent Memory

  15. Scaling Human-AI Coding Collaboration Requires a Governable Consensus Layer

    Apr 20, 2026Tianfu Wang, Zhezheng Hao, Yin Wu +5AI Agent BenchmarksAI Agent Governance

  16. Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots

    Apr 20, 2026Shiquan Zhang, Tianyi Zhang, Le Fang +3Mobile GUI AutomationLLM Agent Evaluation

  17. SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

    Apr 19, 2026Ziao Zhang, Kou Shi, Shiting Huang +13Lifelong Learning AgentsLong-Horizon Agent Evaluation

  18. HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents

    Apr 19, 2026Chao Jin, Wenkui Yang, Hao Sun +6GUI AgentsHallucination in Language Models

  19. SeekerGym: A Benchmark for Reliable Information Seeking

    Apr 18, 2026Remy Kim, Minseung Lee, Shuo Li +1AI Agent ReliabilityAgentic Search

  20. Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents

    Apr 18, 2026Sukai Huang, Chenyuan Zhang, Fucai Ke +4AI Agent EvaluationAI Agent Benchmarks

  21. Agentic Frameworks for Reasoning Tasks: An Empirical Study

    Apr 17, 2026Zeeshan Rasheed, Abdul Malik Sami, Muhammad Waseem +3LLM Agent OrchestrationAI Agent Evaluation

  22. Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent Workflows

    Apr 17, 2026Luay Gharzeddine, Samer SaabMulti-Agent CoordinationMulti-Agent Orchestration