AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 341

All topics
CardsList
  1. DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

    Aug 11, 2026Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub +3Computer-Use Agent BenchmarksAI Agent Evaluation

  2. The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI

    Aug 10, 2026Srinivas Telukunta, Georgios Nektarios Lilis, Lucio BaronAI Agent EvaluationAI Agent Governance

  3. Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

    Aug 10, 2026Shulin Tian, Ziqi Huang, Fan Zhang +3Tool-Augmented Language Model AgentsAI Agent Evaluation

  4. Software Engineering for and with GUI Agent

    Aug 10, 2026Shengcheng Yu, Yuchen Ling, Junyang Xing +3AI Agent ReliabilityGUI Agents

  5. Evo-Bench: Can Language Models Improve Agent Harness?

    Aug 10, 2026Lisheng Huang, Chen Yang, Hao Zhou +6AI Agent EvaluationLong-Horizon Agent Evaluation

  6. Causal Behavioral Evaluation of AI Agents at Scale via Automated Behavioral Science

    Aug 10, 2026Soo Yong Lee, Jongha Lee, Jaewan Chun +8AI-Assisted Scientific ResearchAI Agent Evaluation

  7. Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

    Aug 9, 2026Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen +15AI in EducationAI Agent Evaluation

  8. OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

    Aug 9, 2026Andrea Caciolai, Pere-Lluís Huguet Cabot, Chierh Cheng +11AI Agent EvaluationAI Agent Benchmarks

  9. ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

    Aug 9, 2026Guo Chen, Ziwen Li, Reed Li +4Multi-Agent LLM SystemsAI Agent Evaluation

  10. Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

    Aug 9, 2026Harshitha Kolukuluru, Reshma Ashok, Kirat Arora +7Deep Research AgentsAI Agent Evaluation

  11. AndroidReality: How Far Are Mobile Agents from the Real World?

    Aug 7, 2026Xiaoou Liu, Longchao Da, Hanyang Chen +2Agent ReliabilityGUI Agents

  12. Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability

    Aug 6, 2026Guanghui Min, Liang Wu, Mayank Darbari +2AI Agent EvaluationLong-Horizon LLM Agents

  13. When Agentic AI Meets Integrated Sensing and Communication

    Aug 6, 2026Kai Li, Conggai Li, Sarah Ali Siddiqui +4Integrated Sensing and CommunicationAI Agent Evaluation

  14. ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study

    Aug 5, 2026Siyuan Li, Peng Shu, Churan Yu +19AI Agent EvaluationAgentic AI

  15. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

    Aug 4, 2026Maksymilian Wolski, Nicholas Hoernle, Johannes Forkel +1Multi-Agent CoordinationAI Agent Evaluation

  16. When Truth Is Distributed: Misinformation Derails Collective Fact Recovery in LLM-Based Multi-Agent Systems

    Aug 4, 2026Chenfei Yan, Zeyang Yue, Feifei Zhao +6AI Agent EvaluationMulti-Agent Systems

  17. Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

    Aug 4, 2026Saqib Shouqi, Abdullah Nazly, Januki Wanniarachchi +1Adversarial Attacks on LLMsLLM Agent Evaluation

  18. Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce

    Aug 3, 2026Shicheng Fan, Mingdai Yang, Duohao Wang +9AI Agent EvaluationAgent Evaluation

  19. Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable

    Aug 2, 2026Wenhui Chen, Jianlin Chen, Ziyao Lin +1Reward HackingAI for Science

  20. FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction

    Aug 1, 2026Chaoqun Yang, Fengbin Zhu, Xinyu Lin +5Quantitative FinanceDeep Research Agents

  21. Beyond Component Testing: Validating Agentic AI Systems

    Jul 31, 2026Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi +4AI Agent EvaluationAI Agent Safety

  22. Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

    Jul 31, 2026Yuxuan Liu, Zhaochen Su, Yuhao Zhang +9AI Agent EvaluationAgent Skill Learning

  23. InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

    Jul 31, 2026Yuan Gao, Zeren Yang, Junnan Li +6AI Agent ReliabilityAI Agent Evaluation