AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 341

All topics
CardsList
  1. An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding

    Jul 30, 2026Yanyu Ren, Yunfeng Bai, Xizheng Wang +2Multi-Agent CoordinationMulti-Agent Orchestration

  2. Can AI agents conduct open-ended AI research? Early evidence from two case studies

    Jul 29, 2026Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21AI-Assisted Scientific ResearchAI Agent Evaluation

  3. Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

    Jul 28, 2026Stefan Krsteski, Charlotte Meyer, Guillaume Allegre +2AI Agent EvaluationAgent Evaluation

  4. Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

    Jul 28, 2026Zheng Tong, Yang Liu, Wanshu Fan +6AI Agent EvaluationAgentic Workflows

  5. PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

    Jul 28, 2026Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou +11AI Agent EvaluationAI Agent Safety

  6. Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

    Jul 28, 2026Abu Bakar SiddikAI Agent EvaluationCybersecurity

  7. GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

    Jul 27, 2026Jiacheng Lu, Sinuo Wang, Wentao Zhao +12Quantitative FinanceFinancial Forecasting

  8. Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

    Jul 27, 2026Diandian Guo, Cong Cao, Fangfang Yuan +3AI Agent EvaluationLLM Agent Self-Improvement

  9. Stress-testing large language model agents in a robotic chemistry laboratory

    Jul 25, 2026Lulu Guo, Yingkai Sun, Xiaobo Li +13AI Agent ReliabilityAI Agents for Scientific Discovery

  10. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

    Jul 24, 2026Jiaqi Shao, Hanck Chen, Wei Zhang +2Computer-Use Agent BenchmarksReward Hacking

  11. Agentic Evaluation of Copyright Law Compliance

    Jul 23, 2026Zheng Hui, Doni Bloomfield, Noam KoltAI Agent EvaluationAI Agent Benchmarks

  12. BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

    Jul 21, 2026Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang +9AI Agent EvaluationAI Agent Benchmarks

  13. A Diagnostic Framework for AI Agent Behavior

    Jul 19, 2026Xichen Zhang, Yingjie Zhang, Tianshu SunAI Agent EvaluationAI Agent Governance

  14. Teach it to stop, not just to click

    Jul 19, 2026Barada Sahu, Shivesh PandeyAI Agent ReliabilityAgent Reliability

  15. OTAP: Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories

    Jul 19, 2026Babak Barazandeh, Subhabrata Majumdar, George MichailidisLLM Agent EvaluationAI Agent Evaluation

  16. EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding

    Jul 19, 2026Yaohan Yang, Minglei Shi, Borui Zhang +2Computer-Use Agent BenchmarksGUI Agents

  17. Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries

    Jul 18, 2026Mohammad Arvan, Amber E. Osterholt, Bailee Rue +5AI Agent EvaluationHuman-in-the-Loop AI

  18. OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

    Jul 16, 2026Chengyu Shen, Yujie Fu, Gangtao Xin +13Computer-Use Agent BenchmarksAI Agent Evaluation

  19. AI Agents Do Not Fail Alone:The Context Fails First

    Jul 15, 2026Fouad BousetouaneAI Agent ReliabilityAgent Reliability

  20. DeepStress: Stress-Testing Deep Search Agents

    Jul 15, 2026Ismael Rousseau, Geraldine Damnati, Frederic BechetAI Agent ReliabilityAI Agent Evaluation

  21. Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

    Jul 11, 2026Dongping Liu, Aoyu Zhang, Luyao ZhangAI Agent EvaluationAI Agent Benchmarks

  22. SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

    Jul 9, 2026Shilin Ou, Yifan Xu, Luyao ZhangAlgorithmic AuditingAI Agent Evaluation

  23. Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction

    Jul 9, 2026Sophia Koehler, Antonia Wüst, Inga Ibs +5AI Agent EvaluationActive Learning