AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 341

All topics
CardsList
  1. Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

    May 20, 2026Haiyang Shen, Jiuzheng Wang, Taian Guo +9AI AccountabilityBenchmark Construction

  2. Open-World Evaluations for Measuring Frontier AI Capabilities

    May 19, 2026Sayash Kapoor, Peter Kirgis, Andrew Schwartz +15AI Agent EvaluationLong-Horizon Agent Evaluation

  3. What Do Evolutionary Coding Agents Evolve?

    May 19, 2026Nico Pelleriti, Sree Harsha Nelaturu, Zhanke Zhou +4Evolutionary OptimizationAI Coding Agents

  4. Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation

    May 19, 2026Yuxuan Gao, Megan Wang, Yi Ling YuUncertainty QuantificationAI Agent Evaluation

  5. OpenComputer: Verifiable Software Worlds for Computer-Use Agents

    May 19, 2026Jinbiao Wei, Qianran Ma, Yilun Zhao +4Computer-Use Agent BenchmarksComputer-Use Agents

  6. Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks

    May 18, 2026Yubin Qu, Ying Zhang, Yanjun Zhang +4Computer-Use Agent BenchmarksCoding Agents

  7. Interactive Evaluation Requires a Design Science

    May 18, 2026Keyang Xuan, Peiyang Song, Pan Lu +10LLM Agent EvaluationAI Agent Evaluation

  8. FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics

    May 17, 2026Qiran Zou, Hou Hei Lam, Wenhao Zhao +11AI Agents for Scientific DiscoveryInference-Time Search

  9. Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

    May 17, 2026Jinhu Qi, Muzhi Li, Jiahong Liu +9AI Agent ReliabilityAI Agent Evaluation

  10. 1GC-7RC: One Graphic Card -- Seven Research Challenges! How Good Are AI Agents at Doing Your Job?

    May 16, 2026Robin-Nico Kampa, Fabian Deuser, Anna Bößendörfer +2AI Coding AgentsAI Agent Evaluation

  11. MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

    May 14, 2026Minghao Guo, Qingyue Jiao, Zeru Shi +14Multimodal MemoryAI Agent Evaluation

  12. Holistic Evaluation and Failure Diagnosis of AI Agents

    May 14, 2026Netta Madvil, Gilad Dym, Alon Mecilati +12Agent Failure AnalysisAI Agent Evaluation

  13. Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment

    May 14, 2026Yuchen Sun, Pei Fu, Shaojie Zhang +6GUI AgentsAI Agent Evaluation

  14. Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction

    May 13, 2026Darius A. Faroughy, Sofia Palacios Schweitzer, Ian Pang +2AI Agents for Scientific DiscoveryAI Agent Evaluation

  15. Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

    May 12, 2026Miaosen Zhang, Xiaohan Zhao, Zhihong Tan +14Computer-Use Agent BenchmarksComputer-Use Agents

  16. No More, No Less: Task Alignment in Terminal Agents

    May 12, 2026Sina Mavali, David Pape, Jonathan Evertz +5Computer-Use Agent BenchmarksAI Alignment

  17. Rollout Cards: A Reproducibility Standard for Agent Research

    May 12, 2026Charlie Masters, Ziyuan Liu, Stefano V. AlbrechtAI Agent EvaluationAgent Evaluation

  18. Toward Modeling Player-Specific Chess Behaviors

    May 12, 2026Loris Sogliuzzo, Aloïs Rautureau, Eric PietteGame-Playing AgentsAI Agent Evaluation

  19. ABRA: Agent Benchmark for Radiology Applications

    May 11, 2026Bulat Maksudov, Vladislav Kurenkov, Kathleen M. Curran +1Radiology Report GenerationHealthcare

  20. Generalizing the Turing Test to Interactive Agents

    May 11, 2026Daniel Mitropolsky, Riccardo Neumarker, Emanuele Rimoldi +2LLM EvaluationAI Agent Evaluation

  21. From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

    May 11, 2026Pedro Conde, Henrique Branquinho, Valerio Mazzone +3Software Vulnerability DetectionAI Agent Evaluation

  22. Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability

    May 11, 2026Harsh Raj, Niranjan Orkat, Suvrorup Mukherjee +3AI Agent ReliabilityAgent Reliability

  23. Agentic Performance at the Edge: Insights from Benchmarking

    May 11, 2026Shiqiang Wang, Herbert WoisetschlägerLLM Agent EvaluationAI Agent Evaluation

  24. Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

    May 11, 2026Haonan Dong, Qiguan Feng, Kehan Jiang +3AI AlignmentAI Agent Evaluation

  25. SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

    May 11, 2026Zonglin Yang, Xingtong Liu, Xinyan XuAI Agents for Scientific DiscoveryAI Agent Evaluation

  26. Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence

    May 10, 2026Alex Bogdan, Adrian de Valois-FranklinAI Agent EvaluationAI Agent Monitoring

  27. MDGYM: Benchmarking AI Agents on Molecular Simulations

    May 9, 2026Vinay Kumar, Satyendra Rajput, Mausam +1AI Agents for Scientific DiscoveryAI Agent Evaluation