AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 341

All topics
CardsList
  1. Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

    May 9, 2026Filippo Ziliotto, Ciro Beneduce, Bruno Lepri +3AI Agent EvaluationVision-Language Grounding

  2. Log analysis is necessary for credible evaluation of AI agents

    May 8, 2026Peter Kirgis, Sayash Kapoor, Stephan Rabanser +8Agent Failure AnalysisAI Agent Evaluation

  3. Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge

    May 8, 2026Dhaval Patel, Chathurangi Shyalika, Suryanarayana Reddy Yarrabothula +4Competitive AnalysisMulti-Agent Orchestration

  4. Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

    May 8, 2026Prasanna Desikan, Harshit Rajgarhia, Shivali Dalmia +1HealthcareBenchmark Design

  5. When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory

    May 8, 2026Jiaqi Shao, Yiyi Lu, Yunzhen Zhang +1Agent MemoryAI Agent Evaluation

  6. EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

    May 8, 2026Yi Liu, TingFeng Hui, Wei Zhang +4AI Agent EvaluationAI Agent Benchmarks

  7. Computer Use at the Edge of the Statistical Precipice

    May 7, 2026Pierluca D'Oro, Sneha Silwal, William Wong +6Computer-Use Agent BenchmarksComputer-Use Agents

  8. Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

    May 7, 2026Roger Creus Castanyer, Pablo Samuel Castro, Glen BersethAI Agent EvaluationAI Agent Benchmarks

  9. When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

    May 7, 2026Vasilis Niarchos, Constantinos Papageorgakis, Alexander G. Stapleton +1AI for ScienceAI Agent Evaluation

  10. Process Matters more than Output for Distinguishing Humans from Machines

    May 7, 2026Milena Rmus, Mathew D. Hardy, Thomas L. Griffiths +1AI Agent EvaluationImitation Learning

  11. Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors

    May 7, 2026Jonas Wiedermann-Möller, Leonard Dung, Maksym AndriushchenkoAI Agent EvaluationAI Agent Safety

  12. BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

    May 7, 2026Jinge Wu, Hongjian Zhou, Mingde Zeng +8AI Agents for Scientific DiscoveryDeep Research Agents

  13. Multi-Dimensional Behavioral Evaluation of Agentic Stock Prediction Systems Using Large Language Model Judges with Closed-Loop Reinforcement Learning Feedback

    May 7, 2026Mohammad Al Ridhawi, Mahtab Haj Ali, Hussein Al OsmanTime Series ForecastingAI Agent Evaluation

  14. Intentionality is a Design Decision: Measuring Functional Intentionality for Accountable AI Systems

    May 6, 2026Allessia Chiappetta, Robert MahariAI AccountabilityAI Agent Evaluation

  15. DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents

    May 6, 2026Zhaorun Chen, Xun Liu, Haibo Tong +14AI Agent EvaluationAI Agent Security

  16. Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

    May 5, 2026Zirui Tang, Xuanhe Zhou, Yumou Liu +19Computer-Use Agent BenchmarksAI Agent Evaluation

  17. Foundation-Model-Based Agents in Industrial Automation: Purposes, Capabilities, and Open Challenges

    May 4, 2026Vincent Henkel, Felix Gehlhoff, David Kube +13AI Agent EvaluationLLM Agent Reliability

  18. STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

    May 4, 2026Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6Annotator DisagreementAI Agent Evaluation

  19. Preregistration for Experiments with AI Agents

    May 3, 2026Michelle VaccaroAI Agent Evaluation

  20. Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents

    May 1, 2026Maksym Nechepurenko, Pavel ShuvalovAI Agent EvaluationForecasting Benchmarks

  21. Agentic AI for Substance Use Education: Integrating Regulatory and Scientific Knowledge Sources

    May 1, 2026Kosar Haghani, Zahra Kolagar, Mohammed AtiquzzamanAI Agent EvaluationGenerative AI in Education

  22. WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments

    Apr 30, 2026Jinchao Li, Yunxin Li, Chenrui Zhao +3Computer-Use Agent BenchmarksComputer-Use Agents

  23. End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians

    Apr 30, 2026Aaryan Shah, Andrew Hines, Alexia Downs +6HealthcareAI Agent Evaluation