AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 341

All topics
CardsList
  1. BIABench: Evaluating AI agents on real-world bioimage analysis tasks

    Sep 28, 2026Zixuan Pan, Davide Panzeri, Lukas Johanns +5AI Agent ReliabilityAI Agent Evaluation

  2. GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

    Sep 28, 2026Shaoqing Zhang, Kehai Chen, Xuefeng Bai +4Agent Failure AnalysisGUI Agents

  3. SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities

    Sep 27, 2026Xiaonan Luo, Yue Huang, Kehan Guo +9Software Vulnerability DetectionAI Agent Evaluation

  4. What Happens During Autonomous Deep Research After the User Steps Away?

    Sep 27, 2026Yimin Liu, Yijia Zhang, Yanmin Li +4Counterfactual EvaluationDeep Research Agents

  5. LabFactory: Building and Evaluating Executable AI Labs

    Sep 23, 2026Jinge Wu, Hongjian Zhou, Mingde Zeng +6AI-Assisted Scientific ResearchAI Agent Evaluation

  6. WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

    Sep 23, 2026Jingjie Ning, Xueqi Li, Yibo Kong +1AI Agent EvaluationActive Learning

  7. XYEval: Agents say yes to bad advice

    Sep 20, 2026Zhengxuan Wu, Yuxuan Li, Oyvind Tafjord +1LLM SycophancyAI Agent Evaluation

  8. Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

    Sep 17, 2026Sho Kawano, Zehang Richard Li, Paul A. ParkerAI Agent EvaluationPrediction-Powered Inference

  9. A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

    Sep 17, 2026Shaina Raza, Ahmed Y. Radwan, Imran Liaquat +1LLM EvaluationAI Agent Evaluation

  10. Do AI Agents Understand Computer Architecture?

    Sep 16, 2026Ambika Sharan, Grigory Chirkov, Soheil AbbaslooElectronic Design AutomationAI Agent Evaluation

  11. Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

    Sep 16, 2026Xinshuai Guo, Junjie Wu, Dolly Deng +4AI Agent EvaluationAI Agent Benchmarks

  12. WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation

    Sep 14, 2026Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy +1Synthetic Data GenerationAI Agent Evaluation

  13. DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents

    Sep 13, 2026Zhichao Shi, Xuhui Jiang, Wenjie Zhang +5LLM Agent EvaluationAI Agent Evaluation

  14. Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

    Sep 12, 2026Yuanchen Bai, Zijian Ding, Angelique TaylorAI Agent ReliabilityAgent Reliability