AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 344

All topics
CardsList
  1. Distilling Agentic Systems: A Roadmap across Models, Artifacts, and Harnesses

    Sep 29, 2026Ziluowen Luo, Senzhang Wang, Chaozhuo Li +9AI Agent EvaluationKnowledge Distillation

  2. From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

    Sep 28, 2026Yaxiao Liu, Pengbo Liu, Yiwen Liu +2Domain AdaptationAI Agent Evaluation

  3. BIABench: Evaluating AI agents on real-world bioimage analysis tasks

    Sep 28, 2026Zixuan Pan, Davide Panzeri, Lukas Johanns +5AI Agent ReliabilityAI Agent Evaluation

  4. GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

    Sep 28, 2026Shaoqing Zhang, Kehai Chen, Xuefeng Bai +4Agent Failure AnalysisGUI Agents

  5. SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities

    Sep 27, 2026Xiaonan Luo, Yue Huang, Kehan Guo +9Software Vulnerability DetectionAI Agent Evaluation

  6. What Happens During Autonomous Deep Research After the User Steps Away?

    Sep 27, 2026Yimin Liu, Yijia Zhang, Yanmin Li +4Counterfactual EvaluationDeep Research Agents

  7. ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

    Sep 24, 2026Ming Zhang, Zhenghao Xiang, Peizhong Gao +17AI Agent EvaluationAI Agent Benchmarks

  8. LabFactory: Building and Evaluating Executable AI Labs

    Sep 23, 2026Jinge Wu, Hongjian Zhou, Mingde Zeng +6AI-Assisted Scientific ResearchAI Agent Evaluation

  9. WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

    Sep 23, 2026Jingjie Ning, Xueqi Li, Yibo Kong +1AI Agent EvaluationActive Learning

  10. XYEval: Agents say yes to bad advice

    Sep 20, 2026Zhengxuan Wu, Yuxuan Li, Oyvind Tafjord +1LLM SycophancyAI Agent Evaluation

  11. Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

    Sep 17, 2026Sho Kawano, Zehang Richard Li, Paul A. ParkerAI Agent EvaluationPrediction-Powered Inference

  12. A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

    Sep 17, 2026Shaina Raza, Ahmed Y. Radwan, Imran Liaquat +1LLM EvaluationAI Agent Evaluation

  13. Do AI Agents Understand Computer Architecture?

    Sep 16, 2026Ambika Sharan, Grigory Chirkov, Soheil AbbaslooElectronic Design AutomationAI Agent Evaluation

  14. Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

    Sep 16, 2026Xinshuai Guo, Junjie Wu, Dolly Deng +4AI Agent EvaluationAI Agent Benchmarks

  15. WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation

    Sep 14, 2026Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy +1Synthetic Data GenerationAI Agent Evaluation