AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 331

All topics
CardsList
  1. SciExam for ENSO: Can AI Agents Build Climate Models?

    Oct 7, 2026Yinling Zhang, Langchen Liu, Dongbin Xiu +4AI for ScienceAI Agent Benchmarks

  2. Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

    Oct 7, 2026Tan Yu, Alexander Bukharin, Khushi Bhardwaj +19AI Coding AgentsAI Agent Evaluation

  3. Stale, Misattributed, or Late: Where Personal Memory Fails Before Generation

    Oct 7, 2026Haonan Deng, Park SinchaisriAgent MemoryAgent Memory Management

  4. Shared and structured inputs undermine collective random choice by reasoning AI agents

    Oct 7, 2026Takahiro Ezaki, Naoto Imura, Katsuhiro NishinariAI Agent EvaluationSelection Bias

  5. The AI Evaluation Ecosystem

    Oct 7, 2026Yash Dave, Sang T. Truong, Serena Wang +1AI Agent EvaluationBenchmark Design

  6. Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

    Oct 6, 2026Wenyu Du, Stephen ChungAutonomous Scientific DiscoveryAI Agent Evaluation

  7. Agent Plasticity: Measuring Self-Improvement Through Experience

    Oct 6, 2026Harman Singh, Anton Bakhtin, Rulin Shao +8Self-Improving AgentsLLM Agent Self-Improvement

  8. Inspect Robots: Evaluating the Capabilities and Safety of Embodied AI

    Oct 5, 2026Christopher Leet, Achu Menon, Sravanthi Machcha +12AI Agent EvaluationRobotic Policy Evaluation

  9. MiniCorp: The Last Mile of the AI Agent Firm

    Oct 5, 2026Jingying Zeng, Zhenwei Dai, Jinning Li +6Synthetic Data GenerationAI Agent Evaluation

  10. Code Owns the Simulation, Jev Owns the Evaluation

    Oct 1, 2026Yaodong Yang, Hongyao Tang, Yi Ma +4AI Agent EvaluationRobotic Policy Evaluation

  11. Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

    Oct 1, 2026Ngoc Phuoc An Vo, Aarya Doshi, Vadim SheininTool-Use EvaluationAI Agent Evaluation

  12. Agents Are Systems, Not Models: Rethinking Agentic Evaluation

    Oct 1, 2026Luis Wiedmann, Leander Girrbach, Cordelia Schmid +1AI Agent ReliabilityAI Agent Evaluation

  13. Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

    Oct 1, 2026Haotian Chen, Bowen Ye, Yuning Zhang +1Multi-Agent LLM SystemsAI Agent Auditing

  14. Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles

    Oct 1, 2026Alexei N. Skurikhin, Emily M. Taylor, Nathan A. DeBardelebenLLM Inference EfficiencyLanguage Model Generation Evaluation

  15. Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

    Sep 30, 2026Michael Hardy, Ruhana Azam, Anka Reuel +2AI Agent EvaluationAgent Evaluation

  16. CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

    Sep 30, 2026Chenmu Zhang, Levi Felix, Jun-Jie Zhang +6AI Agents for Scientific DiscoveryAI Agent Evaluation

  17. No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

    Sep 30, 2026Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian +1LLM Red TeamingAI Agent Evaluation

  18. Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

    Sep 30, 2026Priyanath Maji, Spandan Ghose ChowdhuryThompson SamplingAI Agent Evaluation

  19. AgBench: Agentic AI Benchmarks for Personal AI Devices

    Sep 29, 2026Yizhou Han, Di Wu, Dhananjay Saikumar +1AI Agent ReliabilityAI Agent Evaluation

  20. CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?

    Sep 29, 2026Yue Pan, Jiawei Li, Ziyuan Zhang +2LLM EvaluationLLM-as-a-Judge

  21. Distilling Agentic Systems: A Roadmap across Models, Artifacts, and Harnesses

    Sep 29, 2026Ziluowen Luo, Senzhang Wang, Chaozhuo Li +9AI Agent EvaluationKnowledge Distillation

  22. From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

    Sep 28, 2026Yaxiao Liu, Pengbo Liu, Yiwen Liu +2Domain AdaptationAI Agent Evaluation

  23. BIABench: Evaluating AI agents on real-world bioimage analysis tasks

    Sep 28, 2026Zixuan Pan, Davide Panzeri, Lukas Johanns +5AI Agent ReliabilityAI Agent Evaluation

  24. GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

    Sep 28, 2026Shaoqing Zhang, Kehai Chen, Xuefeng Bai +4Agent Failure AnalysisGUI Agents

  25. SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities

    Sep 27, 2026Xiaonan Luo, Yue Huang, Kehan Guo +9Software Vulnerability DetectionAI Agent Evaluation

  26. What Happens During Autonomous Deep Research After the User Steps Away?

    Sep 27, 2026Yimin Liu, Yijia Zhang, Yanmin Li +4Counterfactual EvaluationDeep Research Agents

  27. LabFactory: Building and Evaluating Executable AI Labs

    Sep 23, 2026Jinge Wu, Hongjian Zhou, Mingde Zeng +6AI-Assisted Scientific ResearchAI Agent Evaluation

  28. WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

    Sep 23, 2026Jingjie Ning, Xueqi Li, Yibo Kong +1AI Agent EvaluationActive Learning

  29. XYEval: Agents say yes to bad advice

    Sep 20, 2026Zhengxuan Wu, Yuxuan Li, Oyvind Tafjord +1LLM SycophancyAI Agent Evaluation

  30. Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

    Sep 17, 2026Sho Kawano, Zehang Richard Li, Paul A. ParkerAI Agent EvaluationPrediction-Powered Inference

  31. A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

    Sep 17, 2026Shaina Raza, Ahmed Y. Radwan, Imran Liaquat +1LLM EvaluationAI Agent Evaluation

  32. Do AI Agents Understand Computer Architecture?

    Sep 16, 2026Ambika Sharan, Grigory Chirkov, Soheil AbbaslooElectronic Design AutomationAI Agent Evaluation

  33. Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

    Sep 16, 2026Xinshuai Guo, Junjie Wu, Dolly Deng +4AI Agent EvaluationAI Agent Benchmarks

  34. WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation

    Sep 14, 2026Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy +1Synthetic Data GenerationAI Agent Evaluation

  35. DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents

    Sep 13, 2026Zhichao Shi, Xuhui Jiang, Wenjie Zhang +5LLM Agent EvaluationAI Agent Evaluation