Agent Evaluation

Momentum

23 papers in the last four weeks, up 229% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 97

All topics
CardsList
  1. Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation

    Oct 5, 2026Dishu Yang, Qi Su, Hongbo Qin +1Benchmark AuditingAI Agent Auditing

  2. Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

    Oct 1, 2026Ngoc Phuoc An Vo, Aarya Doshi, Vadim SheininTool-Use EvaluationAI Agent Evaluation

  3. Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

    Oct 1, 2026Chengguang Gan, Zimeng He, Yoshihiro Tsujii +3AI Agent AuditingWeb Agent Benchmarks

  4. Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

    Sep 30, 2026Michael Hardy, Ruhana Azam, Anka Reuel +2AI Agent EvaluationAgent Evaluation

  5. Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

    Sep 30, 2026Priyanath Maji, Spandan Ghose ChowdhuryThompson SamplingAI Agent Evaluation

  6. CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings

    Sep 28, 2026Shama Gupta, Hoang H Nguyen, Chelsea Huang +2Code-Switching Speech RecognitionASR Evaluation

  7. Reward Hacking Challenges Oversight of Autonomous Research Agents

    Sep 23, 2026Yue Huang, Zhangchen Xu, Yuchen Ma +12Reward HackingAI Agent Auditing

  8. Skill-based Agentic Evaluation for Real-time Data Science Tasks

    Sep 15, 2026Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal +4LLM-as-a-JudgeLLM Agent Evaluation

  9. Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

    Sep 11, 2026Hanhua Hong, Yizhi Li, Luu Gia Huy +3AI Agent EvaluationAgent Evaluation

  10. Qiushi Engine on AstaBench E2E-Bench-Hard

    Sep 8, 2026Wenhao Li, Shuxing Yang, Fujia Chen +13Agent EvaluationAI Agent Benchmarks

  11. APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents

    Sep 7, 2026Jintian Feng, Long Chen, Xiao Yu +7Mobile GUI AutomationGUI Agents

  12. ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

    Sep 2, 2026Peiying Zhu, Sidi ChangClaim VerificationData Provenance

  13. DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

    Sep 1, 2026Haoyuan Shi, Mingtao Chen, Shuo Jiang +12Agent EvaluationTemporal Consistency in Video Generation

  14. Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

    Aug 25, 2026Anupam Purwar, Shashank Singh, Kritika SrivastavaVoice Agent EvaluationLLM-as-a-Judge

  15. QuoteBench: How Matched Scores Can Hide Command-Path Failures

    Aug 13, 2026Shangao Li, Yao Zhang, Volker Tresp +1LLM Agent EvaluationAgent Evaluation

  16. DuplexWorld: Can voice agents help you get through the day?

    Aug 11, 2026Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli +3Voice Agent EvaluationSpoken Dialogue Systems