AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 344

All topics
CardsList
  1. A 3D Characterization Framework for Intelligent Sequential Decision Making

    Oct 8, 2026Sadig Gojayev, Carolina FortunaSequential Decision MakingAI Agent Evaluation

  2. Closed-loop evaluation of LLM agents for embedded software development

    Oct 8, 2026Jorge García-Carrasco, Sergio García-Carrasco, Alejandro Maté +1AI Coding AgentsLLM Agent Evaluation

  3. SciExam for ENSO: Can AI Agents Build Climate Models?

    Oct 7, 2026Yinling Zhang, Langchen Liu, Dongbin Xiu +4AI for ScienceAI Agent Benchmarks

  4. Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

    Oct 7, 2026Tan Yu, Alexander Bukharin, Khushi Bhardwaj +19AI Coding AgentsAI Agent Evaluation

  5. Stale, Misattributed, or Late: Where Personal Memory Fails Before Generation

    Oct 7, 2026Haonan Deng, Park SinchaisriAgent MemoryAgent Memory Management

  6. Shared and structured inputs undermine collective random choice by reasoning AI agents

    Oct 7, 2026Takahiro Ezaki, Naoto Imura, Katsuhiro NishinariAI Agent EvaluationSelection Bias

  7. The AI Evaluation Ecosystem

    Oct 7, 2026Yash Dave, Sang T. Truong, Serena Wang +1AI Agent EvaluationBenchmark Design

  8. Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

    Oct 6, 2026Wenyu Du, Stephen ChungAutonomous Scientific DiscoveryAI Agent Evaluation

  9. Agent Plasticity: Measuring Self-Improvement Through Experience

    Oct 6, 2026Harman Singh, Anton Bakhtin, Rulin Shao +8Self-Improving AgentsLLM Agent Self-Improvement

  10. Inspect Robots: Evaluating the Capabilities and Safety of Embodied AI

    Oct 5, 2026Christopher Leet, Achu Menon, Sravanthi Machcha +12AI Agent EvaluationRobotic Policy Evaluation

  11. MiniCorp: The Last Mile of the AI Agent Firm

    Oct 5, 2026Jingying Zeng, Zhenwei Dai, Jinning Li +6Synthetic Data GenerationAI Agent Evaluation

  12. Code Owns the Simulation, Jev Owns the Evaluation

    Oct 1, 2026Yaodong Yang, Hongyao Tang, Yi Ma +4AI Agent EvaluationRobotic Policy Evaluation

  13. Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

    Oct 1, 2026Ngoc Phuoc An Vo, Aarya Doshi, Vadim SheininTool-Use EvaluationAI Agent Evaluation

  14. Agents Are Systems, Not Models: Rethinking Agentic Evaluation

    Oct 1, 2026Luis Wiedmann, Leander Girrbach, Cordelia Schmid +1AI Agent ReliabilityAI Agent Evaluation

  15. Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

    Oct 1, 2026Haotian Chen, Bowen Ye, Yuning Zhang +1Multi-Agent LLM SystemsAI Agent Auditing

  16. Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles

    Oct 1, 2026Alexei N. Skurikhin, Emily M. Taylor, Nathan A. DeBardelebenLLM Inference EfficiencyLanguage Model Generation Evaluation

  17. Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

    Sep 30, 2026Michael Hardy, Ruhana Azam, Anka Reuel +2AI Agent EvaluationAgent Evaluation

  18. CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

    Sep 30, 2026Chenmu Zhang, Levi Felix, Jun-Jie Zhang +6AI Agents for Scientific DiscoveryAI Agent Evaluation

  19. No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

    Sep 30, 2026Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian +1LLM Red TeamingAI Agent Evaluation

  20. Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

    Sep 30, 2026Priyanath Maji, Spandan Ghose ChowdhuryThompson SamplingAI Agent Evaluation

  21. AgBench: Agentic AI Benchmarks for Personal AI Devices

    Sep 29, 2026Yizhou Han, Di Wu, Dhananjay Saikumar +1AI Agent ReliabilityAI Agent Evaluation

  22. CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?

    Sep 29, 2026Yue Pan, Jiawei Li, Ziyuan Zhang +2LLM EvaluationLLM-as-a-Judge