LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. Disentangling Task Difficulty from Run-Level Failure in Agent Failure Prediction

    Oct 4, 2026Mohsen EsfandyariDoulabi, Lawrence Arkoh, Biruk Tadesse +4Agent Failure AnalysisLLM Agent Evaluation

  2. DAYJOB: A Benchmark for Long-Horizon Professional Work

    Oct 1, 2026Stephanie Finley, Liudas Panavas, Thomas Mikkelson +12LLM Agent EvaluationLong-Horizon Agent Tasks

  3. DeFA: Dependency-Guided Failure Attribution for LLM Agents

    Oct 1, 2026Bo Deng, Xinlei Zheng, Yi Wei +6Agent Failure AnalysisLLM Agent Evaluation

  4. Empty Commitments: When Agents Promise What They Cannot Deliver

    Oct 1, 2026Jiaqi Tang, Bingyu Shen, Lan Wei +6AI Agent ReliabilityLLM Agent Evaluation

  5. Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems

    Oct 1, 2026Shixuan Li, Wei Yang, Peiyu Zhang +3Multi-Agent LLM SystemsAI Agent Auditing

  6. Finding the Right Fit: Model-Harness Interactions across Agent Tasks

    Oct 1, 2026Yixuan Li, Yiyun Zhou, Yao Long Teng +6Tool-Augmented Language Model AgentsLLM Agent Evaluation

  7. How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

    Sep 30, 2026Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel AbbéAI Coding AgentsLLM Agent Evaluation

  8. DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?

    Sep 30, 2026Frances Liu, Manny Silva, Paige Calvert +2Software EngineeringLLM Agent Evaluation

  9. RealWorldShop: Benchmarking and Improving Conversational Shopping Agents in Real-World E-commerce

    Sep 30, 2026Xinwei Yang, Kelong Mao, Yudong Guo +4LLM Agent EvaluationAI Agent Benchmarks

  10. Talk2Agent: Benchmarking Voice Interfaces for Text Agents

    Sep 30, 2026Terumi Chiba, Guangzhi Sun, Zheqi Yuan +1ASR EvaluationVoice Agent Evaluation

  11. Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent

    Sep 29, 2026Zheyuan Zhang, Mengyuan Chao, Ke Xiao +5LLM AlignmentLLM Agent Evaluation

  12. AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction

    Sep 29, 2026Wentao Liu, Xi Chen, Siyu Song +13LLM Agent EvaluationLLM Agent Training

  13. When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents

    Sep 29, 2026Yuqiao Meng, Luoxi Tang, Yingxue Zhang +2Agent MemoryLLM Agent Evaluation

  14. EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making

    Sep 29, 2026Min Yang, Yichen Pan, Jinghua Piao +3AI-Assisted Decision MakingLLM Agent Evaluation

  15. The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents

    Sep 29, 2026Xueqi Li, Jingjie Ning, Yibo KongLLM Agent EvaluationTool-Using Agents

  16. FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

    Sep 28, 2026Hoyoung Lee, Suyeol Yun, Jack Haverty +17LLM Agent EvaluationRubric-Based Evaluation

  17. PrivacySkills: How Privacy Guidance Shapes Source Selection in LLM Agents

    Sep 28, 2026Lucas Biechy, Cédric Eichler, Héber H. Arcolezi +1LLM Agent Evaluation

  18. Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias

    Sep 28, 2026Yinhong Liu, Zhili Tan, Zilin Wang +1LLM Agent EvaluationLLM Tool Use

  19. CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets

    Sep 28, 2026An Yan, Yu Huo, Zhiwei Shang +2Game TheoryLLM Agent Evaluation

  20. AgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud Continuum

    Sep 28, 2026Michalis Kasioulis, Moysis Symeonides, George Pallis +1LLM Agent OrchestrationLLM Agent Evaluation