LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges

    Apr 21, 2026Ali Al-Kaswan, Maksim Plotnikov, Maxim Hájek +3LLM Agent EvaluationAI Agent Security Benchmarks

  2. AutomationBench

    Apr 21, 2026Daniel Shepard, Robin SalimansLLM Agent EvaluationBusiness Process Automation

  3. AI scientists produce results without reasoning scientifically

    Apr 20, 2026Martiño Ríos-García, Nawaf Alampara, Chandan Gupta +5AI Agent ReliabilityScientific Reasoning

  4. AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

    Apr 20, 2026Wentao Shi, Yu Wang, Yuyang Zhao +8LLM-as-a-JudgeLLM Agent Evaluation

  5. WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models

    Apr 20, 2026Xinping Lei, Xinyu Che, Junqi Xiong +16LLM Agent EvaluationCode Language Models

  6. Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots

    Apr 20, 2026Shiquan Zhang, Tianyi Zhang, Le Fang +3Mobile GUI AutomationLLM Agent Evaluation

  7. Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity

    Apr 19, 2026Leon Engländer, Sophia Althammer, Ahmet Üstün +2LLM Agent EvaluationTool-Using Agents

  8. HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents

    Apr 19, 2026Chao Jin, Wenkui Yang, Hao Sun +6GUI AgentsHallucination in Language Models

  9. Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks

    Apr 18, 2026Tyler H. Merves, Michael H. Conaway, Joseph M. Escobar +2LLM Agent EvaluationCybersecurity

  10. SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems

    Apr 17, 2026Hikaru Shindo, Hanzhao Lin, Lukas Helff +2LLM Agent EvaluationAI Agent Benchmarks

  11. Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis

    Apr 16, 2026Ke Zhang, Patricio Gallardo, Maziar Raissi +1Code TranslationLLM Agent Evaluation

  12. CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

    Apr 16, 2026Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita +2Social DilemmasCooperative Game Theory

  13. HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks

    Apr 16, 2026Fan Cui, Hongyuan Hou, Zizhang Luo +2Automated Program RepairElectronic Design Automation

  14. DR3^{3}-Eval: Towards Realistic and Reproducible Deep Research Evaluation

    Apr 16, 2026Qianqian Xie, Qingheng Xiong, He Zhu +16LLM Agent EvaluationLong-Horizon Agent Evaluation

  15. UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

    Apr 13, 2026Yijuan Liang, Xinghao Chen, Yifan Ge +5LLM Agent EvaluationAI Agent Benchmarks

  16. Agentic Tool Use in Large Language Models

    Apr 1, 2026Jinchao Hu, Meizhi Zhong, Kehai Chen +2LLM Agent EvaluationTool-Using Agents

  17. SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

    Mar 31, 2026Kuangshi Ai, Haichao Miao, Kaiyuan Tang +13Scientific VisualizationData Visualization

  18. LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

    Mar 20, 2026Xiang Long, Li Du, Yilong Xu +11Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  19. $OneMillion-Bench: How Far are Language Agents from Human Experts?

    Mar 9, 2026Yang Liu, Jiaqi Li, Jun Bai +20LLM Agent EvaluationAI Agent Benchmarks