LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning

    Jun 10, 2026Yizhou Chi, Eric Chamoun, Zifeng Ding +1Event Time PredictionLLM Agent Evaluation

  2. Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment

    Jun 10, 2026Derek Yohn, Luke Flancher, Mirajul Islam +1Software Vulnerability DetectionLLM Agent Evaluation

  3. SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

    Jun 10, 2026Zhiyu Chen, Zihan Guo, Bo Huang +4LLM Agent Evaluation

  4. AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable

    Jun 9, 2026Meysam Alizadeh, Fabrizio Gilardi, Mohsen Mosleh +1AI Agent ReliabilityAI Coding Agents

  5. From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents

    Jun 9, 2026Yifan Li, Shengbin Yue, Boyu Feng +6LLM Agent EvaluationAI Agent Benchmarks

  6. T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

    Jun 9, 2026Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin +12Customer Support AutomationLLM Agent Evaluation

  7. Skill Coverage: A Test Adequacy Metric for Agent Skills

    Jun 9, 2026Boyin Tan, Xiaowei Huang, Youcheng SunLLM Agent EvaluationAI Agent Benchmarks

  8. STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

    Jun 9, 2026Sirui Liang, Bohan Yu, Peiyu Wang +8Computer-Use Agent BenchmarksLLM Agent Evaluation

  9. Early-Token Confidence Predicts Reasoning Quality in Multi-Agent LLM Debate

    Jun 9, 2026Ali Keramati, Justin Cheok, Jacob Horne +1LLM Agent EvaluationMulti-Agent Debate

  10. The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge

    Jun 9, 2026Ali Keramati, Justin Cheok, Jacob Horne +1LLM Agent EvaluationMulti-Agent Debate

  11. H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions

    Jun 8, 2026Shiping Zhu, Yibo Yang, Zhengyang Wang +3Multimodal MemoryLLM Agent Memory

  12. REFLECT: Intervention-Supported Error Attribution for Silent Failures in LLM Agent Traces

    Jun 8, 2026Xiaofeng Lin, Yingxu Wang, Tung Sum Thomas Kwok +4LLM Agent EvaluationLLM Agents

  13. PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting

    Jun 7, 2026Youran Sun, Xingyu Ren, Kejia Zhang +2Multi-Agent OrchestrationLLM Agent Evaluation

  14. Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework

    Jun 7, 2026Aman Gupta, Kevin Rossell, Edesio Alcobaça +8Customer Support AutomationLLM Agent Evaluation

  15. Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems

    Jun 7, 2026Kuncan Wang, Ziting Wang, Peizhuo Lv +4LLM Agent EvaluationCybersecurity

  16. Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

    Jun 6, 2026Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker +7Multi-Agent CoordinationLLM Agent Evaluation

  17. To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation

    Jun 6, 2026John Chen, Sihan Cheng, Can Gurkan +1Moral Reasoning in Language ModelsLLM Agent Evaluation

  18. Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts

    Jun 6, 2026Po-Ya Angela Wang, Chinmaya Mishra, Aslı Özyürek +2Repeated GamesLLM Agent Evaluation

  19. Does Persona Make LLMs K-pop Fans? A Pilot Study of LLM-Based Online Concert Audience Agents

    Jun 5, 2026Kirak Kim, Hyojin Kim, Yejin Son +2Multi-Agent LLM SystemsLLM Agent Evaluation