LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management

    Jul 22, 2026Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko +1LLM Agent EvaluationAI Agent Benchmarks

  2. LLMs Get Lost in Evolving User Intent

    Jul 22, 2026Jihoon Tack, Philippe Laban, Jennifer NevilleLLM Agent EvaluationAI Agent Benchmarks

  3. OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

    Jul 22, 2026Qiyuan Liu, Tingfeng Hui, Kun Zhan +2LLM Agent EvaluationAI Agent Safety

  4. Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning

    Jul 22, 2026Mohit Jiwatode, Ronja Fuchs, Robin Schmöcker +2Game-Playing AgentsLLM Agent Evaluation

  5. Agents in the Wild: Where Research Meets Deployment

    Jul 21, 2026Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz +3AI Agent ReliabilityMulti-Agent LLM Systems

  6. MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

    Jul 21, 2026Guofeng Zhang, Yizeng Quan, Huaiyi Fang +6LLM Agent EvaluationMulti-Turn Dialogue Evaluation

  7. Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents

    Jul 21, 2026SangJin Park, Myungsub Choi, Jineok Kim +1LLM Agent EvaluationAI Agent Security

  8. Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

    Jul 20, 2026Pin Qian, Su Wang, Yihang Chen +5LLM Agent EvaluationAI Agent Benchmarks

  9. Operational Hallucination and Safety Drift in AI Agents

    Jul 20, 2026Shasha Yu, Fiona Carroll, Barry L. BentleyLLM Agent EvaluationAI Agent Safety

  10. LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications

    Jul 20, 2026Daniela Rojas, Abdulwahab Albassam, Aidan G. Leung +11LLM Agent OrchestrationLLM Agent Evaluation

  11. Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data

    Jul 20, 2026Nursultan Askarbekuly, Mohamad Al Mdfaa, Ahmed Helaly +2AI Agent ReliabilityReward Hacking

  12. Stress Testing Concept Erasure with Large Language Model Agents

    Jul 20, 2026Yuyang Xue, Feng Chen, Zhihua Liu +4LLM Agent EvaluationLLM Safety Evaluation

  13. FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

    Jul 20, 2026Jiacheng Ding, Cong Guo, Jason XuLLM Agent EvaluationForecasting Benchmarks

  14. ProEvent: An Event-centric Benchmark for Proactive Agents

    Jul 20, 2026Guanzhen Li, Liangming Pan, Leye WangLLM Agent EvaluationProactive Assistance

  15. Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows

    Jul 20, 2026Jinyuan Deng, Zhengrui Chen, Xufeng Wei +4Electronic Design AutomationLLM Agent Evaluation

  16. Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions

    Jul 19, 2026Chen Xia, Zexi Kuang, Yuqing HuLLM GroundingHuman Behavior Simulation

  17. Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning

    Jul 19, 2026Zhihao Liu, Tianyu Wang, Xi Vincent Wang +1Multi-Agent OrchestrationLLM Agent Evaluation

  18. OTAP: Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories

    Jul 19, 2026Babak Barazandeh, Subhabrata Majumdar, George MichailidisLLM Agent EvaluationAI Agent Evaluation

  19. Lomekwi: Resource-Bounded Tool Discovery in LLM Agents

    Jul 18, 2026Roshan Klein-Seetharaman, Daniel Wang, Andrew XuLanguage Model Scaling LawsLLM Agent Evaluation

  20. Fantastic Adaptive Taxonomies and How to Use Them

    Jul 17, 2026Mert Cemri, Andrei Cojocaru, Melissa Pan +9Agent Failure AnalysisLLM Agent Evaluation

  21. SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

    Jul 17, 2026Yanze Wang, Pengfei Yao, Tianyi Sun +7LLM Agent EvaluationLLM Agent Skill Retrieval

  22. AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

    Jul 15, 2026Kai Chen, Zichen Ding, Jiaye Ge +20LLM Agent EvaluationAI Agent Benchmarks