LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

    May 14, 2026Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang +3Conversational MemoryLLM Agent Evaluation

  2. Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows

    May 14, 2026Zixin Chen, Peng Liu, Rui Sheng +6Intelligent Tutoring SystemsEducational Assessment

  3. ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

    May 13, 2026Yuxiang Lai, Peng Xia, Haonian Ji +8Computer-Use Agent BenchmarksLLM Agent Evaluation

  4. PBT-Bench: Benchmarking AI Agents on Property-Based Testing

    May 13, 2026Lucas Jing, Xinqi Wang, Liao Zhang +1Automated Software TestingSoftware Engineering Agents

  5. TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate

    May 13, 2026Erica Zhang, Fangzhao Zhang, Aneesh Pappu +5Multi-Agent NegotiationAutomated Negotiation

  6. AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

    May 13, 2026Priyam Sahoo, Gaurav Mittal, Xiaomin Li +4Agent Failure AnalysisSoftware Engineering Agents

  7. Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents

    May 13, 2026Harshita Chopra, Kshitish Ghate, Aylin Caliskan +3Human Behavior SimulationLLM Agent Evaluation

  8. Mechanism Plausibility in Generative Agent-Based Modeling

    May 12, 2026Patrick Zhao, David Huu Pham, Nicholas VincentLLM Agent EvaluationAgent-Based Modeling

  9. Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

    May 12, 2026Muhammad Bilal, Jon Crowcroft, Ruizhi Wang +2LLM Agent EvaluationAI Agent Security

  10. MEME: Multi-entity & Evolving Memory Evaluation

    May 12, 2026Seokwon Jung, Alexander Rubinstein, Arnas Uselis +2Agent MemoryPersistent Memory for Language Models

  11. CAX-Agent: A Lightweight Agent Harness for Reliable APDL Automation

    May 12, 2026Chenying Lin, Yichen Hai, Yi He +3LLM Agent OrchestrationLLM Agent Evaluation

  12. Counterfactual Trace Auditing of LLM Agent Skills

    May 12, 2026Xiaolin Zhou, Jinbo Liu, Li Li +2Software Engineering AgentsSoftware Engineering Benchmarks

  13. Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations

    May 12, 2026Junjue Wang, Weihao Xuan, Heli Qi +7LLM Agent EvaluationGeospatial Reasoning

  14. CTFusion: A CTF-based Benchmark for LLM Agent Evaluation

    May 12, 2026Dongjun Lee, Ga-eun Bae, Insu YunBenchmark ContaminationLLM Agent Evaluation

  15. WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

    May 11, 2026Shuangrui Ding, Xuanlang Dai, Long Xing +14Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  16. Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

    May 11, 2026Tz-Huan Hsu, Jheng-Hong Yang, Jimmy LinTool-Augmented Language Model AgentsAgentic Search

  17. ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

    May 11, 2026Yuanyang Li, Xue Yang, Longyue Wang +2LLM Agent EvaluationAI Agent Benchmarks

  18. CrackMeBench: Binary Reverse Engineering for Agents

    May 11, 2026Isaac David, Arthur GervaisSoftware Reverse EngineeringLLM Agent Evaluation

  19. A Reflective Storytelling Agent for Older Adults: Integrating Argumentation Schemes and Argument Mining in LLM-Based Personalised Narratives

    May 11, 2026Jayalakshmi Baskar, Vera C. Kaelin, Kaan Kilic +1Argument MiningLLM Agent Evaluation

  20. Agentic Performance at the Edge: Insights from Benchmarking

    May 11, 2026Shiqiang Wang, Herbert WoisetschlägerLLM Agent EvaluationAI Agent Evaluation