LLM Agent Evaluation

LLM: Large Language Model

Momentum

109 papers in the last four weeks, up 88% on the four weeks before. 1.3% of all new papers.

Jul 6Week of Sep 21

Latest papers 793

All topics
CardsList
  1. Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

    Jun 25, 2026Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi +1LLM Agent EvaluationLLM Agent Security

  2. Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

    Jun 24, 2026Ching-Yu Lin, Yifan LiuLLM Agent EvaluationLLM Agent Reliability

  3. How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

    Jun 24, 2026David Akinpelu, Akintonde Abbas, Rereloluwa Alimi +1LLM Agent EvaluationAI Agent Benchmarks

  4. SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game

    Jun 24, 2026Yeqi Feng, Yuxin Chen, Tianxing HeMulti-Agent NegotiationGame Theory

  5. Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents

    Jun 24, 2026Yuxin Wang, Paul Thomas, Zhiwei Yu +5Conversational MemoryLLM Agent Evaluation

  6. Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing

    Jun 24, 2026Liwei Yu, Shuo Li, Ming Zhou +2LLM Agent EvaluationAI Agent Security Benchmarks

  7. Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

    Jun 23, 2026Ali Pourghasemi Fatideh, Wilder Baldwin, Maria Dhakal +2Requirements EngineeringLLM Evaluation

  8. Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

    Jun 23, 2026Ayan Antik Khan, Harsh Kohli, Yuekun Yao +2LLM Agent EvaluationMechanistic Interpretability

  9. When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure

    Jun 22, 2026Victor Lockwood, Hasan Mahmud, Mohammad Javad Khojasteh +2LLM Agent EvaluationHuman-Robot Interaction

  10. Counsel: A Meta-Evaluation Dataset for Agentic Tasks

    Jun 19, 2026Sashank Pisupati, Henry Broomfield, Eujeong Choi +5LLM-as-a-JudgeLLM Agent Evaluation

  11. Trip+: Benchmarking Agents in Personalized Interactive Travel Planning

    Jun 19, 2026Junle Chen, Wei Chen, Yehong Xu +6LLM Agent EvaluationAI Agent Benchmarks

  12. AgentMeter: Evaluating Model-CLI Matching for CLI-Based Local Task-Solving Agents

    Jun 19, 2026Han Chi, Jiaxin Qi, Yan Cui +2LLM Agent EvaluationAI Agent Benchmarks

  13. Agentic Time Machine as an Infrastructure for Future-Event Forecasting

    Jun 19, 2026Jingyi Chai, Bingyang Zheng, Xiangrui Liu +5Multi-Agent LLM SystemsLLM Agent Evaluation

  14. Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems

    Jun 18, 2026Zewen LiuMulti-Agent LLM SystemsLLM Agent Evaluation

  15. ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?

    Jun 18, 2026Jiajun Li, Mingshu Cai, Yixuan Li +5LLM Agent EvaluationAI Agent Benchmarks

  16. Benchmarking Agentic Review Systems

    Jun 18, 2026Dang Nguyen, Wanqing Hao, Yanai Elazar +1LLM Agent EvaluationAI Agent Benchmarks

  17. Recursive Self-Evolving Agents via Held-Out Selection

    Jun 17, 2026Michael Nguyen, Quoc Nguyen, Paul VuongLLM Agent EvaluationLLM Agent Workflow Optimization