LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments

    Jul 1, 2026Ananya Mantravadi, Harshit Rajgarhia, Prasanna Desikan +1LLM Agent EvaluationAI Agent Benchmarks

  2. MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

    Jul 1, 2026Zhishang Xiang, Zerui Chen, Yunbo Tang +5Agent MemoryLLM Sycophancy

  3. Calibrating the Instrument: Controllability of an LLM-Driven Synthetic Population

    Jul 1, 2026Mirko Degli EspostiHuman Behavior SimulationSynthetic Data Generation

  4. Evaluating Agentic Harness Systems for Autonomous Computational Pathology

    Jul 1, 2026Jie Lin, Zongyi Chen, Qiaoling Zheng +6Computational PathologyLLM Agent Evaluation

  5. Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

    Jun 30, 2026Ben Slater, Matteo G. Mecattaf, Lucy G. Cheke +2Theory of MindLLM Agent Evaluation

  6. ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

    Jun 30, 2026Kaiwen Xiong, Haonian Ji, Shi Qiu +4LLM Agent OrchestrationLLM Agent Evaluation

  7. What Drives Interactive Improvement from Feedback?

    Jun 29, 2026Bartłomiej Cupiał, Jan Łojek, Mikołaj Garstecki +3LLM Self-RefinementLLM Agent Evaluation

  8. Collective cooperation without individual fidelity in LLM agents

    Jun 29, 2026Henrique Ferraz de Arruda, Carlos Gracia Lázaro, Alberto Aleta +1Multi-Agent LLM SystemsHuman Behavior Simulation

  9. CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

    Jun 29, 2026Bo Qu, Mingguang ChenQuantitative FinanceLLM Agent Evaluation

  10. A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents

    Jun 29, 2026Liu ZewenLLM EvaluationLLM Agent Evaluation

  11. Cognitive World Models for Process-Level Social Influence Evaluation

    Jun 28, 2026Minghui Ma, Bin Guo, Han Wang +4LLM EvaluationCognitive Modeling

  12. A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

    Jun 28, 2026Yuanhong Cai, Xiaohui Nie, Kanglin Yin +8Software EngineeringLLM Agent Evaluation

  13. Agentic Abstention: Do Agents Know When to Stop Instead of Act?

    Jun 27, 2026Han Luo, Bingbing Wen, Lucy Lu WangLLM Agent EvaluationLLM Abstention

  14. SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

    Jun 27, 2026My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak +4LLM Agent EvaluationAI Agent Evaluation

  15. TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

    Jun 26, 2026Shoufa Chen, Luyuan Wang, Xuan Yang +7Computer-Use Agent BenchmarksComputer-Use Agents

  16. An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations

    Jun 26, 2026Dihia Falouz, Aida Douaibia, Amine Bechar +3Retrieval-Augmented GenerationLLM Agent Evaluation

  17. ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents

    Jun 26, 2026Shijing Hu, Liang Liu, Zhu Meng +1Privacy AuditingTool-Use Evaluation

  18. Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game

    Jun 26, 2026Subhendu Bhandary, Federico Carucci, Christos Charalambous +6Multi-Agent LLM SystemsDeception in Language Models

  19. When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

    Jun 26, 2026Yiling Tao, Shihan Deng, Meiling Tao +3Agentic SearchLLM Agent Evaluation

  20. OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

    Jun 25, 2026Aoyang Fang, Yifan Yang, Jin'ao Shang +7LLM Agent EvaluationCausal Attribution

  21. Semantic Early-Stopping for Iterative LLM Agent Loops

    Jun 25, 2026Sahil ShrivastavaEarly StoppingLLM Agent Evaluation

  22. What the LLM Should Not Say: Boundary-Aware Context Grounding for A Seven-Channel EEG Agent

    Jun 25, 2026Zhiyuan Xu, Yueqing Dai, Junling Li +1ElectroencephalographyLLM Grounding

  23. Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

    Jun 25, 2026Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi +1LLM Agent EvaluationLLM Agent Security

  24. Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

    Jun 24, 2026Ching-Yu Lin, Yifan LiuLLM Agent EvaluationLLM Agent Reliability