LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

    Oct 8, 2026Kaiser Sun, Bernal Jimenez Gutierrez, Hongjun Liu +4LLM Agent EvaluationEpistemic Uncertainty

  2. DataSense-Bench: The First Step Toward an AI Scientist

    Oct 8, 2026Yudi Zhang, Mingyu Cao, Lu Yin +2Training Data SelectionAI Agent Benchmarks

  3. A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

    Oct 8, 2026Ming Chen, Rong-Xi Tan, Ke Xue +8Black-Box OptimizationAI Agent Benchmarks

  4. Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models

    Oct 8, 2026Youwei Feng, Yitong Zhang, Yuetong Liu +1LLM Agent SafetyLLM Guardrails

  5. Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System

    Oct 8, 2026Younghwan Joo, Sung-il KimLLM AgentsLLM Agent Evaluation

  6. Evidence-Traceable Dynamic Interviewer Architecture for Expertise-Adaptive Qualitative Interviews Using Local LLMs

    Oct 8, 2026Aisvarya Adeseye, Jouni Isoaho, Adeyemi Adeseye +2Conversational AgentsLLM Agent Evaluation

  7. Evaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility Workflows

    Oct 8, 2026Jorge García-Carrasco, Javier Sanchis, Alejandro Reina-Reina +2LLM Agent EvaluationLLM Agent Reliability

  8. Closed-loop evaluation of LLM agents for embedded software development

    Oct 8, 2026Jorge García-Carrasco, Sergio García-Carrasco, Alejandro Maté +1AI Coding AgentsLLM Agent Evaluation

  9. Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents

    Oct 7, 2026Obada KraishanTool-Using AgentsLLM Agent Reliability

  10. LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

    Oct 7, 2026Jun Zhao, Leiming Fu, Yanbo Wen +9LLM Agent EvaluationAI Agent Benchmarks

  11. Coding-Agent Benchmarks Should Match Their Users' Task Flows

    Oct 7, 2026Igor Slinko, Yaroslav Golubev, Sergey TitovLLM Agent EvaluationAI Agent Benchmarks

  12. Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs

    Oct 7, 2026Taehyeon Yun, Dongho Kim, Geonwoo Kim +3Digital ForensicsData Provenance

  13. DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis

    Oct 7, 2026Yizhi Song, Hang Ni, Weijia Zhang +1Data Analysis AgentsAI Agent Benchmarks

  14. GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks

    Oct 6, 2026Gabriel Diaz-Ireland, Diego Prieto-Herráez, Mario García Peces +3Tool-Use EvaluationLLM Agent Evaluation

  15. Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

    Oct 6, 2026Panagiotis Kasnesis, Christos Chatzigeorgiou, Lazaros Toumanidis +1Small Language ModelsLLM Agent Evaluation

  16. Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

    Oct 6, 2026Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas +1LLM Agent EvaluationLong-Horizon Agent Evaluation

  17. DecepEval: A Benchmark for Evaluating Deception in LLM Agents

    Oct 6, 2026Yiming Xu, Hongyue Yu, Beihua Yang +8Deception in Language ModelsLLM Agent Evaluation

  18. Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents

    Oct 6, 2026Brendan King, Farima Fatahi Bayat, Jean-Flavien Bussotti +2Probability CalibrationConfidence Estimation in Language Models

  19. An Empirical Study of Agent Skills' Downstream Utility

    Oct 6, 2026Yu Cheng, Dehai Zhao, Zhongxin Liu +3LLM Agent EvaluationLLM Agent Skill Retrieval

  20. ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

    Oct 6, 2026Haizhong Zheng, Yizhuo Di, Ranajoy Sadhukhan +2Continual Learning for LLM AgentsLLM Agent Evaluation

  21. From Evidence to Action: How Tool-Using Agents Fail

    Oct 6, 2026Hongzhan Lin, Shidong Cao, Ziyang Luo +3Agent Failure AnalysisEvidence-Grounded Reasoning

  22. AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy

    Oct 5, 2026Shouju Wang, Haopeng ZhangAI Agent AuditingLLM Agent Evaluation

  23. Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads

    Oct 5, 2026Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona ZahidLLM Inference EfficiencyDisaggregated LLM Serving

  24. Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

    Oct 5, 2026Chubin Zhang, Zhenglin Wan, Xingrui Yu +4LLM Agent EvaluationTool-Using Agents