LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. JobBench: Aligning Agent Work With Human Will

    May 25, 2026Yuetai Li, Yichen Feng, Zhangchen Xu +21Computer-Use Agent BenchmarksLLM Agent Evaluation

  2. CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

    May 25, 2026Junlin Yang, Dylan Zhang, Xiangchen Song +7Structural Causal ModelsLLM Agent Evaluation

  3. When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

    May 25, 2026Liyun Zhang, Jiayi GuoCoT ReasoningLLM Agent Evaluation

  4. GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

    May 24, 2026Xiang Cheng, Yulan Hu, Lulu Zheng +4AI-Assisted Decision MakingLLM Planning

  5. ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

    May 22, 2026Xianzhong Ding, Yangyang Yu, Changwei Liu +1Tool-Augmented Language Model AgentsAI Coding Agents

  6. SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

    May 22, 2026Yingtie Lei, Zhongwei Wan, Jiankun Zhang +13LLM Agent EvaluationLLM Agent Skill Learning

  7. From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills

    May 22, 2026Zisu Huang, Jingwen Xu, Yifan Yang +13LLM Agent EvaluationLLM Agent Skill Learning

  8. OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

    May 22, 2026Jiahao Ying, Boxian Ai, Wei Tang +2LLM Agent EvaluationAI Agent Benchmarks

  9. MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

    May 21, 2026Thomson Yen, Julian Poeltl, Harshith Srinivas Gear +9LLM Agent EvaluationAI Agent Benchmarks

  10. Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

    May 21, 2026Asaf Yehudai, Lilach Eden, Michal Shmueli-ScheuerLLM Agent EvaluationAI Agent Benchmarks

  11. SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations

    May 21, 2026Shuaiqi Wang, Aadyaa Maddi, Zinan Lin +1Tool-Use EvaluationLLM Agent Evaluation

  12. TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

    May 21, 2026Zhaoyang Chu, Jiarui Hu, Xingyu Jiang +8Computer-Use Agent BenchmarksTerminal Agents

  13. Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions

    May 21, 2026Jianan Ma, Xiaohu Du, Ruixiao Lin +8LLM Agent EvaluationLLM Agent Security

  14. SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?

    May 20, 2026Kevin Han, Renfei Zhang, Kathy Wei +3LLM Agent EvaluationAI Agent Benchmarks

  15. What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

    May 20, 2026Mahdi Naser Moghadasi, Faezeh GhaderiBenchmark AuditingLLM Auditing

  16. MemGym: a Long-Horizon Memory Environment for LLM Agents

    May 20, 2026Wujiang Xu, Yu Wang, Kai Mei +8LLM Agent MemoryLLM Agent Evaluation

  17. AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

    May 19, 2026Parsa Mazaheri, Kasra MazaheriAgent Failure AnalysisComputer-Use Agent Benchmarks

  18. Modeling Emotional Dynamics in Agent-to-Agent Interactions on Moltbook

    May 19, 2026Syed Mhamudul Hasan, Abdur R. ShahidMulti-Agent LLM SystemsLLM Agent Evaluation

  19. Probing an Embodied LLM: When Higher Observation Fidelity Hurts Problem Solving

    May 19, 2026Oussama Zenkri, Oliver BrockLarge Language Model-Based Robot PlanningLLM Agent Evaluation

  20. When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity

    May 19, 2026Samuel Jacob Chacko, James Hugglestone, Chashi Mahiul Islam +1LLM Agent EvaluationLLM Agent Skill Learning

  21. EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

    May 19, 2026Gioele Molinari, Florian Felten, Soheyl Massoudi +1Multi-Agent LLM SystemsLLM Agent Evaluation

  22. Toward User Comprehension Supports for LLM Agent Skill Specifications

    May 19, 2026Zikai Alex WenLLM Agent Evaluation

  23. PAVE: A Cognitive Architecture for Legitimate Violation in Generative Agent Societies

    May 19, 2026Ahmad Yehia, Abduallah Mohamed, Kun Qian +4Cognitive Architectures for AI AgentsLLM Agent Evaluation

  24. Agentic Trading: When LLM Agents Meet Financial Markets

    May 19, 2026Yihan Xia, Panpan You, Taotao Wang +4LLM Agent EvaluationFinance