LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

    Jun 1, 2026Kuan Li, Shuo Zhang, Huacan Wang +12LLM Agent EvaluationInternet of Things

  2. Characterization of Multi-Model Agentic AI Systems on General Tasks via Trace-Driven Simulation

    Jun 1, 2026Donghwan Kim, Prakhar Singh, Younghoon Min +3LLM Agent EvaluationAI Agent Benchmarks

  3. TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning

    May 31, 2026Yaxuan Kong, Qingren Yao, Yuqi Nie +7Time Series ReasoningLLM Agent Evaluation

  4. Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability

    May 31, 2026Xianyou Li, Weiran Yan, Yichao Wu +4LLM Inference EfficiencyMulti-Agent LLM Systems

  5. A New Framework for Cybersecurity Refusals in AI Agents

    May 31, 2026Eliot Krzysztof Jones, Mateusz Dziemian, Matt Fredrikson +1LLM Agent EvaluationLLM Refusal Behavior

  6. TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

    May 31, 2026Weiyi Chen, Shuaixiong Wang, Ziyun Gao +5Spatiotemporal ReasoningLLM Agent Evaluation

  7. Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

    May 30, 2026Xiqi Hao, Zengqing Wu, Yu-Xuan Qiu +4LLM Agent EvaluationMulti-Agent Debate

  8. ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment

    May 30, 2026Qiuyu Tian, Haojie Yin, Yingce Xia +2LLM Agent EvaluationAI Agent Benchmarks

  9. BAGEN: Are LLM Agents Budget-Aware?

    May 29, 2026Yuxiang Lin, Zihan Wang, Mengyang Liu +9LLM Agent EvaluationAI Agent Benchmarks

  10. Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study

    May 29, 2026Xiaonan Xu, Wenjing WuLLM Agent EvaluationLLM Agent Skill Learning

  11. Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation

    May 29, 2026Grégoire Martinon, Ibrahim Merad, Mohammed RakiLLM-Assisted AnnotationLLM Agent Evaluation

  12. TUX: Measuring Human--AI Tacit Understanding

    May 29, 2026Yueshen Li, Hanyi Min, Vedant Das Swain +1LLM AlignmentLLM Agent Evaluation

  13. BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

    May 29, 2026Srivatsa Kundurthy, Clara Na, Colton Moraine +6LLM Agent EvaluationAI Agent Benchmarks

  14. MAVEN: Improving Generalization in Agentic Tool Calling

    May 29, 2026Omkar Ghugarkar, Vishvesh Bhat, Muhammad Ahmed Mohsin +1Logical ReasoningLLM Agent Evaluation

  15. On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

    May 28, 2026Tong Liu, Cheng Qian, Matej Cief +4Tool-Use EvaluationLLM Agent Evaluation

  16. Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

    May 28, 2026Minhua Lin, Juncheng Wu, Zijun Wang +14LLM Agent EvaluationAgent Harness Optimization

  17. Can LLM Teams Play What? Where? When?

    May 28, 2026Anastasia Kotelnikova, Viktor Byzov, Maria Dolzhenkova +1Multi-Agent LLM SystemsLLM Agent Evaluation

  18. SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents

    May 28, 2026Grant Hamblin, Kevin Song, Zhanda Zhu +4Requirements EngineeringSoftware Engineering Agents

  19. Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams

    May 28, 2026Yinsicheng Jiang, Liang Cheng, Yeqi Huang +6Multi-Agent LLM SystemsLLM Agent Evaluation

  20. HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?

    May 28, 2026Weihan Peng, Chenxu Zhang, Qianao Wang +7Personality Modeling in Language ModelsHuman Behavior Simulation

  21. Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents

    May 28, 2026Alejandra Zambrano, Sara Vera Marjanovic, Imene Kerboua +2LLM Agent EvaluationLLM Agent Planning

  22. Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

    May 28, 2026Minyang Hu, Bo Yang, Zhinuo Zhou +4LLM Agent EvaluationAgent Trajectory Analysis