LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

    May 3, 2026Xiao JiaAI Agent ReliabilityLLM Agent Evaluation

  2. How Personas Can Influence Agents to Play Split or Steal

    May 3, 2026Carlos Leon, Alexandre Rodrigues, Pedro Gamito +1Game-Playing AgentsLLM Agent Evaluation

  3. LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation

    May 2, 2026Dong Xu, Jialun Cao, Guozhao Mo +9LLM Agent EvaluationFormal Verification

  4. AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

    May 1, 2026Ranit Karmakar, Jayita ChatterjeeComputer-Use Agent BenchmarksLLM Agent Evaluation

  5. Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows

    Apr 30, 2026Chenxin Li, Zhengyang Tang, Mingxin Huang +8LLM Agent EvaluationAI Agent Benchmarks

  6. Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception

    Apr 30, 2026Neemias B da Silva, Rodrigo Minetto, Daniel Silver +1VLM EvaluationPersonality Modeling in Language Models

  7. Exploring LLM Agent Designs and Interaction Modalities for Scientific Visualization

    Apr 30, 2026Jackson Vonderhorst, Kuangshi Ai, Haichao Miao +2Scientific VisualizationData Visualization

  8. InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?

    Apr 30, 2026Qiyao Wang, Haoran Hu, Longze Chen +4Web Application GenerationWeb Agent Benchmarks

  9. Evaluating Strategic Reasoning in Forecasting Agents

    Apr 28, 2026Tom Liptay, Dan Schwarz, Rafael Poyiadzi +2LLM Agent EvaluationForecasting Benchmarks

  10. Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital

    Apr 28, 2026T. J. Barton, Chris Constantakis, Patti Hauseman +4AI Agent ReliabilityLLM Agent Evaluation

  11. Agentic clinical reasoning over longitudinal myeloma records: a retrospective evaluation against expert consensus

    Apr 27, 2026Johannes Moll, Jannik Lübberstedt, Christoph Nuernbergk +21HealthcareLLM Agent Evaluation

  12. MarketBench: Evaluating AI Agents as Market Participants

    Apr 26, 2026Andrey Fradkin, Rohit KrishnanMulti-Agent CoordinationLLM Agent Evaluation

  13. ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation

    Apr 26, 2026Boqin Yuan, Yue Su, Renchu Song +2LLM Agent EvaluationLLM Agent Skill Learning

  14. DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

    Apr 26, 2026Nishant Balepur, Malachi Hamada, Varsha Kishore +9Deep Research AgentsLLM Agent Evaluation

  15. Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems

    Apr 24, 2026Mengzhuo Chen, Junjie Wang, Fangwen Mu +4Agent Failure AnalysisMulti-Agent LLM Systems

  16. From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes

    Apr 24, 2026Rubén Garzón, Pauline Baron, Vincent Grari +3LLM Agent EvaluationHuman Behavior Prediction

  17. Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results

    Apr 23, 2026Benjamin Kohler, David Zollikofer, Johanna Einsiedler +2AI Coding AgentsLLM Agent Evaluation

  18. When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors

    Apr 23, 2026Chenghao Yang, Yuning Zhang, Zhoufutu Wen +4Tool-Augmented Language Model AgentsLLM Agent Evaluation

  19. Time Series Augmented Generation for Financial Applications

    Apr 21, 2026Anton Kolonin, Alexey Glushchenko, Evgeny Bochkov +1LLM Agent EvaluationAI Agent Benchmarks