LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

    Sep 21, 2026Weihang Ding, Junfei ZhanLLM Agent EvaluationAI Agent Benchmarks

  2. BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

    Sep 20, 2026Peng Kuang, Yuchun Fan, Jiangnan Li +7Multilingual Language Model EvaluationLLM Agent Evaluation

  3. Quantifying Overclaiming Propensity in Frontier LLM Agents

    Sep 17, 2026Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo +6AI Agent ReliabilityAI Coding Agents

  4. Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

    Sep 16, 2026Mahsa Amani, Seungeon Lee, Abhisek Dash +9Tool-Augmented Language Model AgentsAgentic Search

  5. ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

    Sep 16, 2026Guosen Wu, Huizhen Huang, Guoxiong Long +2Privacy AuditingLLM Agent Evaluation

  6. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

    Sep 16, 2026Mika Okamoto, Ansel Kaplan ErolAI Agent AuditingLLM Agent Evaluation

  7. Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

    Sep 16, 2026Guojun Zhu, Xunheng Huang, Peng Yin +3Counterfactual EvaluationBenchmark Construction

  8. GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

    Sep 15, 2026Sikun Wang, Yixi Zhou, Lei Fan +1LLM Reasoning with GraphsLLM Agent Evaluation

  9. Verifiable Social Reasoning for LLM Assistants

    Sep 15, 2026Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush +6Multi-Agent LLM SystemsTheory of Mind

  10. ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures

    Sep 15, 2026Muhammad Ashar Ishfaq, Glaucia MeloTheory of MindLLM Agent Evaluation

  11. Skill-based Agentic Evaluation for Real-time Data Science Tasks

    Sep 15, 2026Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal +4LLM-as-a-JudgeLLM Agent Evaluation

  12. When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

    Sep 14, 2026Kaiyuan Liu, Qiuyang Mang, Bo Peng +6LLM Agent EvaluationTest-Time Scaling

  13. Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

    Sep 14, 2026Hazel Mak, Susheel Suresh, Sahil Bhatnagar +3Terminal AgentsLLM Agent Evaluation

  14. Simulating Disengaged Students to Evaluate LLM-based Tutors

    Sep 14, 2026Xianghui Meng, Jionghao LinIntelligent Tutoring SystemsHuman Behavior Simulation

  15. ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents

    Sep 14, 2026Bowen Guan, Zhentao Yin, Yanming ShenLLM Agent EvaluationAI Agent Benchmarks

  16. Debiasing as a Measurement Intervention: Calibrated Ties and Resolution Loss in LLM-as-a-Judge Evaluation

    Sep 14, 2026Liang Zhao, Yong Wang, Jiangzhe ChenLLM-as-a-JudgeLLM Agent Evaluation

  17. Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf

    Sep 14, 2026Yu-Yu Yang, Ti-Rong Wu, Hung Guei +2Social Deduction GamesPragmatic Reasoning in Language Models

  18. Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

    Sep 14, 2026Mykhailo Kozyrev, Andrei Kozyrev, Anton PodkopaevSoftware Engineering AgentsSoftware Engineering Benchmarks

  19. What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework

    Sep 14, 2026Ioannis Prokopiou, Athanasios Aidinis, Panagiotis-Christos Kyrmpatsos +1LLM Self-RefinementLLM Agent Evaluation

  20. MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

    Sep 14, 2026Bosi Wen, Cunxiang Wang, Jiayi Gui +6Software Engineering AgentsAI Coding Agents