LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems

    Sep 28, 2026Xijing Wang, Yinsheng Yao, Jinru Ding +5LLM Agent EvaluationAI Agent Benchmarks

  2. AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

    Sep 28, 2026Chanhee Park, Jeongho Yoon, Sungbin Han +2Multi-Hop QATool-Augmented Language Model Agents

  3. Certified Selective Automation of LLM Agent Evaluation

    Sep 28, 2026Chengguang Gan, Yunhao Liang, Qinghao Zhang +1LLM-as-a-JudgeSelective Prediction

  4. ControlScope: Workflow Revision and Reliability in LLM Agents

    Sep 28, 2026Jingjie Ning, Xueqi Li, Yibo Kong +1LLM Agent EvaluationLLM Agent Reliability

  5. Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures

    Sep 28, 2026Dong Xu, Zhangfan Yang, Jiantao Wu +5LLM Agent EvaluationAI Agent Benchmarks

  6. Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models

    Sep 27, 2026Eduardo Ariño de la Rubia, Szilard PafkaAI Coding AgentsLLM Agent Evaluation

  7. SWE-Game: Can Coding Agents Build the Games We Want?

    Sep 27, 2026Xiaoyu Chen, Lai Wei, Jin Wang +8Software Engineering AgentsAI Coding Agents

  8. AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

    Sep 27, 2026Tianzhuo Yang, Zirui Mi, Yantao Huang +4LLM Agent EvaluationLLM Agents

  9. LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs

    Sep 27, 2026Haochen Luo, Yifan Li, Binh Minh An +4Quantitative FinanceLLM Agent Evaluation

  10. TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

    Sep 27, 2026Dehai Min, Daoan Zhang, Yiming Zeng +13Benchmark ConstructionLLM Agent Evaluation

  11. Cheap, open agents make LLM pollution harder to mitigate

    Sep 25, 2026Raluca Rilla, Anne-Marie Nussberger, Rui Mata +1AI-Generated Text DetectionLLM Agent Evaluation

  12. SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support

    Sep 24, 2026Muhammad Muhtasim Shahriar, Abdullah Mohammad Sayem, Tze Hui Liew +2LLM Agent OrchestrationLLM Agent Evaluation

  13. IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

    Sep 24, 2026Suvradip Paul, Chandra Bhushan, Harsh Sharma +4LLM Safety BenchmarksFinancial Services

  14. Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

    Sep 24, 2026Liqin Ye, Haorui Wang, Fardin Ahmed +8LLM Agent EvaluationForecasting Benchmarks

  15. RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

    Sep 23, 2026Mithil Salunkhe, Haochen Ding, Samridhi Verma +1LLM Agent EvaluationAI Agent Benchmarks

  16. Where Cyber Agents Struggle: Bottleneck Analysis of Multi-Stage LLM Agents

    Sep 23, 2026Saeedeh Lohrasbi, Mohammad Mamun, Ahmed Yehia +2Agent Failure AnalysisLLM Agent Evaluation

  17. MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

    Sep 23, 2026Yongjun Jeong, Hanbum Ko, Ye Rin Kim +6Molecular OptimizationLLM Agent Evaluation

  18. REFLEX with Jev for Efficient Selective Control in LLM Agents

    Sep 22, 2026Tiantong Wu, Wei Yang Bryan LimLLM Inference EfficiencyLLM Agent Evaluation

  19. How Strongly Should Task State Influence an LLM Agent?

    Sep 22, 2026Chenyu Zhang, Wonbin Kweon, Jiawei HanLLM Agent EvaluationRuntime Enforcement for AI Agents

  20. Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

    Sep 22, 2026YanZe CaoLLM Agent EvaluationLLM Agent Reliability

  21. Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development

    Sep 21, 2026Haocheng Xia, Eugene Wu, Yongjoo ParkLLM Agent EvaluationAI Agent Benchmarks

  22. The AI Neuroscientist: An Interactive Agentic Interface for Neuroimaging Analysis

    Sep 21, 2026Aakash Patel, Panos Ketonis, Shreya Saxena +2AI Agents for Scientific DiscoveryLLM Agent Evaluation