AI Agent Benchmarks

Momentum

148 papers in the last four weeks, up 185% on the four weeks before. 1.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 1,013

All topics
CardsList
  1. Look Before You Leap: Pre-Action Verification for LLM Agents

    Sep 14, 2026Asaad AlthoubiLLM Agent VerificationAI Agent Benchmarks

  2. Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

    Sep 14, 2026Hazel Mak, Susheel Suresh, Sahil Bhatnagar +3Terminal AgentsLLM Agent Evaluation

  3. ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents

    Sep 14, 2026Bowen Guan, Zhentao Yin, Yanming ShenLLM Agent EvaluationAI Agent Benchmarks

  4. Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents

    Sep 14, 2026Zhutao Lv, Chenhao Dang, Yi Feng +5Remote SensingAI Agent Benchmarks

  5. MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

    Sep 14, 2026Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin +7Spoken Dialogue SystemsAI Agent Benchmarks

  6. Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

    Sep 14, 2026Baoyang Jiang, Fengchun Zhang, Leyuan Wang +9Benchmark ConstructionEmbodied QA

  7. MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

    Sep 14, 2026Bosi Wen, Cunxiang Wang, Jiayi Gui +6Software Engineering AgentsAI Coding Agents

  8. Safety Signals to Verify NetOps Agents with Action-Level Granularity

    Sep 13, 2026Tobias Labarta, Frederik Pahde, Novak Boškov +5Long-Horizon Agent EvaluationAI Agent Safety

  9. BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

    Sep 12, 2026Shenghan Zheng, Zonglin Di, Yimin Liu +19Reward HackingLLM Agent Evaluation

  10. MindTopo: Can Foundation Models Reason in Topological Space?

    Sep 12, 2026Yunfei Ge, Anbang Liu, Qineng Wang +9Spatial Reasoning BenchmarksAI Agent Benchmarks

  11. K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

    Sep 11, 2026Guangsheng Yu, Yanna Jiang, Qin Wang +2LLM Agent EvaluationPrivacy Leakage in Language Models

  12. ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation

    Sep 11, 2026Zesheng Wei, Mengfan Li, Wenhao Liu +3Automated NegotiationLLM Agent Evaluation

  13. Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

    Sep 11, 2026Hanhua Hong, Yizhi Li, Luu Gia Huy +3AI Agent EvaluationAgent Evaluation

  14. Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

    Sep 9, 2026Tianzhu Zhang, Weichen Tao, Changgang Zheng +4AI Agent Benchmarks

  15. Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method

    Sep 8, 2026Boao Yu, Zimo Chen, Junreng Rao +4Aerial RoboticsAI Agent Benchmarks

  16. Qiushi Engine on AstaBench E2E-Bench-Hard

    Sep 8, 2026Wenhao Li, Shuxing Yang, Fujia Chen +13Agent EvaluationAI Agent Benchmarks

  17. CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information

    Sep 8, 2026Dingying Liu, Yunshun Zhong, Wentao Zhang +1Agent Failure AnalysisAI Agent Reliability

  18. AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

    Sep 7, 2026Yunxiang Mo, Tianshi Zheng, Yisen Gao +7AI for ScienceAI Agent Evaluation

  19. Search-to-World: Evaluation of 3D World Delivery from User Request through Web Search

    Sep 7, 2026Zixiao Gu, Yabo Chen, Xunzhi Xiang +5Agentic Search3D Reconstruction

  20. NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management

    Sep 7, 2026Yulin Wei, Xiangchen Wang, Jianhui Pan +5Embodied QAAI Agent Benchmarks

  21. ττ\tau^\tau-Bench: An Environment for End-To-End, Realistic Agent Construction

    Sep 7, 2026Quan Shi, Keshav Dhandhania, Karthik Narasimhan +1AI Coding AgentsLLM Agent Evaluation

  22. ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

    Sep 7, 2026Xinran Zhang, Pengrui Lu, Lyumanshan Ye +1AI-Assisted Decision MakingLLM Agent Evaluation

  23. First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves

    Sep 4, 2026Tianjie Ju, Xinyue Xu, Wanxuan Sun +4RL for Language Model ReasoningAI Agent Benchmarks