Long-Horizon Agent Evaluation

Momentum

48 papers in the last four weeks, up 182% on the four weeks before. 0.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 235

All topics
CardsList
  1. AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents

    May 8, 2026Zhengkang Guo, Yiyang Li, Lin Qiu +7LLM Agent EvaluationLong-Horizon Agent Evaluation

  2. Stop Comparing LLM Agents Without Disclosing the Harness

    May 7, 2026Yunbei Zhang, Janet Wang, Yingqiang Ge +3LLM Agent EvaluationLong-Horizon Agent Evaluation

  3. Agentic Coding Needs Proactivity, Not Just Autonomy

    May 7, 2026Nghi D. Q. Bui, Georgios EvangelopoulosAI Coding AgentsCoding Agents

  4. PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

    May 4, 2026Ruoqi Liu, Imran Q. Mohiuddin, Austin J. Schoeffler +10HealthcareLong-Horizon Agent Evaluation

  5. AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

    May 1, 2026Ranit Karmakar, Jayita ChatterjeeComputer-Use Agent BenchmarksLLM Agent Evaluation

  6. Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

    Apr 27, 2026Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried +1Computer-Use Agent BenchmarksWeb Agent Benchmarks

  7. From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

    Apr 21, 2026Md Nayem Uddin, Kumar Shubham, Eduardo Blanco +2Long-Horizon Agent EvaluationAI Agent Benchmarks

  8. SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

    Apr 19, 2026Ziao Zhang, Kou Shi, Shiting Huang +13Lifelong Learning AgentsLong-Horizon Agent Evaluation

  9. MemEvoBench: Benchmarking Safety Risks from Memory Misevolution in LLM Agents

    Apr 17, 2026Weiwei Xie, Shaoxiong Guo, Fan Zhang +5Long-Horizon Agent EvaluationAI Agent Security Benchmarks

  10. GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

    Apr 17, 2026Jize Wang, Xuanxuan Liu, Yining Li +7Tool-Use EvaluationLong-Horizon Agent Evaluation

  11. DR3^{3}-Eval: Towards Realistic and Reproducible Deep Research Evaluation

    Apr 16, 2026Qianqian Xie, Qingheng Xiong, He Zhu +16LLM Agent EvaluationLong-Horizon Agent Evaluation

  12. HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

    Apr 15, 2026Jiacheng Wang, Jinchang Hou, Fabian Wang +3AI Agent AuditingLong-Horizon Agent Evaluation

  13. EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents

    Jan 23, 2026Xinze Li, Ziyue Zhu, Siyuan Liu +4Episodic MemoryLong-Horizon Agent Evaluation

  14. LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning

    Jan 20, 2026Ye Tian, Zihao Wang, Onat Gungor +2HealthcareLong-Horizon Agent Evaluation

  15. SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

    Nov 20, 2025Juntao Cheng, Wanyue Zhang, Zhiwei Yu +7Long-Horizon Agent EvaluationEgocentric Video Understanding

  16. Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

    Oct 21, 2025Howard Yen, Yoonsang Lee, Ashwin Paranjape +5Agentic SearchLong-Horizon Agent Evaluation

  17. BuilderBench: The Building Blocks of Intelligent Agents

    Oct 7, 2025Raj Ghugare, Roger Creus Castanyer, Catherine Ji +4Long-Horizon Agent EvaluationRobot Manipulation Benchmarks

  18. SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

    Date pendingBingchen Zhao, Dhruv Srikanth, Yuxiang Wu +1Reward HackingAI Coding Agents

  19. Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

    Date pendingMinghao Guo, Meng Cao, Sui Zhao +10Deep Research AgentsLong-Horizon Agent Evaluation