Long-Horizon Agent Evaluation

Momentum

47 papers in the last four weeks, up 176% on the four weeks before. 0.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 232

All topics
CardsList
  1. Agentic Coding Needs Proactivity, Not Just Autonomy

    May 7, 2026Nghi D. Q. Bui, Georgios EvangelopoulosAI Coding AgentsCoding Agents

  2. PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

    May 4, 2026Ruoqi Liu, Imran Q. Mohiuddin, Austin J. Schoeffler +10HealthcareLong-Horizon Agent Evaluation

  3. AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

    May 1, 2026Ranit Karmakar, Jayita ChatterjeeComputer-Use Agent BenchmarksLLM Agent Evaluation

  4. Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

    Apr 27, 2026Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried +1Computer-Use Agent BenchmarksWeb Agent Benchmarks

  5. From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

    Apr 21, 2026Md Nayem Uddin, Kumar Shubham, Eduardo Blanco +2Long-Horizon Agent EvaluationAI Agent Benchmarks

  6. SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

    Apr 19, 2026Ziao Zhang, Kou Shi, Shiting Huang +13Lifelong Learning AgentsLong-Horizon Agent Evaluation

  7. MemEvoBench: Benchmarking Safety Risks from Memory Misevolution in LLM Agents

    Apr 17, 2026Weiwei Xie, Shaoxiong Guo, Fan Zhang +5Long-Horizon Agent EvaluationAI Agent Security Benchmarks

  8. GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

    Apr 17, 2026Jize Wang, Xuanxuan Liu, Yining Li +7Tool-Use EvaluationLong-Horizon Agent Evaluation

  9. DR3^{3}-Eval: Towards Realistic and Reproducible Deep Research Evaluation

    Apr 16, 2026Qianqian Xie, Qingheng Xiong, He Zhu +16LLM Agent EvaluationLong-Horizon Agent Evaluation

  10. HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

    Apr 15, 2026Jiacheng Wang, Jinchang Hou, Fabian Wang +3AI Agent AuditingLong-Horizon Agent Evaluation

  11. EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents

    Jan 23, 2026Xinze Li, Ziyue Zhu, Siyuan Liu +4Episodic MemoryLong-Horizon Agent Evaluation

  12. LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning

    Jan 20, 2026Ye Tian, Zihao Wang, Onat Gungor +2HealthcareLong-Horizon Agent Evaluation

  13. SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

    Nov 20, 2025Juntao Cheng, Wanyue Zhang, Zhiwei Yu +7Long-Horizon Agent EvaluationEgocentric Video Understanding

  14. Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

    Oct 21, 2025Howard Yen, Yoonsang Lee, Ashwin Paranjape +5Agentic SearchLong-Horizon Agent Evaluation

  15. BuilderBench: The Building Blocks of Intelligent Agents

    Oct 7, 2025Raj Ghugare, Roger Creus Castanyer, Catherine Ji +4Long-Horizon Agent EvaluationRobot Manipulation Benchmarks

  16. SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

    Date pendingBingchen Zhao, Dhruv Srikanth, Yuxiang Wu +1Reward HackingAI Coding Agents

  17. Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

    Date pendingMinghao Guo, Meng Cao, Sui Zhao +10Deep Research AgentsLong-Horizon Agent Evaluation