cs.SESep 29, 2026

LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

Authors: Yun Peng, Zihan Wu, Zeyang Zhuang, Xin Zhou, Rui Shu, Xu Han, Chun Yong Chong, Yuan Wang, +1 more

Organizations: Fudan University, China · City University of Hong Kong, Hong Kong · Chinese University of Hong Kong, Hong Kong · Singapore Management University, Singapore · Independent Researcher, Hong Kong · HKUST (GZ), China · Monash University Malaysia, Malaysia · Harbin Institute of Technology, China

Abstract

Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17×\times), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades

    May 15, 2026Xinbo Xu, Ruihan Yang, Haiyang Shen +13Coding AgentsAgentic Benchmarks

  2. LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation

    Jul 31, 2026Han Li, Zhemin Fang, Rili Feng +8Coding AgentsAgent Loop

  3. A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

    Aug 10, 2026Xin Zhou, Chun Yong Chong, Kisub Kim +11Coding AgentsCode Generation