LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
Authors: Yun Peng, Zihan Wu, Zeyang Zhuang, Xin Zhou, Rui Shu, Xu Han, Chun Yong Chong, Yuan Wang, +1 more
Organizations: Fudan University, China · City University of Hong Kong, Hong Kong · Chinese University of Hong Kong, Hong Kong · Singapore Management University, Singapore · Independent Researcher, Hong Kong · HKUST (GZ), China · Monash University Malaysia, Malaysia · Harbin Institute of Technology, China
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17×), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.
Figures & tables
Figure 1: Illustration of tasks with different levels of perception and implementation complexity. Perception complexity captures the difficulty of grounding user intent and high-level design in an existing software system, whereas implementation complexity captures the difficulty of producing correct and regression-free code changes from that understanding.
Figure 2: Construction of LoLBench . We collect enhancement proposals and their corresponding implementations, align proposal sections with implementation changes, construct executable tasks, and augment behavioral tests with incorrect solution mutants and coverage feedback.
Benchmark
#Tasks
Lang.
Type
Repo Size (k LoC)
Solution Size (LoC)
Modified (F/C/Fn)
Target Class Mentioned (%)
SWE-bench Verified
500
Python
Issue
253.6 (277.6)
38 (25)
2.6/1.1/6.7 (2/0/5)
15.9 (0.0)
SWE-bench Pro
731
Multiple
Issue
270.0 (107.6)
300 (173)
7.2/2.4/14.5 (5/1/11)
13.7 (0.0)
Multi-SWE-bench
2,132
Multiple
Issue
205.8 (112.2)
224 (64)
6.6/1.8/11.6 (3/0/5)
14.0 (0.0)
FEA-Bench
1,401
Python
Issue
174.7 (89.4)
222 (155)
5.0/2.8/16.3 (4/2/12)
15.5 (0.0)
DeepSWE
113
Multiple
Issue
66.7 (29.8)
1,670 (1,492)
11.6/11.9/80.7 (8/6/72)
31.8 (6.7)
FeatureBench
200
Python
Specification
490.6 (460.8)
2,050 (1,391)
15.4/11.3/78.3 (10/9/54)
38.7 (30.9)
Table 1: Comparison with recent coding benchmarks. The last four columns report mean (median) values over tasks. “Repo Size” counts source lines excluding comments and blanks. “Solution Size” counts inserted and deleted lines across the implementation PR. “Modified” counts inserted, deleted, and changed files/classes/functions in solutions. “Class Mentioned” is the percentage of target class names that appear in the task description. Lower values indicate less explicit implementation guidance.
Scaffold
Model
Resolved (%)
Completed (%)
F2P Pass (%)
P2P Pass (%)
Avg. Turns
Avg. Time (min)
Input Tokens (m)
Generated Tokens (k)
Avg. Cost ($)
Claude Code
DeepSeek V4 Flash
0.0
4.8
15.0
71.8
176.9
64.4
27.6
115.6
1.4
Opus 5
14.0
26.3
52.7
73.0
165.8
57.7
17.7
151.5
15.3
Kimi K3
4.0
10.1
32.5
76.1
187.9
110.7
36.3
149.9
16.9
GLM-5.2
2.0
4.5
20.6
78.3
172.4
55.4
21.8
104.3
7.7
MiniMax M3
3.0
4.0
12.2
73.5
343.4
30.2
44.6
66.0
4.2
Codex
DeepSeek V4 Flash
0.0
3.3
18.1
72.1
197.5
56.2
22.1
125.3
0.6
Table 2: Main results of 28 agents on LoLBench across 100 tasks. “Resolved”, “Completed”, “F2P Pass”, and “P2P Pass” report percentages. We average “turns”, “time”, “input tokens”, “generated tokens”, and “cost” over tasks. “Time” reports average wall-clock minutes per task. “Input tokens” includes cached input and is reported in millions. “Generated tokens” sums output and reasoning tokens. All models use their second-highest reasoning effort. Bold and underline denote the best and second-best performance values within each scaffold, respectively. Claude Code with GPT-5.6 Sol and Codex with Opus 5 are unavailable due to provider issues.
Figure 3: Average agent turns per task across five phases. Solid segments denote initial solution generation, and hatched segments denote repair loops after failed verification.
Scaffold
Model
Agent Failure
Solution Failure
Requirement Understanding
Task Planning
Code Localization
Code Editing
Code Verification
Self-Repair
Tool Use
Claude Code
DeepSeek V4 Flash
0
100
11.2
10.5
27.8
25.6
13.8
11.0
0.1
Opus 5
6
80
6.9
6.2
25.1
18.1
13.8
9.9
0.0
Kimi K3
5
91
5.1
4.4
26.4
25.9
9.8
18.8
0.6
GLM-5.2
0
98
8.2
7.7
25.9
31.7
12.9
11.6
0.0
MiniMax M3
1
96
9.4
9.2
24.2
29.3
12.2
11.5
0.2
Codex
DeepSeek V4 Flash
0
100
4.5
3.9
39.9
24.6
18.3
8.5
0.3
Table 3: The heat map of failure attribution across the 28 agents on LoLBench . “Agent Failure” indicates failed tasks without analyzable trajectories, and “Solution Failure” indicates failed tasks with analyzable trajectories. The 49 agent failures for OpenCode with GPT-5.6 Sol are primarily due to timeouts before submission. The last seven columns show failure attribution across seven phases, and they sum to the “Solution Failure” column. Darker backgrounds indicate a larger share of total failure attribution for each agent’s solution failures.
Figure 4: Macro-averaged precision and recall for code localization based on inspected (left) and modified (right) files and functions.
Agent
Resolved (%)
Completed (%)
F2P Pass (%)
P2P Pass (%)
Avg. Turns
Avg. Time (min)
Generated Tokens (k)
Avg. Cost ($)
Opus 5 + Claude Code
34.0 (+20.0)
42.9 (+16.6)
72.7 (+20.0)
81.1 (+8.1)
168.0 (+2.2)
50.8 (-6.9)
141.5 (-10.0)
15.2 (-0.1)
GPT-5.6 Sol + Codex
25.0 (+22.0)
32.8 (+19.2)
56.2 (+27.6)
85.7 (+9.7)
123.8 (-11.4)
27.4 (+2.5)
47.8 (+1.4)
10.8 (-0.9)
DeepSeek V4 Flash + OpenCode
17.0 (+16.0)
17.7 (+13.7)
31.6 (+13.3)
69.4 (-4.2)
232.0 (+54.8)
55.1 (+10.7)
128.7 (+32.9)
0.9 (+0.4)
Table 4: Results of three agents on LoLBench with file trees and API specification in EPs. All columns are calculated the same way in Table 2 . Parenthesized values report signed absolute changes computed by subtracting the baseline values in Table 2 from the values here.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Project
Language
#Tasks
Repo Size (kLoC)
GitHub Stars
Maintenance Days
Compilers
CPython
Python
18
1,344.4
77,201
3,507
Cargo
Rust
1
230.1
15,489
4,581
Mypy
Python
1
213.2
20,644
5,033
OpenJDK
Java
14
7,744.6
23,358
2,923
PHP
C
4
2,615.7
40,393
5,573
Roslyn
C#
1
5,981.0
20,672
4,268
Appendix
Table 5: Statistics of projects in LoLBench . “Language” is assigned at the project level and applies to every task from that project. “Repo Size” is the mean source code lines excluding comments and blank lines. “GitHub Stars” are as of September 18, 2026, and “Maintenance Days” is the number of days elapsed from each repository’s creation on GitHub to September 18, 2026.
Dimension
Weight
Score anchors
Requirement coverage
40%
9–10: all concrete requirements are mapped; 7–8: one minor sub-requirement or edge case is missing or indirect; 5–6: an important requirement is only partly covered; 3–4: several central requirements are missing or mapped only to support; 1–2: little valid coverage.
Mapping precision
25%
9–10: no meaningful unrelated behavior; 7–8: only incidental support or shared helpers; 5–6: substantial unrelated behavior; 3–4: many loosely related scopes that belong elsewhere or should be associated; 1–2: mostly unrelated scopes.
Scope correctness
15%
9–10: correct section, direct/associated role, category, and unique scope ownership; 7–8: a few minor helper or category errors; 5–6: several incorrect assignments, but still usable; 3–4: assignment is mostly based on file proximity; 1–2: ownership is largely incorrect.
Granularity
10%
9–10: a fine-grained semantic boundary; 7–8: a few broad rows, but coherent ownership; 5–6: a large mixed set of scopes; 3–4: the section is mostly a catch-all; 1–2: no useful granularity.
Requirement specificity
10%
9–10: concrete APIs, syntax, semantics, configuration, or deliverables; 7–8: only routine details are implicit; 5–6: substantial inference is required; 3–4: very short or vague; 1–2: too vague to justify the mapping without external information.
Appendix
Table 6: Dimensions and score anchors used to evaluate EP-to-PR mappings. Each dimension is scored from 1 to 10, and the final section score is their weighted sum.
Type
Weight
Section semantics or condition
Base
1.0
Core public API, syntax, protocol/schema, runtime semantics, or principal feature behavior.
Base
0.8
Important internal mechanism required by the main behavior.
Base
0.6
Compatibility, migration, configuration, rollout, error handling, or deprecation behavior.
Base
0.4
Explicit test, validation, benchmark, or performance requirement.
Base
0.2
Explicit documentation, build, generated-output, fixture, or support requirement.
Base
0.0
Non-implementable knowledge, context, or process section.
Appendix
Table 7: Semantic importance weights for EP sections. Modifiers are applied independently after choosing the base weight. Implementable section weights are clipped to [0.1,1.0] ; non-implementable sections always receive zero.
Mutation method
Count
Share (%)
Section Revert
158
9.36
Requirement Mismatch
1,205
71.39
Semantic Mutation
325
19.25
Total
1,688
100.00
Appendix
Table 8: Distribution of retained solution mutants, excluding retired or specification-equivalent candidates.
Metric
Original
Overall
F2P tests
1,480
2,234
Added: mutant-killing
—
622
Added: coverage
—
132
Mutants killed
1,031
1,620
Mutation score (%)
61.08
95.97
Mean F2P coverage (%)
60.07
78.50
Appendix
Table 9: F2P test augmentation. “Overall” combines original and added tests. Coverage is the mean per-task coverage of F2P tests. Mutation scores use the 1,688 retained mutants in Table 8 .
Figure 5: Resolved% by programming language.
Figure 6: Completed% by programming language.
Figure 7: F2P Pass% by programming language.
Figure 8: P2P Pass% by programming language.
Figure 9: Phase-attributed token usage per attempt across 28 agents. Non-cached input excludes cache reads but includes cache writes. Bar labels sum the five colored phases and omit usage without phase attribution. Colors denote phases, solid segments denote initial solution generation, and hatched segments denote repair loops after verification failures.
Failure Phase
Failure Mode
Description
Detection Method
Requirement Understanding
misread_requirement
Interprets the stated requirement incorrectly.
LLM-based
missed_constraint
Omits an explicit constraint from the reasoning or solution.
LLM-based
incomplete_requirement_coverage
Addresses only part of the requested behavior.
LLM-based
scope_misjudged
Chooses a boundary for the change that is too narrow or too broad.
LLM-based
Task Planning
flawed_plan
Forms a plan that cannot fully satisfy the requirement.
LLM-based
abandoned_plan
Stops following the stated plan before completing it.
LLM-based
Appendix
Table 10: Failure phases, failure modes, and detection methods for the 24 modes observed across the 28 evaluated agents on LoLBench .
Scaffold
Model
Phase
Failure Modes
Requirement Understanding
Scope Misjudged
Missed Constraint
Incomplete Coverage
Misread Requirement
Claude Code
DeepSeek V4 Flash
11.2
8.9
1.1
1.0
0.2
Opus 5
6.9
3.7
1.8
1.4
0.0
Kimi K3
5.1
2.8
1.2
0.8
0.3
GLM-5.2
8.2
6.4
0.7
1.1
0.0
MiniMax M3
9.4
6.6
0.5
1.6
0.7
Appendix
Table 11: Failure mode distribution for Requirement Understanding failures across the 28 agents. The phase column reproduces the corresponding failure attribution in Table 3 and equals the sum of failure mode values in every row. Darker backgrounds indicate a larger share of total failure attribution for each agent’s solution failures.
Scaffold
Model
Phase
Failure Modes
Task Planning
Flawed Plan
Abandoned Plan
Claude Code
DeepSeek V4 Flash
10.5
10.5
0.0
Opus 5
6.2
6.2
0.0
Kimi K3
4.4
4.4
0.0
GLM-5.2
7.7
7.7
0.0
MiniMax M3
9.2
9.2
0.0
Appendix
Table 12: Failure mode distribution for Task Planning failures across the 28 agents. The phase column reproduces the corresponding failure attribution in Table 3 and equals the sum of failure mode values in every row. Darker backgrounds indicate a larger share of total failure attribution for each agent’s solution failures.
Scaffold
Model
Phase
Failure Modes
Code Localization
Cross-Module Context Missing
Missed Relevant File
Wrong File Localization
Claude Code
DeepSeek V4 Flash
27.8
26.8
0.7
0.3
Opus 5
25.1
24.3
0.8
0.0
Kimi K3
26.4
26.0
0.4
0.0
GLM-5.2
25.9
23.9
1.6
0.4
MiniMax M3
24.2
23.4
0.5
0.3
Appendix
Table 13: Failure mode distribution for Code Localization failures across the 28 agents. The phase column reproduces the corresponding failure attribution in Table 3 and equals the sum of failure mode values in every row. Darker backgrounds indicate a larger share of total failure attribution for each agent’s solution failures.
Scaffold
Model
Phase
Failure Modes
Code Editing
Incorrect Patch
Relevant Change Omitted
No Patch Produced
Gave Up Early
Claude Code
DeepSeek V4 Flash
25.6
14.6
11.0
0.0
0.0
Opus 5
18.1
11.0
7.1
0.0
0.0
Kimi K3
25.9
13.9
11.2
0.8
0.0
GLM-5.2
31.7
14.5
16.7
0.5
0.0
MiniMax M3
29.3
13.1
16.2
0.0
0.0
Appendix
Table 14: Failure mode distribution for Code Editing failures across the 28 agents. The phase column reproduces the corresponding failure attribution in Table 3 and equals the sum of failure mode values in every row. Darker backgrounds indicate a larger share of total failure attribution for each agent’s solution failures.
Scaffold
Model
Phase
Failure Modes
Code Verification
No Validation Attempted
No Test After Final Edit
Authored Test Never Run
Test Result Misinterpreted
Claude Code
DeepSeek V4 Flash
13.8
10.0
2.3
0.7
0.8
Opus 5
13.8
10.1
2.3
0.9
0.5
Kimi K3
9.8
4.9
3.3
0.7
0.9
GLM-5.2
12.9
10.9
1.1
0.3
0.6
MiniMax M3
12.2
9.9
0.6
1.3
0.4
Appendix
Table 15: Failure mode distribution for Code Verification failures across the 28 agents. The phase column reproduces the corresponding failure attribution in Table 3 and equals the sum of failure mode values in every row. Darker backgrounds indicate a larger share of total failure attribution for each agent’s solution failures.
Scaffold
Model
Phase
Failure Modes
Self- Repair
Repair Not Reverified
Repeated Ineffective Attempt
Reverted Own Change
Retry Without Change
Claude Code
DeepSeek V4 Flash
11.0
5.7
3.9
1.4
0.0
Opus 5
9.9
6.7
2.5
0.5
0.2
Kimi K3
18.8
9.1
7.3
2.4
0.0
GLM-5.2
11.6
5.2
4.2
2.2
0.0
MiniMax M3
11.5
1.8
5.3
4.3
0.1
Appendix
Table 16: Failure mode distribution for Self-Repair failures across the 28 agents. The phase column reproduces the corresponding failure attribution in Table 3 and equals the sum of failure mode values in every row. Darker backgrounds indicate a larger share of total failure attribution for each agent’s solution failures.
Scaffold
Model
Phase
Failure Modes
Tool Use
Hallucinated Path Unrecovered
Tool Failure Not Retried
Repeated Patch Apply Failure
Claude Code
DeepSeek V4 Flash
0.1
0.1
0.0
0.0
Opus 5
0.0
0.0
0.0
0.0
Kimi K3
0.6
0.1
0.3
0.2
GLM-5.2
0.0
0.0
0.0
0.0
MiniMax M3
0.2
0.1
0.1
0.0
Appendix
Table 17: Failure mode distribution for Tool Use failures across the 28 agents. The phase column reproduces the corresponding failure attribution in Table 3 and equals the sum of failure mode values in every row. Darker backgrounds indicate a larger share of total failure attribution for each agent’s solution failures.
Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many files. However, most existing benchmarks focus predominantly on single-issue bug fixes from Python repositories, with coarse pass/fail evaluation outcomes, and thus fail to capture long-horizon, multi-target development at real engineering scale. To address this gap, we present RoadmapBench, a benchmark of 115 long-horizon coding tasks grounded in real open-source version upgrades across 17 repositories and 5 programming languages. Each task places the agent on a source-version code snapshot and provides a multi-target roadmap instruction requiring it to implement the functionality introduced in the target version, with a median modification of 3,700 lines across 51 files. We conduct a systematic evaluation on thirteen frontier models and find that even the strongest, Claude-Opus-4.7, resolves only 39.1% of tasks, while the weakest achieves merely 5.2%, in stark contrast to existing bug-fix benchmarks, suggesting that long-horizon software development remains a largely unsolved problem.
Xinbo Xu, Ruihan Yang, Haiyang Shen +13
UniPat AI · Peking University · Fudan University +4
Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a long-horizon benchmark for loop engineering in coding agent evaluation. Each task is a dependency DAG over separately testable development units with source-evidenced prerequisite edges. LOOPSBENCH comprises 112 tasks from authentic sources spanning 8 programming languages and 9 domains. Its flow-aware runtime releases tests along the ready frontier and retains completed nodes as regression obligations. We evaluate frontier coding agents paired with widely used loop implementations. The strongest configuration, Opus-4.7 with Claude Code and outer continuation, resolves 25.00% of tasks. Recorded plans recover only part of the source-recovered prerequisite DAG, and regression events remain visible across the evaluated loop profiles. We open source the benchmark data and code, including all tasks, more than 5,300 development units, and executable tests, at microsoft/Loopsbench.
Han Li, Zhemin Fang, Rili Feng +8
Microsoft · Nanjing University · Shanghai Jiao Tong University +1
Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs. Yet existing repository-level benchmarks typically evaluate only whether the final patch passes tests. Satisfying a user request requires a long chain of interdependent reasoning and decisions: an agent must recover explicit and implicit requirements, formulate a repository-grounded implementation plan, and translate it into correct code. A pass/fail outcome cannot characterize how an unsuccessful trajectory diverges from the requirements and implementation process needed for a correct patch. To address this gap, we introduce SWE-RPG, a repository-level benchmark that combines executable patch evaluation with validated ground-truth references (GTs) for (1) Requirement Clarification and (2) Implementation Planning. These intermediate GTs support retrospective, GT-aligned diagnosis of complete coding-agent trajectories across clarification, planning, code generation, and artifact submission. SWE-RPG comprises 163 tasks from 31 Python and Java repositories, including 113 bug fixes and 50 feature additions. We evaluate 3 coding agents, including Claude Code, Codex, and OpenCode, with 6 large language model backends, including Claude-Sonnet-5 and GPT-5.6-Terra. Results show that the evaluated popular coding agents still struggle to implement user requests in existing repositories, achieving an average resolved rate of only 31.5% on SWE-RPG. Intermediate-GT diagnosis further identifies implicit requirement recovery as the main bottleneck, accounting for 24.5%--46.0% of agent runs. This result suggests implicit-requirement recovery as a key candidate direction for improving coding agents. The benchmark data and evaluation code are available at https://github.com/Xin-Zhou-smu/SWE-RPG-Bench.
Xin Zhou, Chun Yong Chong, Kisub Kim +11
Singapore Management University, Singapore · Monash University Malaysia, Malaysia · Daegu Gyeongbuk Institute of Science and Technology (DGIST), Republic of Korea +5