LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
Organizations: Fudan University, China · City University of Hong Kong, Hong Kong · Chinese University of Hong Kong, Hong Kong · Singapore Management University, Singapore · Independent Researcher, Hong Kong · HKUST (GZ), China · Monash University Malaysia, Malaysia · Harbin Institute of Technology, China
Abstract
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.
Figures & tables
| Benchmark | #Tasks | Lang. | Type | Repo Size (k LoC) | Solution Size (LoC) | Modified (F/C/Fn) | Target Class Mentioned (%) |
| SWE-bench Verified | 500 | Python | Issue | 253.6 (277.6) | 38 (25) | 2.6/1.1/6.7 (2/0/5) | 15.9 (0.0) |
| SWE-bench Pro | 731 | Multiple | Issue | 270.0 (107.6) | 300 (173) | 7.2/2.4/14.5 (5/1/11) | 13.7 (0.0) |
| Multi-SWE-bench | 2,132 | Multiple | Issue | 205.8 (112.2) | 224 (64) | 6.6/1.8/11.6 (3/0/5) | 14.0 (0.0) |
| FEA-Bench | 1,401 | Python | Issue | 174.7 (89.4) | 222 (155) | 5.0/2.8/16.3 (4/2/12) | 15.5 (0.0) |
| DeepSWE | 113 | Multiple | Issue | 66.7 (29.8) | 1,670 (1,492) | 11.6/11.9/80.7 (8/6/72) | 31.8 (6.7) |
| FeatureBench | 200 | Python | Specification | 490.6 (460.8) | 2,050 (1,391) | 15.4/11.3/78.3 (10/9/54) | 38.7 (30.9) |
| Scaffold | Model | Resolved (%) | Completed (%) | F2P Pass (%) | P2P Pass (%) | Avg. Turns | Avg. Time (min) | Input Tokens (m) | Generated Tokens (k) | Avg. Cost ($) |
| Claude Code | DeepSeek V4 Flash | 0.0 | 4.8 | 15.0 | 71.8 | 176.9 | 64.4 | 27.6 | 115.6 | 1.4 |
| Opus 5 | 14.0 | 26.3 | 52.7 | 73.0 | 165.8 | 57.7 | 17.7 | 151.5 | 15.3 | |
| Kimi K3 | 4.0 | 10.1 | 32.5 | 76.1 | 187.9 | 110.7 | 36.3 | 149.9 | 16.9 | |
| GLM-5.2 | 2.0 | 4.5 | 20.6 | 78.3 | 172.4 | 55.4 | 21.8 | 104.3 | 7.7 | |
| MiniMax M3 | 3.0 | 4.0 | 12.2 | 73.5 | 343.4 | 30.2 | 44.6 | 66.0 | 4.2 | |
| Codex | DeepSeek V4 Flash | 0.0 | 3.3 | 18.1 | 72.1 | 197.5 | 56.2 | 22.1 | 125.3 | 0.6 |
| Scaffold | Model | Agent Failure | Solution Failure | Requirement Understanding | Task Planning | Code Localization | Code Editing | Code Verification | Self-Repair | Tool Use |
| Claude Code | DeepSeek V4 Flash | 0 | 100 | 11.2 | 10.5 | 27.8 | 25.6 | 13.8 | 11.0 | 0.1 |
| Opus 5 | 6 | 80 | 6.9 | 6.2 | 25.1 | 18.1 | 13.8 | 9.9 | 0.0 | |
| Kimi K3 | 5 | 91 | 5.1 | 4.4 | 26.4 | 25.9 | 9.8 | 18.8 | 0.6 | |
| GLM-5.2 | 0 | 98 | 8.2 | 7.7 | 25.9 | 31.7 | 12.9 | 11.6 | 0.0 | |
| MiniMax M3 | 1 | 96 | 9.4 | 9.2 | 24.2 | 29.3 | 12.2 | 11.5 | 0.2 | |
| Codex | DeepSeek V4 Flash | 0 | 100 | 4.5 | 3.9 | 39.9 | 24.6 | 18.3 | 8.5 | 0.3 |
| Agent | Resolved (%) | Completed (%) | F2P Pass (%) | P2P Pass (%) | Avg. Turns | Avg. Time (min) | Generated Tokens (k) | Avg. Cost ($) |
| Opus 5 + Claude Code | 34.0 (+20.0) | 42.9 (+16.6) | 72.7 (+20.0) | 81.1 (+8.1) | 168.0 (+2.2) | 50.8 (-6.9) | 141.5 (-10.0) | 15.2 (-0.1) |
| GPT-5.6 Sol + Codex | 25.0 (+22.0) | 32.8 (+19.2) | 56.2 (+27.6) | 85.7 (+9.7) | 123.8 (-11.4) | 27.4 (+2.5) | 47.8 (+1.4) | 10.8 (-0.9) |
| DeepSeek V4 Flash + OpenCode | 17.0 (+16.0) | 17.7 (+13.7) | 31.6 (+13.3) | 69.4 (-4.2) | 232.0 (+54.8) | 55.1 (+10.7) | 128.7 (+32.9) | 0.9 (+0.4) |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Project | Language | #Tasks | Repo Size (kLoC) | GitHub Stars | Maintenance Days |
| Compilers | CPython | Python | 18 | 1,344.4 | 77,201 | 3,507 |
| Cargo | Rust | 1 | 230.1 | 15,489 | 4,581 | |
| Mypy | Python | 1 | 213.2 | 20,644 | 5,033 | |
| OpenJDK | Java | 14 | 7,744.6 | 23,358 | 2,923 | |
| PHP | C | 4 | 2,615.7 | 40,393 | 5,573 | |
| Roslyn | C# | 1 | 5,981.0 | 20,672 | 4,268 |
| Dimension | Weight | Score anchors |
| Requirement coverage | 40% | 9–10: all concrete requirements are mapped; 7–8: one minor sub-requirement or edge case is missing or indirect; 5–6: an important requirement is only partly covered; 3–4: several central requirements are missing or mapped only to support; 1–2: little valid coverage. |
| Mapping precision | 25% | 9–10: no meaningful unrelated behavior; 7–8: only incidental support or shared helpers; 5–6: substantial unrelated behavior; 3–4: many loosely related scopes that belong elsewhere or should be associated; 1–2: mostly unrelated scopes. |
| Scope correctness | 15% | 9–10: correct section, direct/associated role, category, and unique scope ownership; 7–8: a few minor helper or category errors; 5–6: several incorrect assignments, but still usable; 3–4: assignment is mostly based on file proximity; 1–2: ownership is largely incorrect. |
| Granularity | 10% | 9–10: a fine-grained semantic boundary; 7–8: a few broad rows, but coherent ownership; 5–6: a large mixed set of scopes; 3–4: the section is mostly a catch-all; 1–2: no useful granularity. |
| Requirement specificity | 10% | 9–10: concrete APIs, syntax, semantics, configuration, or deliverables; 7–8: only routine details are implicit; 5–6: substantial inference is required; 3–4: very short or vague; 1–2: too vague to justify the mapping without external information. |
| Type | Weight | Section semantics or condition |
| Base | 1.0 | Core public API, syntax, protocol/schema, runtime semantics, or principal feature behavior. |
| Base | 0.8 | Important internal mechanism required by the main behavior. |
| Base | 0.6 | Compatibility, migration, configuration, rollout, error handling, or deprecation behavior. |
| Base | 0.4 | Explicit test, validation, benchmark, or performance requirement. |
| Base | 0.2 | Explicit documentation, build, generated-output, fixture, or support requirement. |
| Base | 0.0 | Non-implementable knowledge, context, or process section. |
| Mutation method | Count | Share (%) |
| Section Revert | 158 | 9.36 |
| Requirement Mismatch | 1,205 | 71.39 |
| Semantic Mutation | 325 | 19.25 |
| Total | 1,688 | 100.00 |
| Metric | Original | Overall |
| F2P tests | 1,480 | 2,234 |
| Added: mutant-killing | — | 622 |
| Added: coverage | — | 132 |
| Mutants killed | 1,031 | 1,620 |
| Mutation score (%) | 61.08 | 95.97 |
| Mean F2P coverage (%) | 60.07 | 78.50 |
| Failure Phase | Failure Mode | Description | Detection Method |
| Requirement Understanding | misread_requirement | Interprets the stated requirement incorrectly. | LLM-based |
| missed_constraint | Omits an explicit constraint from the reasoning or solution. | LLM-based | |
| incomplete_requirement_coverage | Addresses only part of the requested behavior. | LLM-based | |
| scope_misjudged | Chooses a boundary for the change that is too narrow or too broad. | LLM-based | |
| Task Planning | flawed_plan | Forms a plan that cannot fully satisfy the requirement. | LLM-based |
| abandoned_plan | Stops following the stated plan before completing it. | LLM-based |
| Scaffold | Model | Phase | Failure Modes | |||
| Requirement Understanding | Scope Misjudged | Missed Constraint | Incomplete Coverage | Misread Requirement | ||
| Claude Code | DeepSeek V4 Flash | 11.2 | 8.9 | 1.1 | 1.0 | 0.2 |
| Opus 5 | 6.9 | 3.7 | 1.8 | 1.4 | 0.0 | |
| Kimi K3 | 5.1 | 2.8 | 1.2 | 0.8 | 0.3 | |
| GLM-5.2 | 8.2 | 6.4 | 0.7 | 1.1 | 0.0 | |
| MiniMax M3 | 9.4 | 6.6 | 0.5 | 1.6 | 0.7 | |
| Scaffold | Model | Phase | Failure Modes | |
| Task Planning | Flawed Plan | Abandoned Plan | ||
| Claude Code | DeepSeek V4 Flash | 10.5 | 10.5 | 0.0 |
| Opus 5 | 6.2 | 6.2 | 0.0 | |
| Kimi K3 | 4.4 | 4.4 | 0.0 | |
| GLM-5.2 | 7.7 | 7.7 | 0.0 | |
| MiniMax M3 | 9.2 | 9.2 | 0.0 | |
| Scaffold | Model | Phase | Failure Modes | ||
| Code Localization | Cross-Module Context Missing | Missed Relevant File | Wrong File Localization | ||
| Claude Code | DeepSeek V4 Flash | 27.8 | 26.8 | 0.7 | 0.3 |
| Opus 5 | 25.1 | 24.3 | 0.8 | 0.0 | |
| Kimi K3 | 26.4 | 26.0 | 0.4 | 0.0 | |
| GLM-5.2 | 25.9 | 23.9 | 1.6 | 0.4 | |
| MiniMax M3 | 24.2 | 23.4 | 0.5 | 0.3 | |
| Scaffold | Model | Phase | Failure Modes | |||
| Code Editing | Incorrect Patch | Relevant Change Omitted | No Patch Produced | Gave Up Early | ||
| Claude Code | DeepSeek V4 Flash | 25.6 | 14.6 | 11.0 | 0.0 | 0.0 |
| Opus 5 | 18.1 | 11.0 | 7.1 | 0.0 | 0.0 | |
| Kimi K3 | 25.9 | 13.9 | 11.2 | 0.8 | 0.0 | |
| GLM-5.2 | 31.7 | 14.5 | 16.7 | 0.5 | 0.0 | |
| MiniMax M3 | 29.3 | 13.1 | 16.2 | 0.0 | 0.0 | |
| Scaffold | Model | Phase | Failure Modes | |||
| Code Verification | No Validation Attempted | No Test After Final Edit | Authored Test Never Run | Test Result Misinterpreted | ||
| Claude Code | DeepSeek V4 Flash | 13.8 | 10.0 | 2.3 | 0.7 | 0.8 |
| Opus 5 | 13.8 | 10.1 | 2.3 | 0.9 | 0.5 | |
| Kimi K3 | 9.8 | 4.9 | 3.3 | 0.7 | 0.9 | |
| GLM-5.2 | 12.9 | 10.9 | 1.1 | 0.3 | 0.6 | |
| MiniMax M3 | 12.2 | 9.9 | 0.6 | 1.3 | 0.4 | |
| Scaffold | Model | Phase | Failure Modes | |||
| Self- Repair | Repair Not Reverified | Repeated Ineffective Attempt | Reverted Own Change | Retry Without Change | ||
| Claude Code | DeepSeek V4 Flash | 11.0 | 5.7 | 3.9 | 1.4 | 0.0 |
| Opus 5 | 9.9 | 6.7 | 2.5 | 0.5 | 0.2 | |
| Kimi K3 | 18.8 | 9.1 | 7.3 | 2.4 | 0.0 | |
| GLM-5.2 | 11.6 | 5.2 | 4.2 | 2.2 | 0.0 | |
| MiniMax M3 | 11.5 | 1.8 | 5.3 | 4.3 | 0.1 | |
| Scaffold | Model | Phase | Failure Modes | ||
| Tool Use | Hallucinated Path Unrecovered | Tool Failure Not Retried | Repeated Patch Apply Failure | ||
| Claude Code | DeepSeek V4 Flash | 0.1 | 0.1 | 0.0 | 0.0 |
| Opus 5 | 0.0 | 0.0 | 0.0 | 0.0 | |
| Kimi K3 | 0.6 | 0.1 | 0.3 | 0.2 | |
| GLM-5.2 | 0.0 | 0.0 | 0.0 | 0.0 | |
| MiniMax M3 | 0.2 | 0.1 | 0.1 | 0.0 | |