Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar's exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.
Figures & tables
Model
Total
Active
Arch.
Qwen2.5-Coder-3B-Instruct
3B
3B
Dense
Qwen2.5-Coder-7B-Instruct
7B
7B
Dense
Qwen2.5-Coder-14B-Instruct
14B
14B
Dense
DeepSeek-Coder-V2-Lite-Instruct
16B
2.4B
MoE
TABLE I: MODELS EVALUATED
Model
Correct.
Maint.
Security
Effic.
Qwen2.5-Coder-3B
0.154
0.856
0.878
1.000
Qwen2.5-Coder-7B
0.308
0.897
0.923
0.480
Qwen2.5-Coder-14B
0.400
0.880
0.900
0.240
DeepSeek-Coder-V2-Lite
0.323
0.877
0.900
0.810
TABLE II: PER-MODEL SCORES ACROSS ALL FOUR QUALITY DIMENSIONS ( n=130 BUGS)
Fig. 1: Per-model scores across the four quality dimensions.
Model
Equal
Corr.-Priority
Other-Priority
Qwen2.5-Coder-3B
0.722
0.533
0.760
Qwen2.5-Coder-7B
0.652
0.537
0.675
Qwen2.5-Coder-14B
0.605
0.537
0.619
DeepSeek-Coder-V2-Lite
0.727
0.593
0.754
TABLE III: QUALITY INDEX UNDER DIFFERENT WEIGHTING SCHEMES
Fig. 2: Quality Index scores under the three weighting schemes.
Fig. 3: Generation speed versus total and active parameter count.
Comparison
p -value
Significant?
DeepSeek vs. Qwen2.5-Coder-14B
0.0525
No
DeepSeek vs. Qwen2.5-Coder-7B
0.8318
No
DeepSeek vs. Qwen2.5-Coder-3B
0.0001
Yes
Qwen2.5-Coder-14B vs. 7B
0.0118
Yes
Qwen2.5-Coder-14B vs. 3B
<0.0001
Yes
Qwen2.5-Coder-7B vs. 3B
<0.0001
Yes
TABLE IV: PAIRWISE MCNEMAR’S EXACT TEST FOR CORRECTNESS
Reinforcement learning for program repair is hindered by sparse execution feedback and coarse sequence-level rewards that obscure which edits actually fix bugs. We present BoostAPR, a three-stage framework addressing these challenges: (1) supervised fine-tuning on execution-verified demonstrations with reasoning traces, (2) training dual reward models--a sequence-level assessor and a line-level credit allocator--from execution outcomes, and (3) PPO optimization where the line-level model redistributes rewards to critical edit regions. This line-level credit assignment operates at an intermediate granularity naturally suited to code changes. Trained on SWE-Gym and evaluated on four benchmarks, BoostAPR achieves 40.7% on SWE-bench Verified (+22.9pp over base model), 24.8% on Defects4J (Python-to-Java transfer), 84.5% on HumanEval-Java, and 95.0% on QuixBugs, achieving competitive results among open-source models with strong cross-language generalization.
Yuanhao Li, Hongbo Wang, Xiaotang Shang +3
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing 100876, China · University of Luxembourg
Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large and repetitive to serve as effective model context. Second, repeated patch sampling may produce different implementations without yielding distinct root-cause hypotheses or repair strategies. We present CT-Repair, an agentic APR framework representing static and dynamic evidence as queryable Code Property Graph (CPG) and Temporal Execution Graph (TEG). CT-Repair applies a three-stage filtering pipeline to construct compact TEGs. Three finite-state-machine-guided agents analyze each bug from static, dynamic, and hybrid perspectives and independently produce evidence-grounded repair strategies. A strategy-guided generation procedure instantiates these strategies as candidate patches and uses validation feedback to refine the most promising strategy. We evaluate CT-Repair on 854 Java bugs from Defects4J v3.0. In the mixed-model configuration, CT-Repair correctly repairs 489 bugs. Under a controlled GPT-5.4-mini configuration, it repairs 388 bugs, 19 and 30 more than ReinFix and RepairAgent, respectively. The union of the three evidence perspectives repairs 99 more bugs than the strongest individual perspective. The filtering pipeline also compacts runtime evidence, with execution filtering narrowing the candidate method scope by 94.85% on average and behavior filtering further reducing retained runtime records by 55.97%. These results show that structured runtime evidence and multi-perspective reasoning can improve repair effectiveness without relying solely on a larger patch-generation budget.
Large Language Model (LLM)-based Automated Program Repair systems are advancing rapidly, yet their performance remains inconsistent. Even when provided with the same contextual information, an LLM may generate a correct patch for one bug but fail on another closely related bug. Why this happens remains poorly understood, and it is unclear how LLMs prioritize the diverse information in bug reports and whether model attention affects repair success. In this paper, we present the first empirical study of attention patterns in LLM-based program repair, providing interpretable insights into how models process bug reports and where their attention is concentrated during repair. We analyze 319 real-world Python and Java bugs from SWE-bench Verified and Multi-SWE-bench to study (RQ1) how model attention is distributed across bug report sections, (RQ2) how attention patterns within each section differ between successful and unsuccessful repairs, and (RQ3) how these patterns compare to information developers consider important for bug fixing. We find that successful repairs are characterized by diffused attention across multiple diagnostic components such as bug descriptions, stacktraces, and test cases, while failures often exhibit over-localized attention toward metadata such as version information. We further observe that stronger alignment between model attention and developer-identified key sections and phrases is associated with higher repair success. Our results provide the first empirical evidence that attention misallocation is a key factor in LLM-based APR failures, and offer actionable insights for designing more interpretable and reliable future APR systems.
Ramtin Ehsani, Irene Manotas, Saurabh Pujar +2
Drexel University Philadelphia, PA, USA · IBM Research Yorktown, NY, USA