Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar's exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.
Figures & tables
Model
Total
Active
Arch.
Qwen2.5-Coder-3B-Instruct
3B
3B
Dense
Qwen2.5-Coder-7B-Instruct
7B
7B
Dense
Qwen2.5-Coder-14B-Instruct
14B
14B
Dense
DeepSeek-Coder-V2-Lite-Instruct
16B
2.4B
MoE
TABLE I: MODELS EVALUATED
Model
Correct.
Maint.
Security
Effic.
Qwen2.5-Coder-3B
0.154
0.856
0.878
1.000
Qwen2.5-Coder-7B
0.308
0.897
0.923
0.480
Qwen2.5-Coder-14B
0.400
0.880
0.900
0.240
DeepSeek-Coder-V2-Lite
0.323
0.877
0.900
0.810
TABLE II: PER-MODEL SCORES ACROSS ALL FOUR QUALITY DIMENSIONS ( n=130 BUGS)
Fig. 1: Per-model scores across the four quality dimensions.
Model
Equal
Corr.-Priority
Other-Priority
Qwen2.5-Coder-3B
0.722
0.533
0.760
Qwen2.5-Coder-7B
0.652
0.537
0.675
Qwen2.5-Coder-14B
0.605
0.537
0.619
DeepSeek-Coder-V2-Lite
0.727
0.593
0.754
TABLE III: QUALITY INDEX UNDER DIFFERENT WEIGHTING SCHEMES
Fig. 2: Quality Index scores under the three weighting schemes.
Fig. 3: Generation speed versus total and active parameter count.
Comparison
p -value
Significant?
DeepSeek vs. Qwen2.5-Coder-14B
0.0525
No
DeepSeek vs. Qwen2.5-Coder-7B
0.8318
No
DeepSeek vs. Qwen2.5-Coder-3B
0.0001
Yes
Qwen2.5-Coder-14B vs. 7B
0.0118
Yes
Qwen2.5-Coder-14B vs. 3B
<0.0001
Yes
Qwen2.5-Coder-7B vs. 3B
<0.0001
Yes
TABLE IV: PAIRWISE MCNEMAR’S EXACT TEST FOR CORRECTNESS
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing 100876, China · University of Luxembourg