cs.SEOct 6, 2026

Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair

Authors: Anvi Kalpesh Shah, Umamaheswara Sharma B

Organizations: National Institute of Technology, Calicut

Abstract

Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar's exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.

Figures & tables

Explore similar work

CardsList
  1. BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models

    May 9, 2026Yuanhao Li, Hongbo Wang, Xiaotang Shang +3Automated Program RepairRepair

  2. Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs

    Jul 14, 2026Zhili Huang, Ling Xu, Hongyu ZhangAutomated Program RepairRepair

  3. How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair

    Jul 28, 2026Ramtin Ehsani, Irene Manotas, Saurabh Pujar +2Automated Program RepairBug