cs.CLSep 24, 2026

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

Authors: Alexandru Stefan Stoica, Traian Rebedea, Marian Cristian Mihaescu

Organizations: University Politehnica of Bucharest Bucharest, Romania · University of Craiova Craiova, Romania

Abstract

Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between these two capabilities. This work focuses on determining how much the LLM deviates from a buggy solution to fix the bug compared to a human-written patch, and if there is a bias towards generating entirely new solutions. We construct a dataset with all the submissions (∼\sim 3000) from a couple of users from Codeforces, and we match each buggy submission with its corresponding human fix. By using the similarity between the buggy solution and the human fix as a baseline, we evaluate the quality of LLM-generated bug fixes on 3 OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1). We check if the generated solutions solve the problem by using the Codeforces-R1 dataset, an openly available dataset that has tests generated with the DeepSeek-R1 model. Our findings suggest that LLMs tend to modify more lines than necessary compared to human fixes and, in some cases, generate entirely new solutions. We also observe that LLMs solve more problems correctly when allowed to generate solutions from scratch rather than patch buggy submissions, even when those submissions are close to the human patch. This has important implications for the design of AI-assisted programming tools, particularly in supporting user debugging processes and promoting incremental problem-solving strategies rather than solution replacement.

Figures & tables

Explore similar work

CardsList
  1. Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming

    Jun 11, 2026Tingqiang Xu, Hangrui Zhou, Tianle Cai +2Raw Judge OutputsCode Generation

  2. If It's Not Buggy, Don't Fix It: On the Dynamics of Iterative Bug-fixing with LLMs

    Sep 9, 2026Xietao Wang-Lin, Anton Isopoussu, Louis MahonBugIterative

  3. Unlocking LLM Code Correction with Iterative Feedback Loops

    Jun 16, 2026Le Zhang, Suresh KothariCode GenerationLarge Language Models Fail