Investigating The Smells of LLM Generated Code
Organizations: School of Engineering, Computing and Mathematics, Oxford Brookes University, Oxford, OX3 0BP, UK
Abstract
Context: Large Language Models (LLMs) are increasingly being used to generate program code. Much research has been reported on the functional correctness of generated code, but there is far less on code quality. Objectives: In this study, we propose a scenario-based method of evaluating the quality of LLM-generated code to identify the weakest scenarios, for which the quality of LLM-generated code should be improved. Methods: The method measures code smells, an important indicator of code quality, and compares them with a baseline formed from reference solutions of professionally written code. The test dataset is divided into various subsets according to the topics of the code and complexity of the coding tasks to represent different scenarios of using LLMs for code generation. We will also present an automated test system for this purpose and report experiments with the Java programs generated in response to prompts given to four state-of-the-art LLMs: Gemini Pro, ChatGPT, Codex, and Falcon. Results: We find that LLM-generated code has a higher incidence of code smells compared to reference solutions. Falcon performed the least badly, with a smell increase of 42.28%, followed by Gemini Pro (62.07%), ChatGPT (65.05%) and finally Codex (84.97%). The average smell increase across all LLMs was 63.34%, comprising 73.35% for implementation smells and 21.42% for design smells. We also found that the increase in code smells is greater for more complex coding tasks and for more advanced topics, such as those involving object-orientated concepts. Conclusion: In terms of code smells, LLMs' performances on various coding task complexities and topics are highly correlated to the quality of human-written code in the corresponding scenarios. However, the quality of LLM-generated code is noticeably poorer than human-written code.
Figures & tables
| Work | Aims | Tools | Language | Usage | Dataset (Size) | LLM | Smell Types |
| Siddiq, et al. 2022 | Investigating the impact of code smells in training datasets on the quality of generated code | Pylint | Python | Detect code smells in training datasets and generated codes | HumanEval (164) | GPT-Neo, GitHub Copilot | Implementation smells |
| Bandit | Detect security smells in training datasets and generated codes | Security smells | |||||
| Moratis, et al. 2024 | Assessing the quality of code generated via iterative conversations in the Write this code scenario | PMD | JavaScript | Detect code smell and measure code quality | DevGPT (47) | ChatGPT | Best practice, code style, error prone |
| Assessing the quality improvement of code generated via iterative conversations in the Improve this code scenario | Comparing code smells before and after refactoring | DevGPT (334) | |||||
| DePalma, et al. 2024 | Evaluating LLM capability for code refactoring | PMD | Java | Assessing the quality of the code before and after refactoring | Ma et al. [ 29 ] (40) | ChatGPT | Best practice, code style, design, documentation, error prone, multi-threading, performance, security |
| Liu, at el. 2024 | Evaluating LLM capability for fixing code quality issues | PMD, CheckStyle | Java | Assessing the quality of generated code | LMDefect + (2033) | ChatGPT | Implementation and design smells |
| Smell Name | Detection Rule(s) | Tool(s) Used |
|---|---|---|
| Inconsistent Naming Convention | Local Variable Naming Convention | PMD |
| /Local Variable Name | /CheckStyle | |
| Formal Parameter Naming Convention | PMD | |
| Method Naming Convention | PMD | |
| /Method Name | /CheckStyle | |
| Class Naming Convention | PMD |
| Smell Name | Detection Rules | Tool Used |
|---|---|---|
| Modularity | God Class | PMD |
| Data Class | PMD | |
| Too Many Methods | PMD | |
| Too Many Fields | PMD | |
| Use Utility Class | PMD | |
| Hide Utility Class Constructor | CheckStyle |
| Name | Year | Version | Size |
|---|---|---|---|
| Gemini Pro | 2023 | Gemini Pro 1.0 | Unknown |
| Falcon | 2023 | Falcon-7B | 7B |
| ChatGPT | 2023 | GPT-3.5-turbo | Unknown |
| Codex | 2021 | GPT-3 (Codex) | 12B |
| Feature | #Textbook Tasks | #Real Tasks |
|---|---|---|
| # Topics | 25 | 18 |
| # Tasks per Topic (average) | 20.00 | 27.78 |
| # Tasks per Topic (max) | 79 | 61 |
| # Tasks per Topic (min) | 8 | 5 |
| Input Length (average) | 18.55 | 21.54 |
| Input Length (max) | 35 | 31 |
| LLM Model | Avg VS | Inc.Rate (%) | MoE(%) | p-value |
|---|---|---|---|---|
| Reference | 19.378 | N/A | 6.045 | N/A |
| Falcon | 27.571 | 42.28% | 4.654 | 1.0749E-19 |
| Gemini Pro | 31.406 | 62.07% | 4.284 | 6.0923E-38 |
| ChatGPT | 31.789 | 64.05% | 4.029 | 3.0557E-42 |
| Codex | 35.844 | 84.97% | 3.708 | 3.9380E-68 |
| Model | Best Topics (VS) | Worst Topics (VS) | Most Improved Toipics (Inc %) | Most Worsened Topics (Inc %) |
|---|---|---|---|---|
| Baseline | Basic Exercise (1.96) | Searching & Sorting (5.67) | N/A | N/A |
| DateTime (2.18) | Polymorphism (5.72) | N/A | N/A | |
| String (2.23) | Inheritance (6.91) | N/A | N/A | |
| Falcon | Basic Exercise (1.99) | Polymorphism (6.91) | Interfaces (-23.34) | Array (58.16) |
| Lambda (2.41) | Inheritance (7.59) | Collections (-12.03) | OOP (96.78) | |
| Collections (2.56) | Encapsulation (9.91) | Lambda (-11.07) | Encapsulation (172.25) |
| Complexity | Baseline | Falcon | GeminiPro | ChatGPT | Codex | Average |
|---|---|---|---|---|---|---|
| Cyclomatic | 0.9344 | 0.9555 | 0.9916 | 0.9485 | 0.9962 | 0.9653 |
| Cognitive | 0.8800 | 0.9754 | 0.6312 | 0.3025 | 0.6578 | 0.6894 |
| Lines of Code | 0.9952 | 0.8694 | 0.9532 | 0.6347 | 0.8900 | 0.8685 |
| LLM | Avg of Code Smells | Increase | Margin of Error (%) | p-value | ||
|---|---|---|---|---|---|---|
| Correct | Incorrect | Rate (%) | Correct | Incorrect | of T-Test | |
| ChatGPT | 31.3279 | 32.9344 | 5.1278 | 5.0862 | 6.5124 | 0.2391 |
| Gemini Pro | 30.6299 | 33.8833 | 10.6216 | 5.0835 | 7.8972 | 0.0400 |
| Falcon | 32.0511 | 23.6830 | -26.1085 | 6.5659 | 6.2138 | 0.0000 |
| Codex | 35.6416 | 36.4379 | 2.2342 | 4.5385 | 6.4342 | 0.5839 |