Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes
Organizations: Bilkent University, Ankara, Turkey
Abstract
A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.
Figures & tables
| Backbone LLM | ||||||
| Pipeline | Gemini 3.1 Pro Preview | GPT-5.4 | Claude Opus 4.6 | GLM-5 | MiniMax M2.7 | Kimi K2.5 |
| Vanilla | V-Gemini | V-GPT | V-Claude | V-GLM | V-MiniMax | V-Kimi |
| Ablation (no subclass prior) | Abl-Gemini | – | – | – | – | Abl-Kimi |
| RAG | RAG-Gemini | – | – | – | – | RAG-Kimi |
| Agentic web search | Agentic-Gemini | – | – | – | – | Agentic-Kimi |
| Executor | Resolution (%) | Applicability (%) | Resolution among applicable (%) |
| Claude Sonnet 5 | 49.7 (160) | 66.1 (213) | 69.5 (148) |
| OpenCUA-72B | 45.7 (147) | 99.1 (319) | 45.8 (146) |
| Muse Glimmer | 14.6 (47) | 70.5 (227) | 18.1 (41) |
| Vanilla | Ablation | RAG | Agentic | ||||||||||
| Gem. | GPT | Cla. | GLM | MiniM. | Kimi | Gem. | Kimi | Gem. | Kimi | Gem. | Kimi | ||
| Res. | S5 | 44.4 | 40.7 | 74.1 | 44.4 | 25.9 | 44.4 | 59.3 | 51.9 | 48.1 | 44.4 | 63.0 | 56.0 |
| OC | 51.9 | 48.1 | 51.9 | 37.0 | 44.4 | 33.3 | 44.4 | 40.7 | 48.1 | 55.6 | 55.6 | 36.0 | |
| MG | 11.1 | 11.1 | 25.9 | 18.5 | 11.1 | 0.0 | 14.8 | 11.1 | 7.4 | 11.1 | 25.9 | 28.0 | |
| App. | S5 | 55.6 | 55.6 | 88.9 | 51.9 | 44.4 | 66.7 | 70.4 | 74.1 | 63.0 | 70.4 | 81.5 | 72.0 |
| OC | 100.0 | 96.3 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 96.3 | 96.3 | 100.0 | 100.0 | 100.0 | |
| Sonnet 5 | OpenCUA-72B | Muse Glimmer | ||||||
| Subclass | Iss. | Fixes | Res. | App. | Res. | App. | Res. | App. |
| Faulty Configuration | 5 | 178 | 43.8 | 66.3 | 65.7 | 98.9 | 13.5 | 69.7 |
| Wrong Version | 1 | 36 | 50.0 | 44.4 | 8.3 | 100.0 | 0.0 | 83.3 |
| External System | 3 | 108 | 59.3 | 73.1 | 25.0 | 99.1 | 21.3 | 67.6 |
| Executor | Run 1 | Run 2 | Run 3 | Any run | All runs agree |
| Claude Sonnet 5 | 136 (42.2%) | 125 (38.8%) | 130 (40.4%) | 160 (49.7%) | 82.3% |
| OpenCUA-72B | 100 (31.1%) | 109 (33.9%) | 100 (31.1%) | 147 (45.7%) | 73.9% |
| Muse Glimmer | 34 (10.6%) | 35 (10.9%) | 32 (9.9%) | 47 (14.6%) | 92.2% |
| Faulty Configuration | Wrong Version | External System | ||||||||
| Issue | 20537 | 22987 | 25057 | 37372 | 45912 | 25379 | 22252 | 24747 | 31108 | Total |
| Fixes | 36 | 36 | 35 | 36 | 35 | 36 | 36 | 36 | 36 | 322 |
| Sonnet 5 | 9 | 13 | 10 | 20 | 26 | 18 | 16 | 19 | 29 | 160 |
| OpenCUA-72B | 19 | 20 | 19 | 36 | 23 | 3 | 12 | 0 | 15 | 147 |
| Muse Glimmer | 3 | 9 | 0 | 12 | 0 | 0 | 10 | 0 | 13 | 47 |
| Executor | Human 1 | Human 2 | Consensus | (consensus) |
| Claude Sonnet 5 | 81.3 [71.1, 88.5] | 78.7 [68.1, 86.4] | 88.1 [77.5, 94.1] | 0.74 |
| OpenCUA-72B | 68.0 [56.8, 77.5] | 57.3 [46.1, 67.9] | 66.1 [53.4, 76.9] | 0.28 |
| Muse Glimmer | 78.7 [68.1, 86.4] | 68.0 [56.8, 77.5] | 79.7 [67.7, 88.0] | 0.48 |
| Human 1 vs. Human 2 | 78.7 [68.1, 86.4] | 0.54 | ||