A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.
Figures & tables
Backbone LLM
Pipeline
Gemini 3.1 Pro Preview
GPT-5.4
Claude Opus 4.6
GLM-5
MiniMax M2.7
Kimi K2.5
Vanilla
V-Gemini
V-GPT
V-Claude
V-GLM
V-MiniMax
V-Kimi
Ablation (no subclass prior)
Abl-Gemini
–
–
–
–
Abl-Kimi
RAG
RAG-Gemini
–
–
–
–
RAG-Kimi
Agentic web search
Agentic-Gemini
–
–
–
–
Agentic-Kimi
Table 1 . The 12 LLM-based fix-generation configurations reused from Gön et al. ( Gon et al., 2026 ) , with the short identifiers used throughout this paper. Rows are pipelines and columns are backbone LLMs; a dash marks a combination that was not run.
Figure 1 . Overview of the evaluation pipeline. Flow diagram: a bug report, the maintainer's fix and 12 LLM fixes enter a Docker environment; two sanity checks bracket a rollback from fixed to buggy state; each LLM fix is applied by an executor and checked.
Executor
Resolution (%)
Applicability (%)
Resolution among applicable (%)
Claude Sonnet 5
49.7 (160)
66.1 (213)
69.5 (148)
OpenCUA-72B
45.7 (147)
99.1 (319)
45.8 (146)
Muse Glimmer
14.6 (47)
70.5 (227)
18.1 (41)
Table 2 . Resolution and applicability per executor over the 322 fixes (fix counts in parentheses). The last column restricts the resolution rate to the fixes the executor found applicable.
Vanilla
Ablation
RAG
Agentic
Gem.
GPT
Cla.
GLM
MiniM.
Kimi
Gem.
Kimi
Gem.
Kimi
Gem.
Kimi
Res.
S5
44.4
40.7
74.1
44.4
25.9
44.4
59.3
51.9
48.1
44.4
63.0
56.0
OC
51.9
48.1
51.9
37.0
44.4
33.3
44.4
40.7
48.1
55.6
55.6
36.0
MG
11.1
11.1
25.9
18.5
11.1
0.0
14.8
11.1
7.4
11.1
25.9
28.0
App.
S5
55.6
55.6
88.9
51.9
44.4
66.7
70.4
74.1
63.0
70.4
81.5
72.0
OC
100.0
96.3
100.0
100.0
100.0
100.0
100.0
96.3
96.3
100.0
100.0
100.0
Table 3 . Resolution (Res.) and applicability (App.) rate (%) per configuration and executor. Each configuration has 27 fixes (Agentic-Kimi: 25). Backbones: Gem. = Gemini 3.1 Pro, Cla. = Claude Opus 4.6, MiniM. = MiniMax M2.7. Executors: S5 = Sonnet 5, OC = OpenCUA-72B, MG = Muse Glimmer.
Sonnet 5
OpenCUA-72B
Muse Glimmer
Subclass
Iss.
Fixes
Res.
App.
Res.
App.
Res.
App.
Faulty Configuration
5
178
43.8
66.3
65.7
98.9
13.5
69.7
Wrong Version
1
36
50.0
44.4
8.3
100.0
0.0
83.3
External System
3
108
59.3
73.1
25.0
99.1
21.3
67.6
Table 4 . Resolution and applicability rate (%) per subclass.
Executor
Run 1
Run 2
Run 3
Any run
All runs agree
Claude Sonnet 5
136 (42.2%)
125 (38.8%)
130 (40.4%)
160 (49.7%)
82.3%
OpenCUA-72B
100 (31.1%)
109 (33.9%)
100 (31.1%)
147 (45.7%)
73.9%
Muse Glimmer
34 (10.6%)
35 (10.9%)
32 (9.9%)
47 (14.6%)
92.2%
Table 5 . Resolved fixes per individual executor run, fixes resolved in at least one run, and the share of fixes on which all three runs agreed ( n=322 ).
Faulty Configuration
Wrong Version
External System
Issue
20537
22987
25057
37372
45912
25379
22252
24747
31108
Total
Fixes
36
36
35
36
35
36
36
36
36
322
Sonnet 5
9
13
10
20
26
18
16
19
29
160
OpenCUA-72B
19
20
19
36
23
3
12
0
15
147
Muse Glimmer
3
9
0
12
0
0
10
0
13
47
Table 6 . Resolved fixes per issue and executor. Column headers are Brave GitHub issue numbers, grouped by subclass.
Executor
Human 1
Human 2
Consensus
κ (consensus)
Claude Sonnet 5
81.3 [71.1, 88.5]
78.7 [68.1, 86.4]
88.1 [77.5, 94.1]
0.74
OpenCUA-72B
68.0 [56.8, 77.5]
57.3 [46.1, 67.9]
66.1 [53.4, 76.9]
0.28
Muse Glimmer
78.7 [68.1, 86.4]
68.0 [56.8, 77.5]
79.7 [67.7, 88.0]
0.48
Human 1 vs. Human 2
78.7 [68.1, 86.4]
0.54
Table 7 . Agreement (%) between each executor and human execution on the 75 randomly sampled fixes, with 95% Wilson confidence intervals. Consensus: the 59 fixes on which both humans reached the same outcome.