The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, we propose RuleEvolve, a self-evolving framework for coding rules. RuleEvolve maintains a pool of candidate coding rules and iteratively improves them. In each iteration, it employs an LLM-powered mutator module to generate variants from existing candidates, and then uses a judge module to evaluate these variants and update the pool with the best-performing ones. Extensive evaluations across two coding-agent frameworks, four backbone LLMs, and three benchmarks demonstrate that RuleEvolve outperforms both manual engineering and existing prompt optimization baselines in terms of functional correctness of the generated code, code length, and/or generation cost (e.g., tokens used).
Figures & tables
Figure 1 : Pipeline of RuleEvolve.
Method
PR ↑
CLn ↓
CC ↓
Time ↓
Tok ↓
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
No rule
0.47
0.25
0.97
35.7
42.3
26.4
1156
1406
812.4
6.00
5.44
4.24
623.5
666.5
502.9
Manual rule
0.45
0.23
0.98
69.0
84.8
44.4
2372
2946
1136
8.75
10.5
5.08
1022
1147
663.8
Prompt-Ops
0.48
0.23
0.89
13.8
15.7
4.50
474.4
583.0
126.5
2.30
2.90
1.33
581.7
588.7
406.3
RuleEvolve
0.49
0.26
0.99
12.6
20.8
4.40
453.1
745.3
114.4
3.02
2.90
1.48
544.5
595.8
367.5
Table 1 : OpenAI SDK results. For each method, we report the performance of the backbone that achieved the highest pass rate across benchmarks. Metrics include Pass Rate ( PR ), Avg. Code Lines ( CLn ), Avg. Code Characters ( CC ), Avg. Generation Time ( Time ), and Avg. Token Usage ( Tok ).
Method
PR ↑
CLn ↓
CC ↓
Time ↓
Tok ↓
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
No rule
0.49
0.29
0.99
35.4
43.2
17.5
1060
1324
503.4
10.6
12.8
6.89
1143
1238
697.5
Manual rule
0.49
0.29
0.99
42.9
53.5
21.7
1350
1722
577.6
12.2
14.1
7.14
1293
1443
727.1
Prompt-Ops
0.43
0.25
0.99
20.1
27.5
19.3
615.2
880.2
548.7
40.2
41.4
12.6
903.1
986.9
725.5
RuleEvolve
0.52
0.27
0.99
26.6
31.8
12.7
814.6
1009
373.5
9.76
11.6
7.15
1016
1075
612.6
Table 2 : OpenHands results. For each method, we report the performance of the backbone that achieved the highest pass rate across benchmarks. Metrics include Pass Rate ( PR ), Avg. Code Lines ( CLn ), Avg. Code Characters ( CC ), Avg. Generation Time ( Time ), and Avg. Token Usage ( Tok ).
Method
PR ↑
CLn ↓
CC ↓
Time (s) ↓
Tok ↓
No rule
0.502±0.015
32.5±0.4
1028±8.1
8.66±0.74
1181±9.3
Public rule
0.498±0.032
44.9±0.7
1482±25
10.21±1.15
2049±35
Prompt-Ops
0.502±0.027
26.3±0.4
811.4±10
6.81±0.62
1130±15
RuleEvolve (OCBA)
0.502±0.026
24.3±0.4
794.4±14
9.67±1.90
1019±18
Table 3 : Multi-seed results on BigCodeBench using GPT-4.1-mini as backbone LLM. Values are mean ± standard deviation (SD) across 10 seeds.
Method
PR ↑
CLn ↓
CC ↓
Time (s) ↓
Tok ↓
No rule
0.330
8.5
380
131.17
911
Public rule
0.190
9.1
413
166.45
867
Prompt-Ops
0.080
5.4
236
175.24
684
RuleEvolve
0.340
5.7
254
297.46
722
Table 4 : SWE-bench Lite results with GPT-4.1-mini as backbone LLM.
Figure 2 : Ablation study results.
Figure 3 : Pass@k results.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Method
PR ↑
CLn ↓
CC ↓
Time ↓
Tok ↓
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
gpt-4.1 -mini
No rule
0.47
0.25
0.97
35.7
42.3
26.4
1156
1406
812.4
6.00
5.44
4.24
623.5
666.5
502.9
Manual rule
0.46
0.20
0.92
58.4
68.6
37.5
2035
2448
1036
12.7
9.61
5.36
852.0
939.1
547.9
Prompt-Ops
0.45
0.24
0.92
20.9
24.1
6.30
650.8
790.7
180.0
3.06
3.94
1.79
540.5
552.8
335.4
RuleEvolve
0.46
0.24
0.93
16.9
19.3
5.10
545.2
674.8
134.7
2.86
3.25
1.37
479.3
484.5
284.5
gpt-5 -mini
No rule
0.36
0.20
0.95
88.2
94.6
46.4
3053
3377
1248
29.5
28.5
13.7
2185
2257
1197
Appendix
Table 5 : Comprehensive OpenAI SDK results on BigCodeBench (BCB), BigCodeBench-Hard (BCB-H), and HumanEval (HE). Metrics include Pass Rate ( PR ), Avg. Code Lines ( CLn ), Avg. Code Characters ( CC ), Avg. Generation Time ( Time ), and Avg. Token Usage ( Tok ).
Backbone
Method
PR ↑
CLn ↓
CC ↓
Time ↓
Tok ↓
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
BCB
BCB-H
HE
gpt-4.1 -mini
No rule
0.48
0.27
0.96
29.8
34.6
19.2
937.8
1127
608.1
8.28
12.6
20.8
1079
1124
767.5
Manual rule
0.48
0.28
0.95
42.3
47.2
22.9
1402
1597
691.8
9.11
13.7
18.3
1320
1383
813.4
Prompt-Ops
0.43
0.22
0.71
20.1
27.5
17.7
615.2
880.2
544.1
40.2
41.4
57.9
903.1
986.9
707.7
RuleEvolve
0.50
0.28
0.98
23.8
28.5
16.9
768.9
944.6
536.3
9.54
14.6
23.0
991.6
1028
725.0
gpt-5 -mini
No rule
0.40
0.18
0.97
66.1
74.3
30.2
2203
2509
854.7
12.9
14.8
7.10
1793
1930
944.4
Appendix
Table 6 : Comprehensive OpenHands results on BigCodeBench (BCB), BigCodeBench-Hard (BCB-H), and HumanEval (HE). Metrics include Pass Rate ( PR ), Avg. Code Lines ( CLn ), Avg. Code Characters ( CC ), Avg. Generation Time ( Time ), and Avg. Token Usage ( Tok ).
LLM-based coding agents repeat the same classes of mistakes across sessions because they lack a mechanism to retain corrections from human review feedback. We present a closed-loop framework in which every accepted review comment is codified as a persistent behavioral rule, progressively expanding the set of error classes the agent can self-detect. The framework combines an accumulating rule set in a version-controlled instruction file, a self-review checklist executed before code submission, and automated validation that ensures rule set integrity as it grows. In deployment across a 35+ service microservices platform, the rule set grew from 5 to 18 behavioral rules, 15+ language-specific standards, and a 15-item self-review checklist, all derived from real review feedback. We present empirical results from 11 recorded working sessions spanning code generation, PR review, incident investigation, and cross service refactoring. We observe that accumulated rules shift review effort from low-level correctness toward design-level validation, achieve a measured 0% recurrence rate for ruled-against error classes, and transfer across heterogeneous agent interfaces. We compare our approach against related work in experiential LLM learning (Reflexion, ExpeL, Voyager) and automated code review (CodeReviewer, SWE-bench agents), showing that our framework achieves persistent cross-session learning without weight updates, operates on production codebases rather than synthetic benchmarks, and addresses an orthogonal dimension (behavioral consistency over time) that existing benchmarks do not measure. The result is a coding agent that improves with every review cycle, accumulating the engineering wisdom of its human collaborators without changing a single model weight.
Coding agents produce rich trajectories while solving software-engineering tasks. To enable agent self-evolution, these trajectories can be distilled into reusable procedural skills that compactly encode experience to guide future behavior. However, existing skill construction and maintenance methods often rely on fixed prompts and heuristic update rules, leaving it unclear how knowledge should be selected, abstracted, and maintained to best serve downstream agents. We propose CODESKILL, an LLM-based framework that reformulates skill extraction and skill-bank maintenance as a learnable management policy. CODESKILL extracts multi-granularity procedural skills from coding-agent trajectories, evolves skills with new experience, and maintains a compact skill bank for future task solving. We train CODESKILL with reinforcement learning, using a hybrid reward that combines dense rubric-based skill-quality feedback with sparse verifiable execution feedback from the frozen downstream agent. Experiments on EnvBench, SWE-Bench Verified, and Terminal-Bench 2 show that CODESKILL improves average pass rate by 9.69 over the no-skill baseline and by 4.01 over the strongest prompt-based or memory baseline, while maintaining the skill bank at a stable size during iterative construction.
Yanzhou Li, Yiran Zhang, Xiaoyu Zhang +2
Nanyang Technological University · 2Zhejiang University
Large Language Models (LLMs) excel at code generation but remain heavily reliant on large-scale annotated solutions and verification-based supervision, which constrains scalability and hinders sustained self-improvement. Recent solver--verifier frameworks exploit program execution as an automatic supervision signal, but their effectiveness degrades as solvers become moderately strong: verifier-generated tests increasingly confirm semantic correctness rather than exposing the remaining failure modes. We propose \textbf{ACE}, a self-evolving code generation framework based on a solver--adversary architecture that prioritizes active failure discovery through execution-centric supervision. A single LLM alternates between generating candidate programs and producing adversarial unit test inputs optimized to induce execution-level failures, such as runtime errors, exceptions, or non-termination. Supervision is derived solely from execution outcomes: robust programs are selected for supervised fine-tuning, while adversarial tests are optimized via Kahneman--Tversky Optimization using execution-derived preferences. Notably, the entire training loop requires no ground-truth code or external reward models. Experiments on CodeContests, MBPP, and LiveCodeBench demonstrate that ACE consistently outperforms strong solver--verifier baselines, achieving 3--7% absolute gains in pass@1, with larger improvements on out-of-distribution benchmarks, while maintaining competitive or improved inference efficiency.
Yixu Huang, Xinglei Yu, Zhongyu Wei
School of Data Science Fudan University Shanghai, China