Self-Evolving Coding Rules for AI Coding Agents
Organizations: Duke University
Abstract
The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, we propose RuleEvolve, a self-evolving framework for coding rules. RuleEvolve maintains a pool of candidate coding rules and iteratively improves them. In each iteration, it employs an LLM-powered mutator module to generate variants from existing candidates, and then uses a judge module to evaluate these variants and update the pool with the best-performing ones. Extensive evaluations across two coding-agent frameworks, four backbone LLMs, and three benchmarks demonstrate that RuleEvolve outperforms both manual engineering and existing prompt optimization baselines in terms of functional correctness of the generated code, code length, and/or generation cost (e.g., tokens used).
Figures & tables
| Method | PR | CLn | CC | Time | Tok | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | |
| No rule | 0.47 | 0.25 | 0.97 | 35.7 | 42.3 | 26.4 | 1156 | 1406 | 812.4 | 6.00 | 5.44 | 4.24 | 623.5 | 666.5 | 502.9 |
| Manual rule | 0.45 | 0.23 | 0.98 | 69.0 | 84.8 | 44.4 | 2372 | 2946 | 1136 | 8.75 | 10.5 | 5.08 | 1022 | 1147 | 663.8 |
| Prompt-Ops | 0.48 | 0.23 | 0.89 | 13.8 | 15.7 | 4.50 | 474.4 | 583.0 | 126.5 | 2.30 | 2.90 | 1.33 | 581.7 | 588.7 | 406.3 |
| RuleEvolve | 0.49 | 0.26 | 0.99 | 12.6 | 20.8 | 4.40 | 453.1 | 745.3 | 114.4 | 3.02 | 2.90 | 1.48 | 544.5 | 595.8 | 367.5 |
| Method | PR | CLn | CC | Time | Tok | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | |
| No rule | 0.49 | 0.29 | 0.99 | 35.4 | 43.2 | 17.5 | 1060 | 1324 | 503.4 | 10.6 | 12.8 | 6.89 | 1143 | 1238 | 697.5 |
| Manual rule | 0.49 | 0.29 | 0.99 | 42.9 | 53.5 | 21.7 | 1350 | 1722 | 577.6 | 12.2 | 14.1 | 7.14 | 1293 | 1443 | 727.1 |
| Prompt-Ops | 0.43 | 0.25 | 0.99 | 20.1 | 27.5 | 19.3 | 615.2 | 880.2 | 548.7 | 40.2 | 41.4 | 12.6 | 903.1 | 986.9 | 725.5 |
| RuleEvolve | 0.52 | 0.27 | 0.99 | 26.6 | 31.8 | 12.7 | 814.6 | 1009 | 373.5 | 9.76 | 11.6 | 7.15 | 1016 | 1075 | 612.6 |
| Method | PR | CLn | CC | Time (s) | Tok |
|---|---|---|---|---|---|
| No rule | |||||
| Public rule | |||||
| Prompt-Ops | |||||
| RuleEvolve (OCBA) |
| Method | PR | CLn | CC | Time (s) | Tok |
|---|---|---|---|---|---|
| No rule | |||||
| Public rule | |||||
| Prompt-Ops | |||||
| RuleEvolve |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Backbone | Method | PR | CLn | CC | Time | Tok | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | ||
| gpt-4.1 -mini | No rule | 0.47 | 0.25 | 0.97 | 35.7 | 42.3 | 26.4 | 1156 | 1406 | 812.4 | 6.00 | 5.44 | 4.24 | 623.5 | 666.5 | 502.9 |
| Manual rule | 0.46 | 0.20 | 0.92 | 58.4 | 68.6 | 37.5 | 2035 | 2448 | 1036 | 12.7 | 9.61 | 5.36 | 852.0 | 939.1 | 547.9 | |
| Prompt-Ops | 0.45 | 0.24 | 0.92 | 20.9 | 24.1 | 6.30 | 650.8 | 790.7 | 180.0 | 3.06 | 3.94 | 1.79 | 540.5 | 552.8 | 335.4 | |
| RuleEvolve | 0.46 | 0.24 | 0.93 | 16.9 | 19.3 | 5.10 | 545.2 | 674.8 | 134.7 | 2.86 | 3.25 | 1.37 | 479.3 | 484.5 | 284.5 | |
| gpt-5 -mini | No rule | 0.36 | 0.20 | 0.95 | 88.2 | 94.6 | 46.4 | 3053 | 3377 | 1248 | 29.5 | 28.5 | 13.7 | 2185 | 2257 | 1197 |
| Backbone | Method | PR | CLn | CC | Time | Tok | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | BCB | BCB-H | HE | ||
| gpt-4.1 -mini | No rule | 0.48 | 0.27 | 0.96 | 29.8 | 34.6 | 19.2 | 937.8 | 1127 | 608.1 | 8.28 | 12.6 | 20.8 | 1079 | 1124 | 767.5 |
| Manual rule | 0.48 | 0.28 | 0.95 | 42.3 | 47.2 | 22.9 | 1402 | 1597 | 691.8 | 9.11 | 13.7 | 18.3 | 1320 | 1383 | 813.4 | |
| Prompt-Ops | 0.43 | 0.22 | 0.71 | 20.1 | 27.5 | 17.7 | 615.2 | 880.2 | 544.1 | 40.2 | 41.4 | 57.9 | 903.1 | 986.9 | 707.7 | |
| RuleEvolve | 0.50 | 0.28 | 0.98 | 23.8 | 28.5 | 16.9 | 768.9 | 944.6 | 536.3 | 9.54 | 14.6 | 23.0 | 991.6 | 1028 | 725.0 | |
| gpt-5 -mini | No rule | 0.40 | 0.18 | 0.97 | 66.1 | 74.3 | 30.2 | 2203 | 2509 | 854.7 | 12.9 | 14.8 | 7.10 | 1793 | 1930 | 944.4 |
| Category | Sections | Chars attributable | Per section |
|---|---|---|---|
| Output format / self-containment | 3 | 153.0 | 51.0 |
| Algorithm knowledge | 4 | 83.0 | 20.7 |
| Brevity style | 1 | 61.6 | 61.6 |
| API / library idiom | 2 | 22.8 | 11.4 |
| Correctness tactics | 2 | 9.9 | 5.0 |
| Process / budget | 4 | -5.0 | -1.2 |
| Metric | No rule | RuleEvolve |
|---|---|---|
| Duplicated-block ratio | 0.10% | 0.00% |
| Cyclomatic complexity (total) | 5.37 | 3.95 |
| Halstead volume (total) | 53.0 | 28.0 |
| Maintainability index | 85.1 | 78.3 |
| Cyclomatic complexity per SLOC | 0.213 | 0.260 |
| Docstring rate | 47.0% | 1.00% |