SkillEvoLean: Mutation-enhanced skill evolution for Lean provers
Organizations: Peking University
Abstract
Skill evolution offers a promising way to improve large language model agents without updating their parameters, but its use in formal theorem proving remains underexplored. Existing methods mainly target natural-language reasoning, improving skills by analyzing successful and failed trajectories and incrementally revising solving strategies. Although the Lean verifier provides reliable execution feedback, when all sampled trajectories fail, existing skill evolution methods lack successful trajectories from which to infer effective update directions. Furthermore, these methods also focus mainly on the root instruction file, thus underexploring the evolution of reference knowledge including mathematical concepts and proving techniques. To address these limitations, we propose a mutation-enhanced skill self-evolution framework for building skill-augmented Lean provers. The framework jointly evolves a high-level solving policy and its reference knowledge through progressive and mutation-based updates. Progressive evolution derives local improvements from successful and failed trajectories, while mutation is triggered when no complete proof can be generated, sampling mathematical concepts to produce and select new skill candidates under verifier feedback. We evaluate our method on MiniF2F, PutnamBench, the 2025 International Mathematical Olympiad (IMO 2025), and the 2026 USA Mathematical Olympiad (USAMO 2026). Under the same backbone model, trajectorysampling budget, and test-time compute, our method achieves proof success rates of 100.0%, 90.6%, 4/6, and 4/6, respectively, with GPT-5.5, outperforming the baseline methods. Further analysis shows that concept-guided mutation outperforms random-text-guided mutation by 6.9 and 8.2 percentage points on MiniF2F and PutnamBench, respectively, while solving one additional problem on both IMO 2025 and USAMO 2026.
Figures & tables
| Operator | Function |
|---|---|
| Generates the initial high-level solving policy stored in SKILL.md . | |
| Merges, deduplicates, and filters candidate proof patterns. | |
| Compares successful and failed trajectories to identify behavioral differences and possible causes of failure. | |
| Locates the skill component or reference entry that should be updated. | |
| Generates concrete modification suggestions from trajectory analysis and verifier feedback. | |
| Rewrites the current skill according to the generated suggestions while preserving its overall structure. |
| Backbone Model | Method | MiniF2F | PutnamBench | IMO 2025 | USAMO 2026 |
|---|---|---|---|---|---|
| DeepSeek-V4-flash | No Skill | 61.5 | 55.3 | 0/6 | 0/6 |
| Human Skill | 82.0 | 58.2 | 1/6 | 1/6 | |
| LLM Skill | 70.5 | 51.8 | 0/6 | 0/6 | |
| Trace2Skill | 77.0 | 55.3 | 0/6 | 0/6 | |
| SkillOpt | 77.9 | 58.8 | 0/6 | 1/6 | |
| Ours | 94.7 | 77.1 | 2/6 | 3/6 |
| Variant | MiniF2F | PutnamBench | IMO | USAMO |
|---|---|---|---|---|
| Full Method | 98.4 | 81.8 | 2/6 | 3/6 |
| w/o Prog. | 92.1 | 74.5 | 1/6 | 2/6 |
| w/o Mut | 82.3 | 63.5 | 0/6 | 0/6 |
| Strategy | MiniF2F | PutnamBench | IMO | USAMO |
|---|---|---|---|---|
| UM | 83.7 | 62.4 | 0/6 | 1/6 |
| RTGM | 91.5 | 73.6 | 1/6 | 2/6 |
| CGM (Ours) | 98.4 | 81.8 | 2/6 | 3/6 |
| Method | Success rate (%) | Input tokens (K/problem) | Output tokens (K/problem) | Lean calls /problem | Time (min/problem) | Tokens/solved problem (K) |
|---|---|---|---|---|---|---|
| No Skill | 61.2 | 80.0 | 20.0 | 10.0 | 20.0 | 163.4 |
| Human Skill | 64.1 | 88.0 | 18.0 | 9.0 | 19.5 | 165.4 |
| LLM Skill | 58.2 | 92.0 | 22.0 | 11.0 | 22.0 | 195.9 |
| Trace2Skill | 61.8 | 90.0 | 18.0 | 9.5 | 20.0 | 174.8 |
| SkillOpt | 66.5 | 84.0 | 17.0 | 8.5 | 18.0 | 151.9 |
| Ours | 81.8 | 82.0 | 14.0 | 8.0 | 17.0 | 117.4 |
| Method | Retained (per 100) | Gained (per 100) | Lost (per 100) | Retention (%) | Net gain (pp) | Historical forgetting (%) |
|---|---|---|---|---|---|---|
| MiniF2F-test (initial success assumption: 77.5%) | ||||||
| Frozen (control) | 76.5 | 1.0 | 1.0 | 98.7 | 0.0 | 3.1 |
| w/o Prog. | 73.5 | 18.6 | 4.0 | 94.8 | 14.6 | 5.6 |
| w/o Mut. | 76.0 | 6.3 | 1.5 | 98.1 | 4.8 | 2.9 |
| Full Method | 76.5 | 21.9 | 1.0 | 98.7 | 20.9 | 1.4 |
| PutnamBench (initial success assumption: 58.2%) | ||||||