Do Self-Evolving Skills Generalize to Held-Out Tasks?
Organizations: SANKEN, Osaka University
Abstract
AI agents can externalize what they learn from past tasks into reusable \emph{skills}, such as procedures, checklists, code, or other executable artifacts, that can be retrieved and reused when solving new tasks. Self-evolving skill methods keep rewriting these skills after each round of practice on training tasks, and the skill is then used on new tasks of the same kind. We ask a question: does the improvement a skill shows on its training tasks carry over to new test tasks? We test five self-evolving methods and a one-shot skill on six benchmarks, with the same model, the same agent, and the same train/test split for every method. Of the 21 skills that improve on their training tasks, 5 keep all of that improvement on the test tasks, 13 keep part of it, and 3 keep none of it. No existing method is best everywhere. When we read the skills, the ones that carry over badly often fix details that should depend on the task, such as column names and output files, or turn a fix for one failure into a rule for every task. An LLM judge that reads the skill content can often see this: it ranks finished skills the same way the test results do in 86% of pairs. But it predicts the effect of a single edit poorly, so edits still have to be tested by running them. Based on these findings, we describe Generalizable Skill Optimization (GSO), which keeps only a guide for writing skills and writes a new skill for each task; it scores highest on all six benchmarks.
Figures & tables
| Dimension | Weight | Question the judge answers |
|---|---|---|
| Generalizability | 30% | Does the skill describe a reusable procedure, or does it hard-code details of specific tasks? |
| Applicability | 20% | Does it say when and where its rules apply? |
| Executability | 20% | Can an agent turn the instructions into concrete actions? |
| Robustness | 20% | Does it plan for failures without imposing harmful global rules? |
| Clarity/efficiency | 10% | Is it short enough to follow without losing needed detail? |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Panel A: Spreadsheet and RBioBench | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| SpreadsheetBench | RBio Clinical | RBio Omics | |||||||
| Method | Train | Val. | Test | Train | Val. | Test | Train | Val. | Test |
| No Skill | 42.5 | 30.0 | 27.5 | 42.5 | 45.0 | 58.8 | 42.5 | 45.0 | 43.5 |
| One-shot LLM Skill | 52.5 | 50.0 | 32.5 | 42.5 | 45.0 | 64.7 | 42.5 | 45.0 | 39.1 |
| SkillGen | 55.0 | 40.0 | 32.5 | 42.5 | 50.0 | 58.8 | 42.5 | 50.0 | 39.1 |
| Trace2Skill | 55.0 | 35.0 | 37.5 | 42.5 | 40.0 | 58.8 | 42.5 | 40.0 | 43.5 |
| Method | Persistent artifact | Update from training feedback | Native candidate retention |
|---|---|---|---|
| SkillGen | one skill | contrast successful and failed trajectories; iterative refinement | paired repairs minus regressions; deployment gate |
| Trace2Skill | skill directory | parallel trajectory-local patches; hierarchical consolidation | deduplication and conflict resolution |
| SkillOpt | one compact document | minibatch reflection; bounded add/delete/replace edits | strict validation; rejected-edit memory |
| EvoSkill | repository of skill folders | failure diagnosis; create or edit one skill | fixed-capacity validation frontier |
| GEPA | any textual component | reflective mutation from trajectory feedback; optional merge | Pareto frontier of candidates |
| GSO (Sec. 4 ) | metaskill; transient task-local skill | first-fault attribution; single-module edit | regression-aware paired gate |
| Stage | Information available | Output of the stage |
|---|---|---|
| Training | Task prompt, permitted inputs, execution feedback, and method-native update state | Candidate skill artifacts and checkpoints |
| Selection | Validation tasks and their executable outcomes | Frozen artifact, native stop decision, and artifact identifier |
| Test | Test prompt, permitted inputs, and execution environment | Final task outcomes; no subsequent method update |
| Domain | Tasks | Held-out boundary | Output and evaluation |
|---|---|---|---|
| SpreadsheetBench | 40/20/40 | Task; shared operation types | Edited workbook; workbook content and structure checks |
| RBioBench | 40/20/40 | Task; shared package families | Scientific workflow outputs; RBioBench evaluator for Clinical and Omics tracks |
| SWE-bench Verified | 40/20/40 | Test repository | Source-code patch; repository-specific test suite |
| HealthBench | 40/20/40 | Medical specialty | Retrieved evidence and response; family-balanced rubric score plus artifact gate |
| SciVisAgentBench | 66/21/20 | Data family | Visualization artifact; family-balanced artifact evaluator |
| Figure | Primary input | Evidence role | Interpretation boundary |
|---|---|---|---|
| 1 | Two candidate edits from one SkillOpt run | Illustrates similar training gains with opposite test effects | Two edits from one run; not a prevalence estimate |
| 2 | Stored optimization checkpoints | Shows development and test trajectories | Descriptive post-freeze analysis; test does not select artifacts |
| 6 | Main result table train/test columns | Shows signed train-to-test score differences | Reports the two displayed split scores directly |
| 3 | Artifact excerpts and execution traces | Illustrates task-local binding and validation | Qualitative execution evidence from the displayed cases |
| 4 | Method description | Shows the GSO loop | Schematic; no data |
| 5 | Main result table | Compares final test performance across domains | Preserves benchmark-specific metrics and denominators |