AI agents can externalize what they learn from past tasks into reusable \emph{skills}, such as procedures, checklists, code, or other executable artifacts, that can be retrieved and reused when solving new tasks. Self-evolving skill methods keep rewriting these skills after each round of practice on training tasks, and the skill is then used on new tasks of the same kind. We ask a question: does the improvement a skill shows on its training tasks carry over to new test tasks? We test five self-evolving methods and a one-shot skill on six benchmarks, with the same model, the same agent, and the same train/test split for every method. Of the 21 skills that improve on their training tasks, 5 keep all of that improvement on the test tasks, 13 keep part of it, and 3 keep none of it. No existing method is best everywhere. When we read the skills, the ones that carry over badly often fix details that should depend on the task, such as column names and output files, or turn a fix for one failure into a rule for every task. An LLM judge that reads the skill content can often see this: it ranks finished skills the same way the test results do in 86% of pairs. But it predicts the effect of a single edit poorly, so edits still have to be tested by running them. Based on these findings, we describe Generalizable Skill Optimization (GSO), which keeps only a guide for writing skills and writes a new skill for each task; it scores highest on all six benchmarks.
Figures & tables
Figure 1: (a) Existing methods optimize skills across training tasks and reuse them on new tasks. We instead optimize a metaskill that generates a new skill for each task. (b) Two edits from the same SkillOpt run on RBioBench, quoted verbatim. Both improve training performance slightly, yet one helps on test tasks while the other hurts. An LLM judge that reads only the edit predicts that both will hurt. Scores are changes in success rate from the previous skill (percentage points).
Figure 2: Training and test success rates of SkillOpt and EvoSkill over twenty rounds on three benchmarks. Dots are per-round values and lines a three-round moving average; Δ is the change from round 1 to round 20, and shading marks the gap between training and test. On SWE-bench and HealthBench the test gain follows the training gain; on SpreadsheetBench it lags far behind.
Figure 3: What an overfit skill looks like, and what GSO writes instead. (a) Three skills learned by existing methods, quoted word for word. Each one fixes something that should depend on the task: column names and an output file, a checklist that applies to every task regardless of what the task asks, and a rule that every run must create a summary file. (b) Two skills that GSO wrote for two tasks it had never seen. Both name the actual objects, functions, parameters, and checks of the task in front of them, and both passed. The metaskill on the left ( GSO ’s guide for writing skills, Section 4 ) stays the same; only the skill on the right changes with the task.
Figure 4: How GSO works. The only thing kept across tasks is a metaskill, a guide for how to write a skill. For each new task, a fresh skill is written from the task’s own files and tools. The agent reads it and works in the environment in a loop, taking actions (such as running code) and reading what comes back, until the output is checked. The skill is then thrown away. When a task fails, the failure is traced to one part of the metaskill and only that part is edited; the edit stays only if it fixes more than twice as many validation tasks as it breaks.
Figure 5: Test score (%) of final skills on the six benchmarks, with the same model and the same tasks for every method. Labels mark GSO and the best existing method; the dashed line marks No Skill. SciVisAgentBench and HealthBench use family-balanced scores (Section 5 ).
Dimension
Weight
Question the judge answers
Generalizability
30%
Does the skill describe a reusable procedure, or does it hard-code details of specific tasks?
Applicability
20%
Does it say when and where its rules apply?
Executability
20%
Can an agent turn the instructions into concrete actions?
Robustness
20%
Does it plan for failures without imposing harmful global rules?
Clarity/efficiency
10%
Is it short enough to follow without losing needed detail?
Table 1: What the judge scores. Each dimension is rated 1–5 from the skill alone; the weighted sum is the judge score.
Figure 6: Training and test score of each final skill; each arrow goes from training to test (Appendix Table 2 ). No Skill’s own arrow shows how much harder or easier the test tasks are.
Figure 7: Judge scores for the final SpreadsheetBench skills, from the skill alone (questions in Table 1 ). (a) Weighted score: the GSO metaskill scores 92 and the other skills 64 to 82, a loose comparison because the questions were written for skills, not for a guide. (b) The five 1–5 dimensions. SkillOpt loses on clarity, Trace2Skill on applicability, EvoSkill on executability.
Figure 8: Does the skill the judge prefers also score higher on test? Each point is one of the 126 pairs of final skills from the same benchmark, with A the skill the judge prefers. The horizontal axis is the test score of A minus B; the vertical axis is the judge score of A minus B. Points to the right agree, points to the left disagree, and hollow points on the zero line are ties on test; ties on test or on judge score count as disagreements. No Skill is left out because it has no skill to score.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Panel A: Spreadsheet and RBioBench
SpreadsheetBench
RBio Clinical
RBio Omics
Method
Train
Val.
Test
Train
Val.
Test
Train
Val.
Test
No Skill
42.5
30.0
27.5
42.5
45.0
58.8
42.5
45.0
43.5
One-shot LLM Skill
52.5
50.0
32.5
42.5
45.0
64.7
42.5
45.0
39.1
SkillGen
55.0
40.0
32.5
42.5
50.0
58.8
42.5
50.0
39.1
Trace2Skill
55.0
35.0
37.5
42.5
40.0
58.8
42.5
40.0
43.5
Appendix
Table 2: Main results (%). Task benchmarks report accuracy; HealthBench reports the family-balanced physician-rubric score, and SciVisAgentBench reports family-balanced official evaluator score. Clinical and Omics are two author-constructed benchmarks in the RBioBench suite and share its task format and verifier. The two panels split the six reporting columns for readability; values are unchanged.
Method
Persistent artifact
Update from training feedback
Native candidate retention
SkillGen
one skill
contrast successful and failed trajectories; iterative refinement
reflective mutation from trajectory feedback; optional merge
Pareto frontier of candidates
GSO (Sec. 4 )
metaskill; transient task-local skill
first-fault attribution; single-module edit
regression-aware paired gate
Appendix
Table 3: Self-evolving skill methods as instances of one train–select–test loop. Native update and retention rules are preserved; under our protocol every final artifact is additionally selected on validation and frozen before test.
Stage
Information available
Output of the stage
Training
Task prompt, permitted inputs, execution feedback, and method-native update state
Candidate skill artifacts and checkpoints
Selection
Validation tasks and their executable outcomes
Frozen artifact, native stop decision, and artifact identifier
Test
Test prompt, permitted inputs, and execution environment
Final task outcomes; no subsequent method update
Appendix
Table 4: Information boundary across the three evaluation stages.
Domain
Tasks
Held-out boundary
Output and evaluation
SpreadsheetBench
40/20/40
Task; shared operation types
Edited workbook; workbook content and structure checks
RBioBench
40/20/40
Task; shared package families
Scientific workflow outputs; RBioBench evaluator for Clinical and Omics tracks
SWE-bench Verified
40/20/40
Test repository
Source-code patch; repository-specific test suite
HealthBench
40/20/40
Medical specialty
Retrieved evidence and response; family-balanced rubric score plus artifact gate