Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improvements for new tasks can overwrite procedures needed for earlier ones. In continual learning, Orthogonal Gradient Descent (OGD) addresses analogous interference by projecting a new-task gradient onto a subspace that locally preserves prior predictions. Natural-language skill revisions, however, have neither gradients nor a canonical vector space in which such a projection can be performed. We introduce \emph{Semantic-Scope Projected Evolution} (SSPE), which transfers the functional principle of gradient projection from parameter space to behavior space. SSPE treats an unconstrained skill revision as a proposed update, identifies acquired capabilities with which it may interfere, and uses the observed gains and regressions to construct a compatible revision rather than merely rejecting the update. This enables one shared skill to evolve across latent and recurring task contexts without exposing semantic domain identities to the evolution model. Across controlled synthetic streams and heterogeneous real-agent benchmarks, SSPE improves final cross-domain competence and mitigates forgetting relative to strong skill-evolution baselines. The evolved skill also retains the strongest average performance after transfer to a different executor model. These results establish semantic projection as a promising principle for stable and adaptive self evolution of language agents.
Figures & tables
OGD
SSPE
Current gradient gt
Unconstrained revision St
Protected output sensitivities
Capability capsules and audit tasks
Alignment with a protected direction
Measured historical regression
Orthogonal projection
Revision conditioned on observed conflicts
Zero projected step
Explicit no update
Table 1: SSPE transfers the operational invariant of OGD, not its vector operations.
Figure 1: SSPE expands one update in a heterogeneous stream. Current evidence produces an unconstrained revision; capability memory routes a small audit that reveals current gains and historical regressions. An inadmissible candidate is repaired and re-audited for at most Kproj rounds. The first admissible revision is committed; otherwise, the parent skill is retained.
Method
SpreadsheetBench
SearchQA
BFCL
DocVQA
Average
No skill
78.75
73.75
45.00
90.00
71.88
Static skill
75.00 −3.75
76.25 +2.50
47.50 +2.50
91.25 +1.25
72.50 +0.62
SkillGrad
72.50 −6.25
76.25 +2.50
55.00 +10.00
92.50 +2.50
74.06 +2.18
SkillOpt
73.75 −5.00
77.50 +3.75
52.50 +7.50
88.75 −1.25
73.13 +1.25
SSPE
73.75 −5.00
76.25 +2.50
65.00 +20.00
95.00 +5.00
77.50 +5.62
Table 2: Accuracy (%) on held-out tasks for the no-skill executor and final frozen skills. Colored suffixes show change from the no-skill executor.
Frozen skill
SpreadsheetBench
SearchQA
BFCL
DocVQA
Average
No skill
50.00
61.25
57.50
72.50
60.31
Static skill
38.75 −11.25
61.25 +0.00
57.50 +0.00
75.00 +2.50
58.13 −2.18
SkillGrad
41.25 −8.75
65.00 +3.75
60.00 +2.50
81.25 +8.75
61.88 +1.57
SkillOpt
36.25 −13.75
61.25 +0.00
56.25 −1.25
72.50 +0.00
56.56 −3.75
SSPE
47.50 −2.50
57.50 −3.75
66.25 +8.75
83.75 +11.25
63.75 +3.44
Table 3: Cross-model transfer of GPT-5.4-evolved skills to GPT-4.1 on the same held-out tasks. Values are accuracy (%); colored suffixes show change from the no-skill executor.
Variant
Kproj
SpreadsheetBench
SearchQA
BFCL
DocVQA
Average
SSPE-NoHistory
1
75.00
73.75
50.00
95.00
73.44
SSPE (default)
1
73.75
76.25
65.00
95.00
77.50
SSPE
2
76.25
73.75
50.00
92.50
73.13
SSPE
5
71.25
73.75
46.25
90.00
70.31
Table 4: Held-out accuracy (%) when removing historical evidence or changing semantic projection depth.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Hard outcome
Secondary signal
SpreadsheetBench
task success
cell accuracy
SearchQA
exact answer
—
BFCL
task success
turn-prefix accuracy
DocVQA
exact answer
ANLS
Appendix
Table 5: Benchmarks and evaluation signals in the heterogeneous stream. Hard success is the primary metric in every domain; partial scores are diagnostic.
Method
Update mechanism
Verification access
Persistent optimizer state
Static skill
none
none
none
SkillGrad
diagnosis, semantic momentum, patching
none
semantic momentum
SkillOpt
reflection, bounded editing, validation gate
current bank only
rejected-edit memory
SSPE-NoHistory
free proposal and semantic repair
current bank only
current evidence only
SSPE
risk-directed semantic projection
current and selected historical banks
capability capsules
Appendix
Table 6: Method configurations and information available during evolution.
Kproj
Batches entering k≥2
Later repairs accepted
Total commits
Preq. (%)
1
—
—
4
56.25
2
3
0
3
55.63
5
6
0
3
54.38
Appendix
Table 7: How the projection budget was used during the 20 stream updates. An additional round has index k≥2 , beyond the default first repair.
Kproj
SpreadsheetBench
SearchQA
BFCL
DocVQA
Average
1
48.75
65.00
60.00
87.50
65.31
2
46.25
66.25
53.75
88.75
63.75
5
41.25
61.25
57.50
72.50
58.13
Appendix
Table 8: GPT-4.1 repetition of the projection-budget ablation. Values are held-out accuracy (%) after independently training each lineage with Kproj∈{1,2,5} .
Quantity
Values
Worlds per configuration
100
Latent capabilities K
2,4,8,16
Batches per capability m
2,5,10
Tasks per batch b
8
Interference λ
0,0.10,0.22
Positive-transfer scale β
0.025
Appendix
Table 9: Controlled synthetic sweep. The main factorial uses K=4 and m=5 ; K and m are additionally varied through one-factor scaling slices.
Policy
Update rule
Static skill
Retain qt at every step.
SkillGrad
Commit every proposal; perform no current-validation or historical-capability check. This is the primary baseline.
SkillOpt
Apply a current-capability improvement gate but perform no historical check. Because generated proposals improve the current capability, this policy is usually close to SkillGrad.
Random replay
Audit one uniformly sampled previous capability. Attempt repair if the audit discovers interference; reject the proposal if checked interference remains.
Summary only
Use the noisy scope prediction to identify interference and attempt repair, without empirical historical audits.
SSPE
Use the true anonymous capability context, one risk-directed audit, and one random predicted-safe audit when available; repair measured conflicts and otherwise reject unsafe proposals.
Appendix
Table 10: Policies in the controlled simulator. SkillGrad and SkillOpt denote mechanism-level abstractions of their update decisions, not executions of the corresponding released methods.
Method
Preq.
Final
Forget.
Harm
Static skill
0.475
0.475
0.000
0.000
SkillGrad
0.520
0.462
0.387
0.514
SkillOpt
0.520
0.464
0.386
0.511
Random replay
0.553
0.570
0.271
0.355
Summary only
0.595
0.673
0.206
0.298
SSPE
0.584
0.681
0.171
0.248
Appendix
Table 11: Mean results over the complete 11,300-world sweep. Higher is better for cumulative reward and final macro; lower is better for forgetting and harmful updates.
Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifacts or monolithic documents optimized using aggregate validation scores as feedback. However, representing a skill as a monolithic document restricts optimization to its textual content, without explicitly modeling the structure through which procedural knowledge is retrieved and executed. We identify a key distinction between learning what knowledge to retain and determining how to organize it: textual updates should first be validated through execution evidence, after which the retained knowledge should be structured according to its procedural dependencies and retrieval requirements. To this end, we introduce SkillSpec, a two-phase framework comprising consensus-gated evolution and representation specialization. In the consensus-gated phase, complementary editing intents generate complete candidate skills. An update is committed only when paired evaluations reach consensus, requiring sufficient overall improvement and non-negative aggregate paired gain in every repeated evaluation. In the specialization phase, signals of process and redundancy sensitivity derived from the full optimization trajectory, including accepted and rejected candidates, guide the selection of a flat, graph, or hybrid representation.Across six benchmarks and three target language models, SkillSpec improves average success rate over SkillOpt by 6.89%, averaged across the three models. These results demonstrate that reliable skill evolution and representation specialization address complementary objectives: deciding what knowledge to retain and how to structure it for inference.
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill. We introduce SkillFocus, which decomposes recurring task requirements into a capability space that remains fixed as the skill evolves, separating what tasks require from how the current skill behaves. SkillFocus maps current task outcomes to this space to identify the capability that leaves the most tasks unresolved, then uses that capability to determine what to revise and which evidence to use. Across four benchmarks spanning heterogeneous tasks, SkillFocus achieves the best held-out accuracy on all four, outperforming the strongest competing result by 5.7 points on average while using 24% fewer evolution tokens on average than the closest iterative baseline. Controlled studies further show that capabilities derived from recurring task requirements outperform task-semantic and execution-derived alternatives, while randomizing task--capability assignments reduces final accuracy by up to 20.2 points. Matching evidence to the selected capability increases candidate gain by 4.4 points under prioritized revision.
Ning Wang, Zhiren Gong, Bingdong Li +2
East China Normal University · Nanyang Technological University · Southern University of Science and Technology +1
Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify where productive problem solving begins to break down, rather than reflect coarsely over the entire failure. Based on this insight, we propose SkillPivot, a deviation-point-guided framework for skill self-evolution. SkillPivot detects the transition from a useful prefix to an erroneous suffix using execution validity, goal progress, and action diversity. A stronger teacher then continues from the same prefix and produces a successful alternative under the same interaction history. By contrasting the student's failed suffix with the teacher's successful suffix, SkillPivot generates localized skill updates while preserving already effective guidance. Experiments on ToolQA, LogicBench, and WildClawBench show that SkillPivot consistently outperforms competing skill-evolution methods, improves multiple agent models, and produces compact, transferable skill updates.
Yichun Feng, Jiawei Wang, Haozhe Sun
University of Chinese Academy of Sciences · University of Science and Technology of China · Meituan