Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improvements for new tasks can overwrite procedures needed for earlier ones. In continual learning, Orthogonal Gradient Descent (OGD) addresses analogous interference by projecting a new-task gradient onto a subspace that locally preserves prior predictions. Natural-language skill revisions, however, have neither gradients nor a canonical vector space in which such a projection can be performed. We introduce \emph{Semantic-Scope Projected Evolution} (SSPE), which transfers the functional principle of gradient projection from parameter space to behavior space. SSPE treats an unconstrained skill revision as a proposed update, identifies acquired capabilities with which it may interfere, and uses the observed gains and regressions to construct a compatible revision rather than merely rejecting the update. This enables one shared skill to evolve across latent and recurring task contexts without exposing semantic domain identities to the evolution model. Across controlled synthetic streams and heterogeneous real-agent benchmarks, SSPE improves final cross-domain competence and mitigates forgetting relative to strong skill-evolution baselines. The evolved skill also retains the strongest average performance after transfer to a different executor model. These results establish semantic projection as a promising principle for stable and adaptive self evolution of language agents.
Figures & tables
OGD
SSPE
Current gradient gt
Unconstrained revision St
Protected output sensitivities
Capability capsules and audit tasks
Alignment with a protected direction
Measured historical regression
Orthogonal projection
Revision conditioned on observed conflicts
Zero projected step
Explicit no update
Table 1: SSPE transfers the operational invariant of OGD, not its vector operations.
Figure 1: SSPE expands one update in a heterogeneous stream. Current evidence produces an unconstrained revision; capability memory routes a small audit that reveals current gains and historical regressions. An inadmissible candidate is repaired and re-audited for at most Kproj rounds. The first admissible revision is committed; otherwise, the parent skill is retained.
Method
SpreadsheetBench
SearchQA
BFCL
DocVQA
Average
No skill
78.75
73.75
45.00
90.00
71.88
Static skill
75.00 −3.75
76.25 +2.50
47.50 +2.50
91.25 +1.25
72.50 +0.62
SkillGrad
72.50 −6.25
76.25 +2.50
55.00 +10.00
92.50 +2.50
74.06 +2.18
SkillOpt
73.75 −5.00
77.50 +3.75
52.50 +7.50
88.75 −1.25
73.13 +1.25
SSPE
73.75 −5.00
76.25 +2.50
65.00 +20.00
95.00 +5.00
77.50 +5.62
Table 2: Accuracy (%) on held-out tasks for the no-skill executor and final frozen skills. Colored suffixes show change from the no-skill executor.
Frozen skill
SpreadsheetBench
SearchQA
BFCL
DocVQA
Average
No skill
50.00
61.25
57.50
72.50
60.31
Static skill
38.75 −11.25
61.25 +0.00
57.50 +0.00
75.00 +2.50
58.13 −2.18
SkillGrad
41.25 −8.75
65.00 +3.75
60.00 +2.50
81.25 +8.75
61.88 +1.57
SkillOpt
36.25 −13.75
61.25 +0.00
56.25 −1.25
72.50 +0.00
56.56 −3.75
SSPE
47.50 −2.50
57.50 −3.75
66.25 +8.75
83.75 +11.25
63.75 +3.44
Table 3: Cross-model transfer of GPT-5.4-evolved skills to GPT-4.1 on the same held-out tasks. Values are accuracy (%); colored suffixes show change from the no-skill executor.
Variant
Kproj
SpreadsheetBench
SearchQA
BFCL
DocVQA
Average
SSPE-NoHistory
1
75.00
73.75
50.00
95.00
73.44
SSPE (default)
1
73.75
76.25
65.00
95.00
77.50
SSPE
2
76.25
73.75
50.00
92.50
73.13
SSPE
5
71.25
73.75
46.25
90.00
70.31
Table 4: Held-out accuracy (%) when removing historical evidence or changing semantic projection depth.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Hard outcome
Secondary signal
SpreadsheetBench
task success
cell accuracy
SearchQA
exact answer
—
BFCL
task success
turn-prefix accuracy
DocVQA
exact answer
ANLS
Appendix
Table 5: Benchmarks and evaluation signals in the heterogeneous stream. Hard success is the primary metric in every domain; partial scores are diagnostic.
Method
Update mechanism
Verification access
Persistent optimizer state
Static skill
none
none
none
SkillGrad
diagnosis, semantic momentum, patching
none
semantic momentum
SkillOpt
reflection, bounded editing, validation gate
current bank only
rejected-edit memory
SSPE-NoHistory
free proposal and semantic repair
current bank only
current evidence only
SSPE
risk-directed semantic projection
current and selected historical banks
capability capsules
Appendix
Table 6: Method configurations and information available during evolution.
Kproj
Batches entering k≥2
Later repairs accepted
Total commits
Preq. (%)
1
—
—
4
56.25
2
3
0
3
55.63
5
6
0
3
54.38
Appendix
Table 7: How the projection budget was used during the 20 stream updates. An additional round has index k≥2 , beyond the default first repair.
Kproj
SpreadsheetBench
SearchQA
BFCL
DocVQA
Average
1
48.75
65.00
60.00
87.50
65.31
2
46.25
66.25
53.75
88.75
63.75
5
41.25
61.25
57.50
72.50
58.13
Appendix
Table 8: GPT-4.1 repetition of the projection-budget ablation. Values are held-out accuracy (%) after independently training each lineage with Kproj∈{1,2,5} .
Quantity
Values
Worlds per configuration
100
Latent capabilities K
2,4,8,16
Batches per capability m
2,5,10
Tasks per batch b
8
Interference λ
0,0.10,0.22
Positive-transfer scale β
0.025
Appendix
Table 9: Controlled synthetic sweep. The main factorial uses K=4 and m=5 ; K and m are additionally varied through one-factor scaling slices.
Policy
Update rule
Static skill
Retain qt at every step.
SkillGrad
Commit every proposal; perform no current-validation or historical-capability check. This is the primary baseline.
SkillOpt
Apply a current-capability improvement gate but perform no historical check. Because generated proposals improve the current capability, this policy is usually close to SkillGrad.
Random replay
Audit one uniformly sampled previous capability. Attempt repair if the audit discovers interference; reject the proposal if checked interference remains.
Summary only
Use the noisy scope prediction to identify interference and attempt repair, without empirical historical audits.
SSPE
Use the true anonymous capability context, one risk-directed audit, and one random predicted-safe audit when available; repair measured conflicts and otherwise reject unsafe proposals.
Appendix
Table 10: Policies in the controlled simulator. SkillGrad and SkillOpt denote mechanism-level abstractions of their update decisions, not executions of the corresponding released methods.
Method
Preq.
Final
Forget.
Harm
Static skill
0.475
0.475
0.000
0.000
SkillGrad
0.520
0.462
0.387
0.514
SkillOpt
0.520
0.464
0.386
0.511
Random replay
0.553
0.570
0.271
0.355
Summary only
0.595
0.673
0.206
0.298
SSPE
0.584
0.681
0.171
0.248
Appendix
Table 11: Mean results over the complete 11,300-world sweep. Higher is better for cumulative reward and final macro; lower is better for forgetting and harmful updates.