Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-space self-evolution, where each recurring task type maintains a persistent tool space which is constructed from accumulated output experience. We identify three limitations of existing methods: (1) output-unaware selection: they rely primarily on tool descriptions or model priors rather than observed tool outputs; (2) statelessness across request: they select tools independently for each request without consolidating prior output experience into persistent task-specific state; (3) inference cost: they repeatedly search, rank, or reason over candidate tools for subsequent requests of the same task. We address these limitations through output-aware tool scoring, persistent task-specific tool spaces, amortized tool selection, and reusable configurations across models. We introduce LOTS (Likelihood-Only Tool Scoring), which evolves an agent's tool space from accumulated output experience while keeping model parameters fixed. After each request, LOTS holds the model's generated answer and estimates each tool's contribution by measuring how much the answer likelihood changes when its observed output is removed. These contributions are aggregated within each recurring task to rank tools and update its persistent space. Across three benchmarks, LOTS improves task performance while substantially reducing tool context. More importantly, sequential experiments demonstrate that task-specific spaces persist and continue to improve over time, while cross-model experiments show that learned configurations transfer across different models.
Figures & tables
Figure 1: A task-specific tool space that evolves from the agent’s own outputs. (a) A fixed model is served full registry on every request of a task. (b) Likelihood-Only Tool Scoring ( LOTS ) turns the output experience of a task’s earlier requests into a compact space Et . Later requests of the task reuse it, and their outputs update it. (c) Accuracy as the registry grows from 30 to 240 tools (TGB, Qwen3.5-9B): serving every tool falls to 34.5%, while the LOTS space stays at 53.8% or above (Appendix F.1 ).
Figure 2: LOTS evolves persistent task-specific tool spaces from output traces. (a) Requests are assigned to recurring task types, each of which maintains its own space Ec,k . Requests of task c collected for an update are served under the current space. (b) One LOTS update freezes each generated output, scores tools by leave-one-out likelihood, aggregates the scores within the task, and uses the ranking to prune the task’s tool space.
Algorithm 1: Task-specific tool-space construction
Input: Task ct , request batch Qt , candidate space Etc⊆S , budget K , fixed model pθ .
1. Serve. Execute Qt under the fixed space Etc and record traces (hx,yx) .
2. Score. For each request x and exposed, scoreable tool s , keep yx fixed and compute Ix(s)=Lθ(yx∣hx)−Lθ(yx∣hx−s) . Here hx−s removes s ’s calls and observations; where a later tool consumes s ’s output, the chain is re-executed without s (Appendix B.2 ).
3. Aggregate. Let Qt(s) contain the requests providing scores for s , and At={s∈Etc:∣Qt(s)∣>0} . Compute Iˉt(s)=∣Qt(s)∣1∑x∈Qt(s)Ix(s) .
4. Prune. Set Et=TopKs∈AtIˉt(s) , retaining all eligible tools if ∣At∣<K .
5. Deploy. Store Et for subsequent task- ct requests; repeat with new evidence when updating.
Table 3
Qwen2.5-7B
Qwen2.5-14B
Llama3.1-8B
Mistral-7B
Qwen3.5-9B
Qwen3-8B
Phi4-14B
Gemma4-12B
No tools
10.8
11.4
12.9
10.0
10.5
10.4
11.0
12.8
Keep all
69.6
88.0
77.7
46.4
50.9
56.5
79.7
86.7
Random
23.4
19.3
41.3
11.5
27.5
10.1
24.3
20.0
Search-Based
BM25
19.3
36.3
40.0
13.4
24.9
10.0
39.3
37.5
Dense
20.5
36.3
37.6
17.2
18.5
18.5
11.1
38.8
Table 1: TGB dataset accuracy (%). One tool space fitted per task family. Search-based baselines include BM25, Dense, and Tool2Vec; the LLM-based router selects tools per request. † marks methods that read correctness labels on the fitting requests. Bold marks the best result in each column.
Menu
Qwen3-4B
Qwen3.5-4B
Qwen3.5-9B
Gemma4-12B
Phi-4-mini-3.8B
No tools
0.00
0.00
0.00
0.00
0.00
Keep all
17.71±0.95
14.37±0.63
23.12±0.63
24.79±0.36
3.33±0.36
Random
10.21±1.06
2.71±1.06
13.96±0.29
13.54±0.29
3.54±0.29
Search-Based
BM25
18.13±0.51
12.08±0.29
27.08±0.78
25.83±1.93
3.33±0.29
Dense
25.62±1.77
13.96±0.29
27.71±0.29
26.04±0.29
3.75±0.00
Table 2: BFCL v4 multi_turn accuracy. For each task, the first 10 records fit the space and its remaining 40 records evaluate it (160 records in total). † marks methods that read correctness labels for all fitting requests. Bold marks the best result in each column; underlining marks the second best. Results are averaged over three seeds.
Figure 3: Accumulating task-specific spaces as tasks arrive. (a) Seven of the eight TGB families in this stream arrive on one request stream (Qwen2.5-7B, registry of 120 tools); a step is a batch of four requests, and dotted lines mark arrivals. The curve is credited accuracy Acct , it rises in a jump when a family’s space is added. (b) Final accuracy and space size for each family. (c) Four BFCL API classes arrive in sequence (Qwen3-4B). Bars give each class’s accuracy on its own evaluation records; lines give Acct over the four classes, which ends at the mean of the bars.
Method
Qwen2.5-3B
Qwen2.5-7B
Qwen3-8B
Qwen2.5-14B
Qwen3.5-9B
Phi-4-14B
G4-12B
No tools
1.79
2.68
4.46
4.46
6.25
4.46
2.68
Keep all
7.14
7.14
10.71
8.04
40.18
11.61
24.11
Random
6.25
5.36
11.61
6.25
16.96
8.93
14.29
Search-Based
BM25
1.79
2.68
6.25
3.57
14.29
2.68
14.29
Dense
0.00
1.79
1.79
1.79
12.50
0.89
5.36
Table 3: GTA-Atomic accuracy (%) . † marks methods that read correctness labels on the fitting requests. Bold denotes the best result in each column; underlining denotes the second best.
Figure 4: Improving an existing task space through successive request batches. TGB, 120-tool registry. Left: four family updates showing accuracy on the task’s fixed evaluation requests. Right: mean accuracy after each update over all 24 runs per model.
Figure 9
Figure 7: Accuracy against serving cost on GTA-Atomic read_arith . Prompt tokens per test task over all ReAct turns against accuracy. The arrow runs from Keep all to LOTS .
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
execution model
from Qwen3-4B
from Qwen3.5-9B
from Gemma4-12B
Keep all
Qwen3-4B
21.5
18.1
19.4
18.1
Qwen3.5-9B
27.7
26.5
29.4
23.3
Gemma4-12B
26.9
29.1
27.5
25.0
Appendix
Table 6: Cross-model reuse on BFCL with independent fitting and evaluation records: accuracy (%) on the 160 evaluation records at a fixed K=16 , mean over three seeds (Gemma4-12B as execution model: two seeds). Rows are execution models; columns are the source of the ranking.
Figure 8: BFCL documentation-budget sensitivity on the 101 tune records: accuracy change over Keep all, in percentage points, as the number of fully documented tools K varies. Each panel is one deployed model; colors give the source of the ranking. Hollow markers mark K=16 ; at K=31 all documentation is kept and only the menu order changes. Results average three seeds, two for Qwen3-8B.
Figure 9: TGB accuracy against the budget K , every ranking deployed at the same K per family. Macro average over the ten families; Random is the mean over five subsets per family; the dotted line is Keep all.
Baseline
K=2
K=4
Qwen2.5-7B
Tool2Vec
+46.0 [ +44.9 , +47.1 ]
+35.9 [ +34.6 , +37.3 ]
Dense
+55.5 [ +54.3 , +56.8 ]
+60.9 [ +59.5 , +62.4 ]
Random
+55.7 [ +54.5 , +57.0 ]
+56.3 [ +55.0 , +57.7 ]
Qwen3.5-9B
Tool2Vec
+37.2 [ +36.2 , +38.2 ]
+31.9 [ +30.6 , +33.2 ]
Appendix
Table 9: Paired accuracy difference, LOTS minus the baseline, at each fixed budget, with a 95% interval from a family-stratified paired bootstrap over evaluation requests (2,000 resamples). Positive values favor LOTS .
Router
Clusters
Purity
NMI
ARI
Qwen2.5-7B
Qwen3.5-9B
R0: family label
10
1.000
1.000
1.000
82.84
65.53
R1: k -means, k=10 , E5
10
1.000
1.000
1.000
82.84±0.00
65.53±0.00
R1: k -means, k=10 , BGE
10
1.000
1.000
1.000
82.84±0.00
65.53±0.00
R2: silhouette k , E5 ( k=9 )
9
0.900
0.969
0.898
82.38
65.19
R2: silhouette k , BGE ( k=10 )
10
1.000
1.000
1.000
82.84
65.53
R1: k -means, k=5 , E5
5
0.500
0.758
0.428
81.53±0.71
62.92±0.78
Appendix
Table 10: Routing by request clustering on TGB. Clustering quality is measured on the 3,200 evaluation requests against the ten family labels; accuracy (%) is the macro average over families under the per-cluster fitted budget K , mean ± standard deviation over three clustering seeds. Deterministic routers have one seed.
Figure 10: Tool space, tool use, and performance for one task per benchmark. Tint marks the tools in the space; on BFCL, where every function stays callable, it marks full documentation, and a paler tint for Random gives the share of records in which the tool is documented. Circle area is the share of evaluation requests that call the tool. TGB executes every served tool once per request, so its panel shows the configuration only. Bars give evaluation accuracy; white dots are the three seeds on BFCL.
Figure 11: Evolution of one persistent space: TGB fx_settle , Qwen2.5-7B, 120-tool registry with a cap of 16 tools per request, batches of four requests, ε=0.5 , α=0.3 , K=3 . Top: running score of the ten highest-scoring tools after each batch; outlines mark the tools kept in the next space; bold labels mark the task’s chain. Bottom: accuracy of the updated space on the evaluation requests, and the number of tools and chain tools it keeps.
Figure 12: First tool call and final outcome on GTA-Atomic web_fact with the search tool enabled. Each band links the first tool the agent calls on an evaluation request to whether the request is answered correctly; band width is the number of requests. The figure describes association and omits the later calls of a request.
Model
Families
Seed 0
Seed 1
Seed 2
Prompt tokens
Qwen2.5-7B
8
0.82 / 0.52
0.82 / 0.43
0.79 / 0.46
641 / 1832
Qwen2.5-14B
8
0.87 / 0.76
0.86 / 0.80
0.85 / 0.82
623 / 1820
Qwen3-8B
8
0.42 / 0.21
0.53 / 0.23
0.37 / 0.20
629 / 1834
Qwen3.5-9B
8
0.62 / 0.30
0.59 / 0.32
0.61 / 0.29
659 / 1860
phi-4
8
0.77 / 0.48
0.77 / 0.53
0.74 / 0.49
531 / 1537
Mistral-7B
8
0.62 / 0.20
0.50 / 0.30
0.64 / 0.30
706 / 2026
Appendix
Table 12: Task-arrival streams on seven models and three seeds. Accuracy is credited over arrived families at the end of the stream. Each entry is LOTS / Keep all. Prompt tokens are averaged over served requests, including exploration and excluding update scoring.
Figure 13: From request scores to a task ranking and tool space, Qwen2.5-7B. (a) BFCL TravelAPI , using documentation demotion. (b) GTA-Atomic read_arith , using pruning, fitted on 33 requests; the accuracies are those of Table 8 on the remaining 32.
Figure 14: Task-specific tool spaces on TGB, Qwen2.5-7B, under the main protocol: each row is fitted on a family’s first 80 requests and keeps its top- K tools. Outlined cells mark the family’s gold chain; the right margin reports K and accuracy against Keep all on the remaining 320 requests.
Figure 15: Stationary-task update trajectories, Qwen2.5-7B: 15-tool registry, eight batches of 40 requests. Top: tools retained by the updating arm after each round; outlines mark entries. Bottom: accuracy on each family’s fixed evaluation set for the updating, frozen and Keep-all configurations.
Update rule
Runs
Mean acc.
Final acc.
Size
Gold
LOTS update (top- K )
24
0.162
0.265
4.0
0.41
Random, same K
24
0.010
0.006
4.0
0.00
Union of exposed tools
24
0.023
0.023
52.7
0.29
Appendix
Table 13: Update rules under a serving cap, Qwen2.5-7B: eight families, 120-tool registry, 30 rounds of four requests, ε=0.5 , three seeds; each rule is a separate run with the same request order, exploration probability and seeds. An exploring request receives the current space plus at most 16−∣S∣ further tools, and a space larger than 16 tools is served in full. Accuracy is read on each family’s fixed evaluation set under the current space, averaged over the rounds and at the final round; size is the final space; gold is the share of the family’s required tools in the final space.
Variant
Tools kept
Accuracy
Fixed own output, LOO, τ⋆
3.7
0.724
Fixed own output, LOO, τ=0
10.4
0.698
Per-candidate generation ( Lself )
—
0.648
Gold-answer scoring
12.2
0.796
Appendix
Table 14: Request-level scoring variants on TGB, Qwen2.5-7B. Gold-answer scoring uses labels; per-candidate generation scores a different string under each intervention. These configurations differ in retained size.
Figure 16: Tool scores for one TGB sensor_convert request, Qwen2.5-7B. The full-menu answer is fixed while each tool is removed using the TGB intervention, including downstream propagation where applicable. The two gold-chain tools carry the score.
Qwen2.5-7B
Qwen2.5-14B
Mistral-7B
Qwen3.5-9B
Qwen3-8B
Keep all
0.696
0.880
0.464
0.509
0.565
Trace judge
0.623
0.906
0.558
0.613
0.689
LOTS
0.829
0.918
0.658
0.652
0.815
Appendix
Table 17: Trace-judge ablation on TGB: the space is built from LLM judgments of the frozen traces instead of likelihood scores, with the same task-level aggregation and budget K .
Model
Top- K
Random
Keep all
Qwen2.5-7B
0.829
0.234
0.696
Qwen2.5-14B
0.918
0.193
0.880
Llama3.1-8B
0.848
0.413
0.777
Mistral-7B
0.658
0.115
0.464
Qwen3.5-9B
0.652
0.275
0.509
Qwen3-8B
0.815
0.101
0.565
Appendix
Table 19: Likelihood ranking and execution utility on TGB. Accuracy is averaged across ten task families, using 320 evaluation requests per family. The three configurations are those of Table 1 : Top- K and Random use identical family-specific budgets, and Random uses a single fixed-seed draw. Bold denotes the best configuration for each model.
Figure 17: Scoring-model diagnostic on TGB: accuracy under Keep all and under the spaces fitted with an instruct-model or a base-model scorer, from frozen instruct-model outputs and a fixed support-frequency threshold. The instruct model serves every configuration; the right margin gives the number of tools each scorer keeps.
Figure 18: Incorrect-output diagnostic: 3,000 fitting traces per model and 400 separate evaluation requests. Triangles give label-filtered subsets (correct only, wrong only); lines replace a growing share of answers at random (three draws). Top: accuracy; bottom: tools kept. Support-frequency retention uses a 3% threshold for Llama-3.1-8B and 5% otherwise; mean-score retention uses τ=0.1 .
Figure 19: Trace-count sensitivity, TGB and Qwen2.5-7B. Thirty subsamples per size are compared with the configuration fitted from all 3,000 traces: rank correlation, Jaccard overlap of the space, and the fraction of subsamples with an identical space (left); mean space size (right).
Figure 20: GTA-Atomic read_arith token amortization. LOTS starts with likelihood-scoring input tokens for the 33 fitting requests and accumulates serving input tokens; Keep all starts at zero and accumulates its own serving cost. The plotted account excludes any additional cost of collecting fitting traces.
Figure 21: TGB registry-size diagnostic with one pooled configuration and two rounds of 400 requests. Top: accuracy under Keep all and under the fitted space. Middle: serving prompt tokens per request, with the number of retained tools above each point. Bottom: scoring time on one L40S for round 0 (full registry) and round 1 (fitted space); this is scoring time, not generation latency. The hollow marker uses a different batch cap. Qwen2.5-7B and Qwen2.5-14B average three seeds at 30 to 120 tools and Qwen3.5-9B averages four seeds for accuracy and tokens; the other points, and all timings, use one seed.
College of Computer Science and Technology, Zhejiang University · ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University · ZJU-UIUC Institute, Zhejiang University