Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{$4.03} in search cost, it produces a harness that solves \textbf{82.0}% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.
Figures & tables
Figure 1: SelfSearch. An agent modifies an editable copy of itself using previous episode records, which contain interaction trajectories, code changes, and check results. The successor becomes the next actor, while its episode record informs subsequent revisions. Stacked boxes represent parallel lineages. Lineage superscripts are omitted.
GPT-5.6 family
DeepSeek-V4 family
Initial B0
Bc
Ba
Initial B0
Bc
Ba
SWE-bench Verified
Success (%) ↑
75.0
77.5
75.0
81.7
86.7
86.7
Cost/task ()\downarrow$
0.0096
0.0091
0.0093
0.0214
0.0169
0.0184
SWE-bench Multilingual
Success (%) ↑
53.3
60.0
58.3
68.3
66.7
73.3
Table 1: Downstream success (%) and average execution cost per task (USD) for the initial agent and the capability ( Bc ) and adaptive ( Ba ) agents after ten generations. SWE-bench Verified uses the 120-task evaluation set. Costs exclude search. Bold indicates the best value within each model and metric, including ties.
Figure 2: Accuracy–cost frontier on Terminal-Bench 2.1. Dashed lines connect agents on the accuracy–cost frontier.
GPT-5.6 family
DeepSeek-V4 family
Method
SWE-bench Verified
SWE-bench Multilingual
Search cost
SWE-bench Verified
SWE-bench Multilingual
Search cost
Initial B0
76.4
53.3
—
81.8
68.3
—
Linear search
77.3
60.0
12.35
81.8
73.3
8.59
Archive search
76.4
56.7
7.53
86.4
70.0
7.90
SelfSearch Bc
78.2
60.0
6.52
87.3
66.7
4.03
SelfSearch Ba
75.5
58.3
87.3
73.3
Table 2: Comparison with evaluation-guided search. Success (%) and search cost (USD) on SWE-bench Verified and SWE-bench Multilingual. The SWE-bench Verified comparison excludes the ten development tasks, leaving 110 tasks. Bold marks the best success rate in each setting.
GPT-5.6 Sol
DeepSeek V4 Pro
Model
Initial B0
Bc
Ba
Bc
Ba
DeepSeek V4 Flash
81.7
83.3
85.0
86.7
86.7
GPT-5.6 Luna
75.0
77.5
75.0
79.2
76.7
Table 3: Cross-model transfer on SWE-bench Verified (success %). Rows specify execution models and column groups specify search models. Superscripts c and a denote capability and adaptive lineages. Bold marks the best rate within each search model per row, including ties.
GPT-5.6 family
DeepSeek-V4 family
Method
c
a
Mean
c
a
Mean
Initial agent
75.0
75.0
75.0
81.7
81.7
81.7
w/o episode records
75.0
73.3
74.2
82.5
85.0
83.8
w/ fixed improver
73.3
75.8
74.6
82.5
85.0
83.8
SelfSearch (full)
77.5
75.0
76.2
86.7
86.7
86.7
Table 4: Mechanism ablations on SWE-bench Verified (success %). Columns c and a denote capability and adaptive lineages. Mean averages their rates. Search conditions use generation-10 agents, while Initial reports B0 . Compare within model settings.
Gen.
Capability lineage
Adaptive lineage
1
Exact-text replacement
Plan revision and failure recovery
2
Line-range file viewing
Exact-text replacement
3
Plan revision and failure recovery
Line-range file viewing
4
Character-range file viewing
General instructions loaded at startup
5
General instructions loaded at startup
Character-range file viewing
6
Text search with output limits
Structured trajectory reader
Table 5: Changes introduced across ten generations of SelfSearch with GPT. Repeated entries may reflect changes incorporated from the other lineage. Bold highlights changes discussed in the analysis.
Figure 3: A clipped episode record motivates a text-search tool, which is later reused to inspect downstream task code.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Task IDs
astropy__astropy-7166
django__django-11333
django__django-13964
astropy__astropy-7606
django__django-11477
django__django-14011
astropy__astropy-8707
django__django-11551
django__django-14017
astropy__astropy-8872
django__django-11555
django__django-14089
astropy__astropy-12907
django__django-11740
django__django-14140
astropy__astropy-13236
django__django-12741
django__django-14238
Appendix
Table 6: The 60 additional SWE-bench Verified task IDs sampled from DGM’s large subset.
Task IDs
apache__druid-14092
laravel__framework-53914
apache__lucene-13494
laravel__framework-53949
apache__lucene-13704
nushell__nushell-13605
astral-sh__ruff-15356
php-cs-fixer__php-cs-fixer-7635
astral-sh__ruff-15443
phpoffice__phpspreadsheet-3570
axios__axios-4731
phpoffice__phpspreadsheet-4114
Appendix
Table 7: The 60 sampled SWE-bench Multilingual task IDs.
Model
Calls
Input (M)
Output (k)
Total cost ($)
GPT-5.6 Sol
281
5.48
101.4
6.52
DeepSeek V4 Pro
832
36.89
578.7
4.03
Appendix
Table 8: Search resources. Input tokens include cached input. M and k denote millions and thousands of tokens. Costs are in USD and exclude downstream evaluation.
Figure 4: Accuracy–cost tradeoffs on SWE-bench Verified and SWE-bench Multilingual. Costs and markers follow Figure 2 . Dashed lines connect non-dominated evaluated agents in each subplot. Axis ranges differ between benchmarks.
GPT-5.6 Luna
DeepSeek V4 Flash
Benchmark
Agent
Calls
Input (k)
Output (k)
Calls
Input (k)
Output (k)
SWE-bench Verified
B0
14.3
139.7
2.9
44.1
1518.8
20.5
Bc
14.1
154.9
2.7
33.2
1150.5
18.1
Ba
14.3
163.1
2.7
35.8
1312.0
19.1
SWE-bench Multilingual
B0
17.1
218.8
3.9
56.3
2451.9
25.2
Bc
18.3
263.7
3.9
41.3
1274.1
19.1
Appendix
Table 9: Mean execution resources per task, including unsuccessful tasks. Tokens are in thousands, and input includes cached tokens.
Capability ( Bc )
Adaptive ( Ba )
Model
Benchmark
Shared tasks
Reduction (%)
Shared tasks
Reduction (%)
GPT-5.6 Luna
SWE-bench Verified
86
8.7
84
13.6
SWE-bench Multilingual
31
4.1
32
8.3
Terminal-Bench 2.1
33
-2.4
31
15.2
DeepSeek V4 Flash
SWE-bench Verified
95
36.1
95
19.7
SWE-bench Multilingual
37
52.9
39
38.5
Appendix
Table 10: Execution cost reductions relative to B0 on tasks solved by both agents. Negative values indicate increased cost.
Gen.
Capability lineage
Adaptive lineage
1
Search, file viewing, and exact replacement
Search, viewing, replacement, and guidance
2
Verification and recovery guidance
Shell exit status and output excerpts
3
Shell exit status and output excerpts
Unique shell capture files and guidance
4
Unique shell capture files
File views with head and tail excerpts
5
File views with head and tail excerpts
Search output limits
6
Search excerpts centered on matches
Directory listing limits and filters
Appendix
Table 11: Changes introduced across ten generations of SelfSearch with DeepSeek. Repeated entries may reflect changes incorporated from the other lineage. Bold highlights changes discussed in the analysis.
Agent
Operation
SWE-bench Verified
SWE-bench Multilingual
Terminal- Bench 2.1
All
GPT Bc
Text search
9.2
11.7
3.4
7.8
Line-range viewing
48.3
45.0
21.3
38.7
Character-range viewing
0.0
1.7
0.0
0.4
Trajectory reader
0.0
0.0
0.0
0.0
Exact-text replacement
98.3
98.3
32.6
76.6
GPT Ba
Text search
13.3
6.7
5.6
9.3
Appendix
Table 12: Downstream use of introduced operations by the final capability and adaptive agents. Entries give the percentage of tasks with at least one invocation. Character-range viewing and the trajectory reader were not introduced in the DeepSeek agents and are omitted for those agents.
Method
Evolving task agent
Evolving meta-agent
Unified meta and task agent
Reward-free search
DGM ( Zhang et al., 2026a )
✓
—
—
—
HGM ( Wang et al., 2026 )
✓
—
—
—
Hyperagents ( Zhang et al., 2026b )
✓
✓
—
—
SICA ( Robeyns et al., 2025 )
✓
✓
✓
—
SelfSearch (ours)
✓
✓
✓
✓
Appendix
Table 13: Comparison of agent search designs. A checkmark denotes the property as defined below, and a dash denotes its absence.