SelfSearch: Reward-Free Search for Self-Improving Agents
Organizations: Graduate School of Data Science, Seoul National University
Abstract
Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{$4.03} in search cost, it produces a harness that solves \textbf{82.0}% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.
Figures & tables
| GPT-5.6 family | DeepSeek-V4 family | |||||
| Initial | Initial | |||||
| SWE-bench Verified | ||||||
| Success (%) | 75.0 | 77.5 | 75.0 | 81.7 | 86.7 | 86.7 |
| Cost/task (\downarrow$ | 0.0096 | 0.0091 | 0.0093 | 0.0214 | 0.0169 | 0.0184 |
| SWE-bench Multilingual | ||||||
| Success (%) | 53.3 | 60.0 | 58.3 | 68.3 | 66.7 | 73.3 |
| GPT-5.6 family | DeepSeek-V4 family | |||||
|---|---|---|---|---|---|---|
| Method | SWE-bench Verified | SWE-bench Multilingual | Search cost | SWE-bench Verified | SWE-bench Multilingual | Search cost |
| Initial | 76.4 | 53.3 | — | 81.8 | 68.3 | — |
| Linear search | 77.3 | 60.0 | 12.35 | 81.8 | 73.3 | 8.59 |
| Archive search | 76.4 | 56.7 | 7.53 | 86.4 | 70.0 | 7.90 |
| SelfSearch | 78.2 | 60.0 | 6.52 | 87.3 | 66.7 | 4.03 |
| SelfSearch | 75.5 | 58.3 | 87.3 | 73.3 | ||
| GPT-5.6 Sol | DeepSeek V4 Pro | ||||
|---|---|---|---|---|---|
| Model | Initial | ||||
| DeepSeek V4 Flash | 81.7 | 83.3 | 85.0 | 86.7 | 86.7 |
| GPT-5.6 Luna | 75.0 | 77.5 | 75.0 | 79.2 | 76.7 |
| GPT-5.6 family | DeepSeek-V4 family | |||||
|---|---|---|---|---|---|---|
| Method | Mean | Mean | ||||
| Initial agent | 75.0 | 75.0 | 75.0 | 81.7 | 81.7 | 81.7 |
| w/o episode records | 75.0 | 73.3 | 74.2 | 82.5 | 85.0 | 83.8 |
| w/ fixed improver | 73.3 | 75.8 | 74.6 | 82.5 | 85.0 | 83.8 |
| SelfSearch (full) | 77.5 | 75.0 | 76.2 | 86.7 | 86.7 | 86.7 |
| Gen. | Capability lineage | Adaptive lineage |
|---|---|---|
| 1 | Exact-text replacement | Plan revision and failure recovery |
| 2 | Line-range file viewing | Exact-text replacement |
| 3 | Plan revision and failure recovery | Line-range file viewing |
| 4 | Character-range file viewing | General instructions loaded at startup |
| 5 | General instructions loaded at startup | Character-range file viewing |
| 6 | Text search with output limits | Structured trajectory reader |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Task IDs | ||
|---|---|---|
| astropy__astropy-7166 | django__django-11333 | django__django-13964 |
| astropy__astropy-7606 | django__django-11477 | django__django-14011 |
| astropy__astropy-8707 | django__django-11551 | django__django-14017 |
| astropy__astropy-8872 | django__django-11555 | django__django-14089 |
| astropy__astropy-12907 | django__django-11740 | django__django-14140 |
| astropy__astropy-13236 | django__django-12741 | django__django-14238 |
| Task IDs | |
|---|---|
| apache__druid-14092 | laravel__framework-53914 |
| apache__lucene-13494 | laravel__framework-53949 |
| apache__lucene-13704 | nushell__nushell-13605 |
| astral-sh__ruff-15356 | php-cs-fixer__php-cs-fixer-7635 |
| astral-sh__ruff-15443 | phpoffice__phpspreadsheet-3570 |
| axios__axios-4731 | phpoffice__phpspreadsheet-4114 |
| Model | Calls | Input (M) | Output (k) | Total cost ($) |
|---|---|---|---|---|
| GPT-5.6 Sol | 281 | 5.48 | 101.4 | 6.52 |
| DeepSeek V4 Pro | 832 | 36.89 | 578.7 | 4.03 |
| GPT-5.6 Luna | DeepSeek V4 Flash | ||||||
|---|---|---|---|---|---|---|---|
| Benchmark | Agent | Calls | Input (k) | Output (k) | Calls | Input (k) | Output (k) |
| SWE-bench Verified | 14.3 | 139.7 | 2.9 | 44.1 | 1518.8 | 20.5 | |
| 14.1 | 154.9 | 2.7 | 33.2 | 1150.5 | 18.1 | ||
| 14.3 | 163.1 | 2.7 | 35.8 | 1312.0 | 19.1 | ||
| SWE-bench Multilingual | 17.1 | 218.8 | 3.9 | 56.3 | 2451.9 | 25.2 | |
| 18.3 | 263.7 | 3.9 | 41.3 | 1274.1 | 19.1 | ||
| Capability ( ) | Adaptive ( ) | ||||
|---|---|---|---|---|---|
| Model | Benchmark | Shared tasks | Reduction (%) | Shared tasks | Reduction (%) |
| GPT-5.6 Luna | SWE-bench Verified | 86 | 8.7 | 84 | 13.6 |
| SWE-bench Multilingual | 31 | 4.1 | 32 | 8.3 | |
| Terminal-Bench 2.1 | 33 | -2.4 | 31 | 15.2 | |
| DeepSeek V4 Flash | SWE-bench Verified | 95 | 36.1 | 95 | 19.7 |
| SWE-bench Multilingual | 37 | 52.9 | 39 | 38.5 | |
| Gen. | Capability lineage | Adaptive lineage |
|---|---|---|
| 1 | Search, file viewing, and exact replacement | Search, viewing, replacement, and guidance |
| 2 | Verification and recovery guidance | Shell exit status and output excerpts |
| 3 | Shell exit status and output excerpts | Unique shell capture files and guidance |
| 4 | Unique shell capture files | File views with head and tail excerpts |
| 5 | File views with head and tail excerpts | Search output limits |
| 6 | Search excerpts centered on matches | Directory listing limits and filters |
| Agent | Operation | SWE-bench Verified | SWE-bench Multilingual | Terminal- Bench 2.1 | All |
|---|---|---|---|---|---|
| GPT | Text search | 9.2 | 11.7 | 3.4 | 7.8 |
| Line-range viewing | 48.3 | 45.0 | 21.3 | 38.7 | |
| Character-range viewing | 0.0 | 1.7 | 0.0 | 0.4 | |
| Trajectory reader | 0.0 | 0.0 | 0.0 | 0.0 | |
| Exact-text replacement | 98.3 | 98.3 | 32.6 | 76.6 | |
| GPT | Text search | 13.3 | 6.7 | 5.6 | 9.3 |
| Method | Evolving task agent | Evolving meta-agent | Unified meta and task agent | Reward-free search |
|---|---|---|---|---|
| DGM ( Zhang et al., 2026a ) | — | — | — | |
| HGM ( Wang et al., 2026 ) | — | — | — | |
| Hyperagents ( Zhang et al., 2026b ) | — | — | ||
| SICA ( Robeyns et al., 2025 ) | — | |||
| SelfSearch (ours) |