Stateless Language Agents: Scaling Long-Horizon Automated Research
Organizations: Stanford University
Abstract
Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay growing histories, duplicate one another's work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that saturate early, leaving these failure modes untested. We trace these failures to two choices: where research state lives and who decides what to try next. We introduce Stateless Language Agents (SLAs), built on the principle of stateful search with stateless agents: no agent carries its conversation across invocations; instead, the harness owns the research state (candidate solutions and measured outcomes) and reconstructs a fresh and role-specific context for every invocation. What each agent sees becomes an explicit design choice rather than a history that grows with the run. We implement this principle in the SLA framework, where a stateless Advisor reads harness-summarized evidence across search directions and assigns concrete experiments to parallel Workers. We evaluate SLA against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets of up to one billion tokens. SLA achieves the best final result on every task and reaches the strongest kernel baseline's final performance with over 84% fewer tokens. Ablations from shared checkpoints show that focused contexts and explicit assignments each contribute to SLA's progress, with effects that can compound over full runs, while the Advisor consumes less than 0.6% of tokens. These results argue for SLAs, which keep durable research state out of agent conversations, and show that short evaluation horizons can misjudge research systems and their components.
Figures & tables
| Codex (3 runs) | Claude Code (1 run) | |||||
| Method | 250M Cycles | 1B Cycles | To match Tokens (M) | 250M Cycles | 1B Cycles | To match Tokens (M) |
| EvoX | 1478.3 112.0 | 1343.3 4.0 | NR | 1174 | 1153 | NR |
| CORAL | 1524.0 101.7 | 1350.0 134.2 | NR | 1205 | 1130 | 888.3 |
| SwarmResearch | 1501.7 185.7 | 1275.7 133.9 | 986.3 | 1648 | 1243 | NR |
| SLA (ours) | 1149.3 18.9 | 1112.0 8.9 | 67.9 | 1079 | 1039 | 138.5 |
| Reduction vs. best baseline | 22.3% | 12.8% | 93.1% | 8.1% | 8.1% | 84.4% |
| libexpat x86-64 Assembly | Git Zig | |||||
| Method | 175M Score | 700M Score | To match Tokens (M) | 175M Score | 700M Score | To match Tokens (M) |
| EvoX | 10.71 | 11.33 | NR | 19.29 | 21.42 | NR |
| CORAL | 8.74 | 12.08 | NR | 19.46 | 22.55 | NR |
| SwarmResearch | 19.67 | 24.58 | 559.9 | 24.32 | 27.98 | 697.3 |
| SLA (ours) | 18.82 | 31.33 | 354.1 | 20.42 | 28.84 | 630.2 |
| vs. best baseline | 0.85 | 6.75 | 36.8% | 3.90 | 0.86 | 9.6% |
| SOL-ExecBench #1 | SOL-ExecBench #58 | |||||
| Method | 125M SOL | 500M SOL | To match Tokens (M) | 125M SOL | 500M SOL | To match Tokens (M) |
| EvoX | 0.710 | 0.710 | 46.1 | 0.773 | 0.804 | 354.1 |
| CORAL | 0.386 | 0.552 | NR | 0.694 | 0.714 | NR |
| SwarmResearch | 0.598 | 0.670 | NR | 0.764 | 0.764 | NR |
| SLA (ours) | 0.706 | 0.763 | 152.7 | 0.836 | 0.862 | 14.5 |
| vs. best baseline | 0.6% | 7.5% | 231.2% | 8.2% | 7.2% | 95.9% |
| Anthropic Kernel | FrontierSWE (libexpat) | |||||||
| Configuration | 100M 200M Cycles | 200M 300M Cycles | 70M 170M Score ( 100) | 140M 240M Score ( 100) | ||||
| Shared checkpoint | 1218 | 1197 | 16.12 | 17.69 | ||||
| SLA (full) | 1192.3 2.5 | 25.7 | 1170.0 14.2 | 27.0 | 18.78 0.92 | 2.66 | 20.13 0.20 | 2.44 |
| Focused contexts | ||||||||
| w/o Advisor reconstruction | 1198.3 2.3 | 23.3% | 1188.7 6.8 | 69.3% | 16.86 0.64 | 72.2% | 19.07 0.31 | 43.4% |
| w/o Worker isolation | 1203.3 7.1 | 42.8% | 1174.7 4.5 | 17.4% | 16.79 0.47 | 74.8% | 19.77 0.12 | 14.8% |
| Token composition (%) | Coordinator share (%) | ||||
| Method | Cached input | Uncached input | Output | Tokens | Cost |
| Anthropic kernel (1B tokens) | |||||
| SwarmResearch | 94.67 | 4.36 | 0.97 | 8.39 | 6.14 |
| SLA (ours) | 93.11 | 5.59 | 1.30 | 0.24 | 1.20 |
| FrontierSWE libexpat (700M tokens) | |||||
| SwarmResearch | 94.69 | 4.36 | 0.95 | 10.70 | 7.35 |
| Anthropic Kernel | FrontierSWE | |||||||||||||
| Workers | Performance Cycles | Time Hours | Dart Score | Git Zig Score | libexpat Score | Lua Score | Mean time Hours | |||||||
| Short horizon (25% of token budget) | ||||||||||||||
| 1151 | 27.5 | 12.1 | 21.8 | 21.6 | 93.4 | 18.0 | ||||||||
| 1149 | 2 | 11.8 | 2.3 | 15.4 | 3.3 | 20.4 | 1.4 | 18.8 | 2.8 | 80.2 | 13.2 | 7.4 | 2.4 | |
| 1488 | 337 | 5.1 | 5.4 | 13.8 | 1.7 | 19.7 | 2.1 | 28.2 | 6.6 | 78.0 | 15.4 | 4.5 | 4.0 | |
| 1325 | 174 | 4.3 | 6.4 | 25.5 | 13.4 | 21.9 | 0.1 | 23.6 | 2.0 | 46.2 | 47.2 | 2.8 | 6.4 | |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Anthropic Kernel Cycles | FrontierSWE (libexpat) Score ( 100) | |||
| Trials per epoch | 100M 200M | 200M 300M | 70M 170M | 140M 240M |
| Shared checkpoint | 1218 | 1197 | 16.12 | 17.69 |
| 1 | 1194.7 7.2 | 1179.3 13.5 | 17.29 1.65 | 20.94 1.39 |
| 3 (default) | 1192.3 2.5 | 1170.0 14.2 | 18.78 0.92 | 20.13 0.20 |
| 5 | 1195.0 9.6 | 1176.0 15.4 | 17.02 1.00 | 19.42 0.16 |
| Attempt | Time (h) | Agent | Event | Best |
| 89 | 3.3 | A | Adds hex-float and unsigned string.format conversions; writes note | 83.52 |
| 116 | 4.4 | B | Lineage without the feature takes the lead | 86.81 |
| 140 | 5.3 | D | Best improves to 88.46; feature still absent | 88.46 |
| 149 | 5.5 | B | Reads A’s note and reimplements the feature | 89.01 |
| 151 | 5.7 | C | Reads the same note and restores the same feature | 89.01 |
| 158 | 6.0 | D | Composes the same feature again | 89.01 |
| Attempt | Time (h) | Agent | Parent | Change | Score |
| 97 | 2.72 | A | 92 | Rejects undeclared namespace prefixes | 8.42 |
| 101 | 2.82 | A | 97 | Rejects duplicate expanded attribute names | 8.74 |
| 102 | 2.85 | B | 97 | Rejects duplicate expanded attribute names | 8.74 |
| 104 | 2.99 | A | 101 | Reserved xml / xmlns prefix and namespace checks | 9.79 |
| 107 | 3.03 | B | 102 | Reserved xml / xmlns prefix and namespace checks | 9.79 |