Frameworks that use large language models for scientific discovery typically rely on a fixed, human-designed algorithm that decides what the model sees at each step, leaving the model only the role of proposer. The model knows nothing of the search beyond what it is shown. As models grow more capable, a question arises: does a search strategy chosen by a human before the run scale better than promoting the model from proposer to planner and letting it own the search? The Bitter Lesson suggests that choosing the strategy in advance is the kind of hand-designed structure that general methods eventually outscale. We introduce AgentDiscover, in which a coding agent plans the search using its context as working memory, runs experiments, and records every attempt in a database of ideas, candidates, and their relations. This database serves as the agent's long-term memory and is structured so that the selection rules of classical algorithms such as MAP-Elites and Monte Carlo tree search each reduce to a single query, which the agent is free to use, combine, or replace. A server maintains the database and steers the agent after every submission, keeping it on course over long runs. In our experiments, AgentDiscover is more cost-efficient than existing frameworks, reaching better scores at lower cost. On tasks in kernel engineering, biology, algorithm design, and mathematics, AgentDiscover outperforms prior discovery frameworks. Its programs would have placed first among human competitors in seven past AtCoder heuristic contests, and on eleven mathematical and systems optimization tasks it matches or exceeds every baseline that uses the same model. Our code is available at https://github.com/mhdfb/AgentDiscover.
Figures & tables
Figure 1: On the kernel builder task, AgentDiscover reaches 960 cycles, the lowest count among the frameworks we compare against, and it reaches comparable cycle counts at lower model spend (more tasks in Appendix E ). All platforms run on Claude Opus 5 with the same budget.
Figure 2: AgentDiscover lets a coding agent direct the evolutionary search. (a) A fixed algorithm decides what the LLM sees in prior systems; in AgentDiscover, the agent queries its memory itself. (b) The agent’s loop: (1) query the graph database, (2) experiment, (3) submit a candidate, (4) receive its score and a steering message. New sessions start with a fresh context and the same memory.
Figure 3: AgentDiscover on a toy task: place 15 circles in a square to maximize the sum of their radii. A coding agent does the search. An MCP server stores all past ideas and candidates in a graph database and scores new candidates. (0) The agent starts in its folder (worktree). A CLAUDE.md file there describes the task and a summary of the database. (1) The agent explores the database with any read-only query. (2) It runs experiments, using available tools or new ones it writes first. (3) It submits a candidate. The server scores it with the task’s evaluator and replies with a steering message. The agent repeats steps 1–3 for each new candidate (inner loop). After the last one, a new agent with a fresh context starts on the same database and keeps all tools written so far (outer loop).
Task
TTT-Discover
CORAL
AgentDiscover
Kernel
TriMul ↓
2,198
2,419
2,080
MLA-Decode ↓
– †
1,582
1,474
Biology
Pancreas ↓
0.1655 ‡
0.1676
0.1603
PBMC ↓
0.154
0.156
0.147
Tabula lung ↓
0.136
0.139
0.132
Algo.
AHC039 ↑
567,062
566,654
577,217
Table 1: Results on kernel engineering, biology, algorithm design, and mathematics tasks. CORAL and AgentDiscover run at matched model spend: GPT-6 Sol at 10pertaskonthefirsttwocategories,ClaudeOpus5at100 on the last two; TTT-Discover results are from Yuksekgonul et al. (2026) . Kernel entries are runtimes in μ s (TriMul on A100, MLA-Decode on H200); biology entries are single-cell denoising MSE. Best in bold. † Published only on AMD MI300X (1,669–1,706 μ s). ‡ Published program re-scored by our evaluator on the search split.
OpenEvolve
EvoX
CORAL
AgentDiscover
Task
Score
Con.
Ext.
Score
Con.
Ext.
Score
Con.
Ext.
Score
Con.
Ext.
AHC006 ↑
2,238,362
1st
13th
2,197,920
4th
36th
2,317,352
1st
1st
2,309,017
1st
1st
AHC012 ↑
99,090,668
5th
39th
98,579,333
9th
55th
99,977,717
1st
4th
99,997,290
1st
3rd
AHC015 ↑
144,369,058
15th
130th
151,853,188
4th
93rd
156,230,796
4th
57th
168,660,525
1st
5th
AHC025 † ↑
92.60
37th
51st
94.46
21st
30th
98.50
2nd
6th
98.76
1st
3rd
AHC041 ↑
76,321,331
10th
58th
75,787,437
31st
127th
77,358,800
1st
6th
78,044,915
1st
1st
Table 2: Results on past AtCoder heuristic contests with Claude Opus 5, scored on the official system tests. Methods are compared at equal model spend on each contest (120–200, depending on the contest). Con. is the rank among the original competitors; Ext. includes post-contest submissions. † Relative score; see Appendix E.2 .
Task
OpenEvolve
ShinkaEvolve
EvoX
CORAL
AgentDiscover
Systems
EPLB ↑
0.127
0.129
0.146
0.149
0.153
PRISM ↑
26.26
26.26
26.26
26.26
26.26
LLM-SQL ↑
0.716
0.724
0.726
0.731
0.737
Txn Sched. ↑
3774
3802
3984
4566
4651
Cloudcast ↓
627.2
627.8
623.5
618.4
618.0
Math
Circle-Pack. ↑
2.6293
2.6001
2.6320
2.6360
2.6360
Table 3: Results on five systems and six mathematical optimization tasks. All methods use Claude Opus 4.6 with a budget of at most 100 candidates. Published baseline results are taken from Qu et al. (2026) . Arrows indicate the direction of improvement; best scores and ties at common reporting precision are bold. For signal processing, the faded row reports scores from the original evaluator, which permits reward hacking; these scores are excluded from performance comparisons. The next row uses the corrected evaluator (Appendix C ).
Budget
OpenEvolve
ShinkaEvolve
EvoX
CORAL
AgentDiscover
$25
1,248
1,344
1,225
1,257
1,016
$50
1,213
1,218
1,153
1,214
976
$75
1,178
1,210
1,119
1,157
973
$100
1,174
1,187
1,098
1,148
960
$118
1,174
1,156
1,098
1,148
960
Table 4: Kernel builder at common model-spend budgets. All methods use Claude Opus 5 with low reasoning effort and the same evaluator. Each entry is the lowest cycle count achieved at or below the indicated budget; lower is better. The best result at each budget is bold.
Figure 4: Context resets. (a) Without a reset, the input cost per call keeps growing. (b) After a reset, the agent sees the search afresh, making new ideas more likely. (c,d) After a stall of over five candidates, a reset raises the chance and size of an improvement; a meta pass raises them further.
Figure 5: (a,b) Four parallel agents find good solutions sooner in wall-clock time than one agent, but not at lower model spend. Lines show the mean over five runs, bands the standard error, and faint lines individual runs. (c) With meta guidance, a new search agent beats the best score so far sooner: within two candidates in 73% of cases, against 33% without meta guidance. (d) With meta guidance, the candidates the agent proposes are better overall, not only its best one: each curve shows the share of candidates that reach at least a given fitness, and the median candidate rises from 0.801 to 0.832.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Label
Key
Holds
User
name
The person who owns the agents.
Agent
key
One branch: a persistent search-agent identity and the path of its worktree.
Session
id
One search-agent session: its iteration, model and briefing, its reasoning text ( thinking_text ), start and end times, the fingerprint of its call token ( token_hash ), its own lesson ( lesson_learned ), and the meta agent’s judgment of it ( meta_lesson ).
Candidate
genome_hash
One unique program ( genome , the full solution.py ), its fitness and evaluation status, one property per evaluator metric, and the agent’s predictions, made before scoring.
Idea
name
A named idea, reused across candidates, with a description.
Resource
name
A file the problem provides, and where agents find it.
Appendix
Table 5: Node labels of the knowledge graph, with what each node holds.
Table 6: Knowledge-graph relationships and their writers. Endpoints are shown as source → target; facts are reconstructed from the session log, whereas claims are the counterpart declared by the search agent.
Tool
Caller
Arguments
Effect
read_graph
both
cypher
Runs a read-only Cypher query and returns rows as JSON.
add_idea
search
name , description , related_to
Creates or reuses an idea and records its relations to existing ideas.
submit_candidate
search
idea , predictions, parents , resources
Evaluates the current solution.py , records its provenance and result, and returns steering feedback.
write_lesson
search
text
Stores the session’s final lesson.
write_meta_lesson
meta
session_id , text
Stores the meta agent’s assessment of one session’s search process.
add_tool
meta
branch , name , description , code , params
Registers a read-only Cypher query as a tool for one branch.
Appendix
Table 7: MCP tools available to search and meta agents.
Setting
Search model
Effort
Web
Budget ceiling
Systems + math
Opus 4.6
default
off
100 candidates
AtCoder (5 contests)
Opus 5
default
off
matched model spend
Multi-domain (AHC039, AHC058)
Opus 5
default
off
matched model spend
Multi-domain (AC1, AC2)
Opus 5
default
on
matched model spend
Multi-domain (TriMul, MLA-Decode, denoising)
GPT-6 Sol
default
off
matched model spend
Kernel optimization
Opus 5
low
off
matched model spend
Appendix
Table 8: Model, access, and system-level budget ceiling by evaluation setting. “On” for AC1 and AC2 applies to both the CORAL and AgentDiscover runs.
Setting
Cand. / session
Max iter.
Search timeout (s)
Systems + math
25
5
28,800
AtCoder (5 contests)
20
10
36,000
Multi-domain (Opus 5 tasks)
20
5
57,600
Multi-domain (GPT-6 Sol tasks)
5
10
28,800
Kernel optimization
10
10
28,800
Parallel-agent ablation
8
10
14,400
Appendix
Table 9: Search schedule and wall-clock ceiling by evaluation setting.
Setting
Meta model
Meta timeout (s)
Systems + math
Opus 4.6
7,200
AtCoder (5 contests)
Opus 5
3,600
Multi-domain (Opus 5 tasks)
Opus 5
5,400
Multi-domain (GPT-6 Sol tasks)
GPT-6 Sol
7,200
Kernel optimization
Opus 5
7,200
Parallel-agent ablation
Haiku 4.5
3,600
Appendix
Table 10: Meta-agent configuration by evaluation setting.
Figure 6: Best-so-far performance against billed model spend on four further AtCoder heuristic contests, measured on the public evaluation cases. Each objective is normalized by the corresponding contest-winner mean, so higher is better and 1× is the reference value.
Figure 7: Best-so-far quality against model spend on the three TTT-Discover tasks, AgentDiscover and CORAL with GPT-6 Sol, cut at the 10budgetofTable1.QualityistheTTT−Discoverreferencedividedbythecandidate’sobjective(runtimeforthekernels,pancreasMSEfordenoising),sohigherisbetterand1\timesmatchesTTT−Discover;forMLA−Decodethereferenceisthe1,700\mu$ s target of its prompt. AgentDiscover reaches each reference sooner and ends higher on all three tasks.
Figure 8: AHC025 fixed-objective comparison at matched model spend. The trajectories and the upper-right summary compare methods by their best mean absolute objective on the same 50 public evaluation cases (lower objective is better). The dashed 1× reference uses the contest winner’s mean over the 5,000 official system-test cases, whereas the trajectories use the 50 public cases. Official leaderboard (relative) scores are reported in Table 2 .
Scientific discovery can be formulated as an iterative search process over the space of hypotheses and experiments. Contemporary methods navigate this space using heuristics such as MCTS. These algorithms conflate the merit of a hypothesis with the quality of its experimental execution. A promising hypothesis with preliminary execution is therefore ranked below a modest hypothesis whose execution is refined. Moreover, prior methods prune the search logs as the search progresses because the accumulated history outgrows the context window. We propose Agentic Reasoning for Tree Search (ARTS), where we deploy a reasoning language model to navigate this space. The model inspects prior execution logs, diagnoses whether earlier failures arose from faulty implementations or bad hypotheses, and selects the hypothesis to build on next. To mitigate challenges with context length, ARTS uses test-time training to instill the knowledge of search tree in the model weights. Across 22 tasks from MLGym and MLEBench, we show that ARTS outperforms leading algorithms, with over 15.3% relative improvement in the normalized score. With test-time training we show that a Qwen3-4B agent can match performance with closed-source frontier models like Gemini-3 Pro and GPT o3-reasoning with upto 5x lower inference cost. We further observe that on partially observable RL tasks, the test-time trained Qwen3-4B scientist surpasses ARTS with the o3 scientist by rediscovering the human-best recurrent-memory solution that heuristic methods prune away.
Large language models have advanced automated algorithm discovery by synthesizing executable code, but existing frameworks trap them in rigid search pipelines with pre-defined control flows. This limitation restricts adaptive reasoning, blocks cross-paradigm transfer, and overlooks richer execution feedback. To bridge this gap, we introduce an end-to-end framework, AlgoEvo, a unified agentic architecture that transforms automated algorithm discovery into an interactive, knowledge-accumulating process. An autonomous agent dynamically inspects, diagnoses, and edits code based on runtime feedback. A design skill hub decouples paradigm-specific knowledge from the core discovery engine, allowing a unified workflow to seamlessly handle single-heuristic, multi-objective, and multi-component design. Meanwhile, a hierarchical experience bank organizes search trajectories into a task-level tree to guide exploration and consolidates cross-task patterns into reusable skills. Across six representative benchmark tasks, AlgoEvo reaches state-of-the-art performance with as little as 7% of the evaluation budget and reduced token consumption, demonstrating strong intra-task accumulation, cross-task transfer, and the ability to reproduce or exceed the strongest existing methods through flexible skill activation.
Junhao Qiu, Qinglong Hu, Ji Cheng +3
Department of Computer Science, City University of Hong Kong · Huawei Noah’s Ark Lab · Institute of Advanced Intelligence and Computing, A*STAR
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent? We present ReASearch, a unified framework for reasoning-driven optimization in which the agent autonomously decides what to evaluate, how to diagnose failures, which edits to make, and when to verify or restart. Rather than serving only as a proposal generator guided by hand-designed heuristics, the agent actively analyzes outcomes, allocates budget, and refines its strategy over long horizons through persistent memory. With a shared agent loop and domain-specific tools, ReASearch instantiates the exact same scaffold to optimize prompts, programs, and ML workflows. Across 14 diverse tasks, it is competitive with and mostly better than specialized optimization systems, achieving gains of 2% to 40% over strong domain-specific baselines, and in some cases discovering solutions that improve on prior human best-known results. Crucially, we observe that complex search behaviors, which are typically implemented by explicit controllers, emerge naturally from the agent's reasoning process.