Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations. To address these challenges, we introduce the Agentic Idea Manager (AIM), a fully autonomous framework for managing and exploring research directions in idea-driven automated research. Inspired by Bayesian optimization, AIM uses an Agentic Surrogate and an Agentic Acquisition mechanism to organize discovered ideas and guide their selection. A Solution Auditor maintains idea-solution integrity, while a Resource Planner adaptively allocates the remaining experimental budget across parallel search branches. Experiments on 10 AutoLab benchmark tasks show that AIM surpasses the strongest baseline by 1.6 percentage points on System Optimization tasks and 4.9 percentage points on long-horizon Model Development & CUDA tasks. Notably, AIM reaches the best baseline performance up to 3.1x faster in wall-clock time. We further provide a theoretical analysis of when searching over ideas becomes beneficial. Our analysis shows that explicit idea-level allocation makes semantic coverage directly controllable, and that broader coverage becomes increasingly valuable when competitive research directions are sparse among many plausible alternatives. Project Page: https://imhgchoi.github.io/agentic-idea-manager/
Figures & tables
Figure 1 : Solution-driven vs Idea-driven comparison. While Solution-driven approaches take the solution code as the target object for optimization, Idea-driven approaches first reason over research ideas, directions, or hypotheses based on evaluator feedback, after which the chosen ideas are implemented by the solver agent.
Figure 2 : The AIM pipeline: the Agentic Surrogate organizes and estimates candidate idea promise and the Agentic Acquisition selects ideas for execution and expands the idea pool. Results are audited by the Solution Auditor , while the Resource Planner dynamically allocates the budget.
Figure 3 : A visualization of a real pipeline example on the Flash Attention task. The x-axis shows all the cluster themes explored, numbers in each cell are the best scores from respective idea clusters at each round, and the stars indicate the trailing best score.
Methods
Scaffold
Flash Attention
Radix Sort
FFT Rust
AES128 Ctr
Z Range Scan
Average
Solution-driven Approaches
EvoX
Evolutionary
55.9 ± 3.3
55.5 ± 1.8
55.2 ± 0.1
63.1 ± 0.1
43.7 ± 1.6
54.7
AdaEvolve
Evolutionary
85.3 ± 7.6
62.4 ± 3.4
55.6 ± 0.1
65.9 ± 1.1
47.3 ± 2.3
63.3
AIRA (Evolutionary)
Evolutionary
76.3 ± 1.0
66.7 ± 2.5
56.2 ± 0.4
62.8 ± 0.5
51.1 ± 0.8
62.6
AIRA (MCTS)
MCTS
77.9 ± 0.2
61.3 ± 5.8
56.4 ± 0.1
62.6 ± 0.3
50.4 ± 1.3
61.7
Idea-driven Approaches
Table 1 : System Optimization Task Results . The mean and standard error of three independent runs are reported. All idea-driven approaches take the same research brief as additional context for a fair comparison.
Methods
Scaffold
MM World Model
Data Select IE
Huffman Dec.
NTT Butterfly
ICP Corr. Step
Average
Solution-driven Approaches
AdaEvolve
Evolutionary
–
82.7 ± 7.1
23.3 ± 1.5
54.9 ± 0.9
18.4 ± 11.8
44.8
AIRA (MCTS)
MCTS
–
38.4 ± 5.9
24.3 ± 10.8
51.6 ± 6.8
54.5 ± 0.9
42.2
Idea-driven Approaches
DeepScientist
List
14.6 ± 8.7
19.4 ± 19.4
35.5 ± 3.7
42.2 ± 7.1
55.7 ± 0.5
33.5
Arbor (max depth = 3)
Tree
33.7 ± 11.9
47.0 ± 16.1
37.0 ± 2.3
47.9 ± 5.0
40.8 ± 10.2
41.3
Table 2 : Model Development & CUDA Task Results . All idea-driven approaches take the same research brief as additional context for a fair comparison. Average scores are computed over all five tasks and omitted for methods with missing results. Missing results with ‘–’ is because the baseline methods could not handle the Moving Mnist World Model task that requires multiple file outputs.
Figure 6
Figure 6 : Semantic coverage required to achieve a fixed probability of discovering an ε -optimal research direction. As competitive directions become sparser, the effective semantic breadth Bε=K/Gε increases, requiring proportionally broader semantic coverage. The curves show the sufficient condition C≥Bεlog(1/δ) for different target success probabilities 1−δ .
Figure 7 : Idea-driven approaches generally show broader coverage of solutions in the embedding space. Embedding cosine similarity-based supplementary analysis is in Appendix F.5 . Plots for all tasks are in Appendix F.3 and Appendix F.4 .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Task
ScientistOne (Claude-Code Solver)
AIM (Claude-Code Solver)
Flash Attention
81.6 ± 3.5
84.5 ± 0.3
Radix Sort
67.4 ± 0.6
67.7 ± 0.2
FFT Rust
54.6 ± 0.3
55.4 ± 0.2
AES128 Ctr
66.9 ± 0.3
67.6 ± 0.6
Z-order Range Scan
52.2 ± 0.1
52.2 ± 0.3
Appendix
Table 3 : Solver Substitution . Comparison of AIM and ScientistOne with Claude Code Solvers.
Method
Flash Attention Scores
Full AIM
90.5 ± 0.8
without Organize
89.0 ± 0.4
without Estimate
87.9 ± 1.2
without Agentic Surrogate (both Organize and Estimate)
85.9 ± 2.7
Embedding-based Organize
83.4 ± 2.8
Direct Score Estimation
89.3 ± 1.2
Appendix
Table 4 : Further ablation studies on the Agentic Surrogate module.
Figure 8 : Time Efficiency Plots Across AutoLab Tasks.
Figure 9 : Time Efficiency Plots Across AutoLab Model Development & CUDA Tasks.
Figure 10 : Execution-to-Score Plots Across AutoLab Tasks.
Figure 11 : Execution-to-Score Plots Across AutoLab Model Development & CUDA Tasks.
Task
Metric
AIM (ours)
ScientistOne
Arbor
AIRA-MCTS
AIRA-EVO
Flash Attention
# LLM calls
1740±86
1550±40
432±73
121±19
83±13
# Input tokens ( ×106 )
90.1±6.0
70.3±8.8
14.4±1.8
0.9±0.2
0.5±0.1
# Output tokens ( ×106 )
2.25±0.12
1.77±0.15
0.29±0.00
3.17±0.06
2.20±0.09
Best score (%)
90.5±0.8
86.2±3.1
77.2±3.8
77.9±0.2
76.3±1.0
Radix Sort
# LLM calls
1899±14
1558±159
286±64
177±5
178±0
# Input tokens ( ×106 )
121.1±6.6
79.2±16.1
7.3±3.3
0.7±0.0
0.6±0.0
Appendix
Table 5 : LLM usage and cost per run on all ten AutoLab tasks (mean ± std over three runs). Output tokens include thinking token.
Figure 12 : Token Cost Comparison between AIM and ScientistOne.
Figure 13 : Agentic Surrogate: ORGANIZE . Heatmap of the best score of each cluster explored in each iteration. The star symbol marks the point when and where the new best scoring solution was discovered.
Figure 14 : Agentic Surrogate: ORGANIZE . Alluvial plots showing how the ideas flow and how they are re-organized throughout iterations.
Figure 15 : Cluster-level Ordinal Estimate’s Average Spearman Correlation.
Figure 16 : Trend of Spearman Correlation across Iterations.
Figure 17 : Explore/Exploit Action Distribution Across Branches.
Figure 18 : Expand operator mode distributions.
Figure 19 : Decrease in idea-mismatch flags after Solution Auditor idea reconstruction.
Figure 20 : Idea-driven approaches generally show broader coverage of solutions in the embedding space. Embedding cosine similarity-based supplementary analysis is in Appendix F.5 .
Figure 21 : Convex Hull Visualization
Method
Flash Attention
Radix Sort
Idea-driven Approaches
AIM
0.1150
0.0600
ScientistOne
0.1199
0.0661
Arbor
0.0396
0.0684
Solution-driven Approaches
AIRA (MCTS)
0.0311
0.0335
Appendix
Table 6: Supplementary Cosine Similarity Analysis. Entries are measured values of ( 47 ). Higher is more diverse. Bold marks the highest value per column.