AIM: Agentic Idea Management for Automated Research
Organizations: Google Cloud AI Research · University of Wisconsin-Madison
Abstract
Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations. To address these challenges, we introduce the Agentic Idea Manager (AIM), a fully autonomous framework for managing and exploring research directions in idea-driven automated research. Inspired by Bayesian optimization, AIM uses an Agentic Surrogate and an Agentic Acquisition mechanism to organize discovered ideas and guide their selection. A Solution Auditor maintains idea-solution integrity, while a Resource Planner adaptively allocates the remaining experimental budget across parallel search branches. Experiments on 10 AutoLab benchmark tasks show that AIM surpasses the strongest baseline by 1.6 percentage points on System Optimization tasks and 4.9 percentage points on long-horizon Model Development & CUDA tasks. Notably, AIM reaches the best baseline performance up to 3.1x faster in wall-clock time. We further provide a theoretical analysis of when searching over ideas becomes beneficial. Our analysis shows that explicit idea-level allocation makes semantic coverage directly controllable, and that broader coverage becomes increasingly valuable when competitive research directions are sparse among many plausible alternatives. Project Page: https://imhgchoi.github.io/agentic-idea-manager/
Figures & tables
| Methods | Scaffold | Flash Attention | Radix Sort | FFT Rust | AES128 Ctr | Z Range Scan | Average |
|---|---|---|---|---|---|---|---|
| Solution-driven Approaches | |||||||
| EvoX | Evolutionary | 55.9 3.3 | 55.5 1.8 | 55.2 0.1 | 63.1 0.1 | 43.7 1.6 | 54.7 |
| AdaEvolve | Evolutionary | 85.3 7.6 | 62.4 3.4 | 55.6 0.1 | 65.9 1.1 | 47.3 2.3 | 63.3 |
| AIRA (Evolutionary) | Evolutionary | 76.3 1.0 | 66.7 2.5 | 56.2 0.4 | 62.8 0.5 | 51.1 0.8 | 62.6 |
| AIRA (MCTS) | MCTS | 77.9 0.2 | 61.3 5.8 | 56.4 0.1 | 62.6 0.3 | 50.4 1.3 | 61.7 |
| Idea-driven Approaches | |||||||
| Methods | Scaffold | MM World Model | Data Select IE | Huffman Dec. | NTT Butterfly | ICP Corr. Step | Average |
|---|---|---|---|---|---|---|---|
| Solution-driven Approaches | |||||||
| AdaEvolve | Evolutionary | – | 82.7 7.1 | 23.3 1.5 | 54.9 0.9 | 18.4 11.8 | 44.8 |
| AIRA (MCTS) | MCTS | – | 38.4 5.9 | 24.3 10.8 | 51.6 6.8 | 54.5 0.9 | 42.2 |
| Idea-driven Approaches | |||||||
| DeepScientist | List | 14.6 8.7 | 19.4 19.4 | 35.5 3.7 | 42.2 7.1 | 55.7 0.5 | 33.5 |
| Arbor (max depth = 3) | Tree | 33.7 11.9 | 47.0 16.1 | 37.0 2.3 | 47.9 5.0 | 40.8 10.2 | 41.3 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | ScientistOne (Claude-Code Solver) | AIM (Claude-Code Solver) |
|---|---|---|
| Flash Attention | 81.6 3.5 | 84.5 0.3 |
| Radix Sort | 67.4 0.6 | 67.7 0.2 |
| FFT Rust | 54.6 0.3 | 55.4 0.2 |
| AES128 Ctr | 66.9 0.3 | 67.6 0.6 |
| Z-order Range Scan | 52.2 0.1 | 52.2 0.3 |
| Method | Flash Attention Scores |
|---|---|
| Full AIM | 90.5 0.8 |
| without Organize | 89.0 0.4 |
| without Estimate | 87.9 1.2 |
| without Agentic Surrogate (both Organize and Estimate) | 85.9 2.7 |
| Embedding-based Organize | 83.4 2.8 |
| Direct Score Estimation | 89.3 1.2 |
| Task | Metric | AIM (ours) | ScientistOne | Arbor | AIRA-MCTS | AIRA-EVO |
|---|---|---|---|---|---|---|
| Flash Attention | # LLM calls | |||||
| # Input tokens ( ) | ||||||
| # Output tokens ( ) | ||||||
| Best score (%) | ||||||
| Radix Sort | # LLM calls | |||||
| # Input tokens ( ) |
| Method | Flash Attention | Radix Sort |
|---|---|---|
| Idea-driven Approaches | ||
| AIM | 0.1150 | 0.0600 |
| ScientistOne | 0.1199 | 0.0661 |
| Arbor | 0.0396 | 0.0684 |
| Solution-driven Approaches | ||
| AIRA (MCTS) | 0.0311 | 0.0335 |