Autoresearch agents tackle open-ended problems by repeatedly proposing candidate solutions, evaluating them, and using feedback to guide subsequent experiments. We show that independent runs of the same agent on the same task often plateau at substantially different scores, with gaps that persist even after considerable additional compute. Embedding their candidate artifacts by functional similarity provides further evidence that trajectories remain in localized regions of the solution space, which we call idea basins. To help agents escape these basins, we study a simple periodic intervention, fork-and-flush. Our method forks the agent into parallel trajectories, each inheriting the accumulated workspace but starting with a fresh chat context. After running each trajectory for a fixed horizon, the agent continues from the highest-scoring one. Across 13 long-horizon research and engineering tasks, with individual agent runs lasting up to several days, fork-and-flush outperformed the single-run and best-of-N baselines by a relative improvement of 66.0% and 44.4%, respectively, on the min-max normalized average score under an equal compute budget.
Figures & tables
Figure 1: Eight independent agents on p171 research task. Left: Best-so-far score against compute spend, with dots representing individual submissions. Runs plateau at different levels, with gaps persisting even after substantial additional compute. Right: The artifacts from these runs embedded by judged functional similarity. Agents converge into five distinct “idea basins” and remain there.
Figure 2: An illustration of fork-and-flush . Left: A single run (grey, dashed) keeps exploring similar ideas within one basin. Fork-and-flush clears the chat context on the same workspace ( ⋄ ), fans out K probes, retains the best ( ⋆ ), and discards the rest ( × ). Right: While the leader’s edits remain localized, a fresh-context probe jumps to a new basin via fork-and-flush and continues there.
single-run
best-of-4
fork-and-flush (ours)
task
max
mean
sd
max
mean
sd
max
mean
sd
Benchmark tasks: 6,000–25,000 credits per run, hours to several days each
polyomino
94.31
88.56
4.95
94.66
90.73
3.90
95.29
94.41
0.64
job-shop
10.36
9.66
0.87
10.27
9.91
0.24
10.92
10.33
0.59
AHC014
39.35
23.71
9.18
25.11
20.66
3.41
105.57
57.46
29.35
p150
130.80
123.47
4.90
133.25
123.47
5.99
135.37
129.74
4.43
Table 1: Final performance for single-run , best-of-4 , and fork-and-flush (ours) agents evaluated under matched compute budgets. We report the maximum score, mean score, and standard deviation across 5 independent runs (4 for tinyCIFAR ), with the highest max and mean per row highlighted in bold . The aggregate row reports the min–max normalized averages.
Method
Flush
Fork
Norm. avg
Single-run
×
×
0.45
Flushing
✓
×
0.55
Forking
×
✓
0.57
Fork-and-flush
✓
✓
0.71
Table 2: Impact of forking and flushing: min–max normalized mean score over 11 tasks. Best in bold , second best in italics .
Figure 3: Fork-and-flush escaping an idea basin on polyomino . Left: A mid-run fan-out overlaid on baseline basins. Rejected branches ( × ) from the fork point ( ⋄ ) remain trapped, but the selected branch (blue arrow) escapes the idea basin to achieve its final score ( ⋆ ). Right: Step distances between consecutive submissions. The selected branch’s jump ( ⋆ ) cleanly breaks out of the standard single-run refinement range (grey band).
Figure 4: Outcome variance is largest at the frontier of what the model can do. Spread vanishes where every run solves or fails the formula and peaks in between.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
task
single-run
flush only
fork only
fork-and-flush
polyomino
88.56
85.91
90.74
94.41
job-shop
9.66
10.00
10.43
10.33
AHC014
23.71
23.04
24.59
57.46
p150
123.47
160.80
128.57
129.74
p171
84.26
79.73
86.68
92.79
sphere-pack
43.87
43.74
44.36
44.94
Appendix
Table 3: Ablation on the impact of forking and flushing: mean final score of each method at the task’s budget, over the eight single-run baselines (seven on AHC014, p171 and LWE-rec, six on p150) and each method’s five draws (four flush-only cells on LLM-SQL; flush only : K=1 with a flush at each forced round; fork only : K=4 branches that continue the trunk’s context). Scores are on each task’s own scale, so rows are not comparable with each other; the last two rows min–max normalize every run of a task over all four methods before averaging. Best per row in bold , second best in italics . The ablation is run at 6,000 credits on Borden-inv , and every column there is read at that budget.
task
source
problem
score
credits
hours
polyomino
FCS algo 0
Pack a given set of polyominoes into a bounding rectangle of minimal area (C++ heuristic).
Ratio to a reference solution.
25k
29
job-shop
FCS algo 46
Job-shop scheduling: order operations on machines to minimize the makespan.
Ratio to known bounds.
10k
15
AHC014
FCS algo 159
Port of AtCoder Heuristic Contest 014.
Ratio to a reference solution.
10k
68
p150
FCS algo 150
Port of an AtCoder Heuristic Contest problem.
Ratio to a reference solution.
10k
70
p171
FCS algo 171
Port of an AtCoder Heuristic Contest problem.
Ratio to a reference solution.
10k
80
kernel
Anthropic take-home ( Anthropic, 2026 )
Anthropic’s performance-engineering take-home: optimize a kernel for a simulated VLIW machine; a deterministic simulator counts cycles.
Simulated-cycle score, higher is better.
100k
64
Appendix
Table 4: The thirteen tasks. “Credits” is the per-run budget, set per task to where single runs plateau. On tinyCIFAR it is the length of the shortest single-run baseline. “Hours” is the median time a single-run baseline takes to reach the budget, counting only time the agent is working and excluding pauses between resumed legs. Most of it is the agent’s own solvers and training jobs, which consume no credits; on LWE-rec and order-add single episodes of this kind run for up to two days.
Figure 5: Independent single runs across four embedded tasks: polyomino , p150 , AHC014 , and job-shop . Left: Best score achieved versus compute spent. Right: Run artifacts embedded by judged functional similarity, with k -means basins shaded.
task
single-run
plateau trigger
one round at 50%
rounds at 33/66%
polyomino
87.54
93.84 (1)
94.43 (1)
93.82 (6)
job-shop
9.23
11.44 (1)
10.24 (2)
10.99 (2)
AHC014
21.32
21.30 (3)
26.98 (4)
40.97 (5)
p150
118.24
126.72 (4)
145.13 (2)
135.52 (7)
p171
83.70
80.16 (4)
85.83 (5)
89.11 (11)
kernel
120.36
–
132.11 (5)
136.77 (2)
Appendix
Table 5: Fork-and-flush under three schedules, development runs ( K=4 ). We report the mean final score at the task’s budget, with the number of runs in parentheses. The single-run column is the mean of the eight baselines from the same period.
Figure 6: Figure 3 under two embeddings of the same judged triplets. Top: t-STE, as in Figure 3 . Bottom: SOE, aligned to the same frame, with basins recomputed on its coordinates; 99% of single-run artifacts keep their basin. In both, the rejected branches ( × ) stay in the branch point’s basin and within about one median single-run step of it, while the selected branch lands in a different basin: 43× the median single-run step under t-STE and 10× under SOE, in each case more than twice the largest step any single run takes.
t-STE
SOE
task
judged triplets
train
held out
train
held out
polyomino
6,375
0.90
0.87
0.90
0.85
AHC014
6,375
0.90
0.87
0.91
0.86
p171
6,375
0.88
0.85
0.88
0.84
p150
5,610
0.87
0.83
0.86
0.80
job-shop
6,375
0.86
0.81
0.85
0.79
Appendix
Table 6: Held-out validation of the idea-space maps. Agreement is the fraction of LLM triplet judgments the fitted two-dimensional map reproduces; chance is 0.50 . t-STE is the embedding used in the figures; SOE refits the same triplets with soft ordinal embedding.
Figure 7: Fork-and-flush trajectories against the single-run baselines of each task.
Figure 8: Nine fork-and-flush runs on their tasks’ idea maps (task-map basins shaded). Blue: the selected path ( ∘ opening round, ⋄ fork with a flush, ⋆ final artifact); vermillion × : rejected branches.
Long-running coding agents such as autoresearch can persistently discover optimizations for open-ended problems. However, they tend to converge onto a single high-level approach, then proceed with low-level edits while missing other superior approaches to the problem. We hypothesize two harness-level design choices contribute to this behavior: accumulating context in a single long-running agent and only exposing a single program state to edit. We introduce SwarmResearch, an orchestrator-subagent harness in which a Shepherd Agent uses global context to steer a population of Search Agents, each operating with local context in their respective git branch. On open-ended optimization tasks, SwarmResearch discovers better or comparable solutions to state-of-the-art LLM-guided evolution and multi-agent techniques on 13/15 tasks, driven by higher-level exploration. Compared with fixed scaling of serial and parallel agents, SwarmResearch's orchestrator-guided scaling discovers better-performing solutions by adapting parallelism at different search depths.
A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have inspired large industry investment, motivated by their potential to automate time-consuming human labor and customize machine learning solutions for specialized applications. In this paper, we study the modeling pipeline at the core of these autoresearch systems and identify common failure modes when they are applied to tabular datasets: (1) they waste compute resolving the same bugs over and over again; (2) they often fail to tune hyperparameters even when they have a large remaining compute budget; (3) the tree-search algorithms that power them do not explore; and (4) they perform data analysis, mimicking the humans whose data they are trained on, but do not use that analysis to make downstream decisions. We explore targeted interventions and find that a global debug consultant that shares discovered runtime constraints across all branches of the search tree, prompt- and control-level enhancements, and refined tree-search algorithms successfully recover wasted compute. Our results show that large gains in autoresearch agent performance are achievable through agentic design alone, holding the underlying language model fixed.
Au Kwok Chun, Abhigyan Acherjee, Amrutha Rao +4
Department of Computer Science, Columbia University · 2AI, Analytics and Future of Work Initiative, Georgetown University · Department of Applied Mathematics, Columbia University +3
Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations. To address these challenges, we introduce the Agentic Idea Manager (AIM), a fully autonomous framework for managing and exploring research directions in idea-driven automated research. Inspired by Bayesian optimization, AIM uses an Agentic Surrogate and an Agentic Acquisition mechanism to organize discovered ideas and guide their selection. A Solution Auditor maintains idea-solution integrity, while a Resource Planner adaptively allocates the remaining experimental budget across parallel search branches. Experiments on 10 AutoLab benchmark tasks show that AIM surpasses the strongest baseline by 1.6 percentage points on System Optimization tasks and 4.9 percentage points on long-horizon Model Development & CUDA tasks. Notably, AIM reaches the best baseline performance up to 3.1x faster in wall-clock time. We further provide a theoretical analysis of when searching over ideas becomes beneficial. Our analysis shows that explicit idea-level allocation makes semantic coverage directly controllable, and that broader coverage becomes increasingly valuable when competitive research directions are sparse among many plausible alternatives. Project Page: https://imhgchoi.github.io/agentic-idea-manager/