Autoresearch agents tackle open-ended problems by repeatedly proposing candidate solutions, evaluating them, and using feedback to guide subsequent experiments. We show that independent runs of the same agent on the same task often plateau at substantially different scores, with gaps that persist even after considerable additional compute. Embedding their candidate artifacts by functional similarity provides further evidence that trajectories remain in localized regions of the solution space, which we call idea basins. To help agents escape these basins, we study a simple periodic intervention, fork-and-flush. Our method forks the agent into parallel trajectories, each inheriting the accumulated workspace but starting with a fresh chat context. After running each trajectory for a fixed horizon, the agent continues from the highest-scoring one. Across 13 long-horizon research and engineering tasks, with individual agent runs lasting up to several days, fork-and-flush outperformed the single-run and best-of-N baselines by a relative improvement of 66.0% and 44.4%, respectively, on the min-max normalized average score under an equal compute budget.
Figures & tables
Figure 1: Eight independent agents on p171 research task. Left: Best-so-far score against compute spend, with dots representing individual submissions. Runs plateau at different levels, with gaps persisting even after substantial additional compute. Right: The artifacts from these runs embedded by judged functional similarity. Agents converge into five distinct “idea basins” and remain there.
Figure 2: An illustration of fork-and-flush . Left: A single run (grey, dashed) keeps exploring similar ideas within one basin. Fork-and-flush clears the chat context on the same workspace ( ⋄ ), fans out K probes, retains the best ( ⋆ ), and discards the rest ( × ). Right: While the leader’s edits remain localized, a fresh-context probe jumps to a new basin via fork-and-flush and continues there.
single-run
best-of-4
fork-and-flush (ours)
task
max
mean
sd
max
mean
sd
max
mean
sd
Benchmark tasks: 6,000–25,000 credits per run, hours to several days each
polyomino
94.31
88.56
4.95
94.66
90.73
3.90
95.29
94.41
0.64
job-shop
10.36
9.66
0.87
10.27
9.91
0.24
10.92
10.33
0.59
AHC014
39.35
23.71
9.18
25.11
20.66
3.41
105.57
57.46
29.35
p150
130.80
123.47
4.90
133.25
123.47
5.99
135.37
129.74
4.43
Table 1: Final performance for single-run , best-of-4 , and fork-and-flush (ours) agents evaluated under matched compute budgets. We report the maximum score, mean score, and standard deviation across 5 independent runs (4 for tinyCIFAR ), with the highest max and mean per row highlighted in bold . The aggregate row reports the min–max normalized averages.
Method
Flush
Fork
Norm. avg
Single-run
×
×
0.45
Flushing
✓
×
0.55
Forking
×
✓
0.57
Fork-and-flush
✓
✓
0.71
Table 2: Impact of forking and flushing: min–max normalized mean score over 11 tasks. Best in bold , second best in italics .
Figure 3: Fork-and-flush escaping an idea basin on polyomino . Left: A mid-run fan-out overlaid on baseline basins. Rejected branches ( × ) from the fork point ( ⋄ ) remain trapped, but the selected branch (blue arrow) escapes the idea basin to achieve its final score ( ⋆ ). Right: Step distances between consecutive submissions. The selected branch’s jump ( ⋆ ) cleanly breaks out of the standard single-run refinement range (grey band).
Figure 4: Outcome variance is largest at the frontier of what the model can do. Spread vanishes where every run solves or fails the formula and peaks in between.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
task
single-run
flush only
fork only
fork-and-flush
polyomino
88.56
85.91
90.74
94.41
job-shop
9.66
10.00
10.43
10.33
AHC014
23.71
23.04
24.59
57.46
p150
123.47
160.80
128.57
129.74
p171
84.26
79.73
86.68
92.79
sphere-pack
43.87
43.74
44.36
44.94
Appendix
Table 3: Ablation on the impact of forking and flushing: mean final score of each method at the task’s budget, over the eight single-run baselines (seven on AHC014, p171 and LWE-rec, six on p150) and each method’s five draws (four flush-only cells on LLM-SQL; flush only : K=1 with a flush at each forced round; fork only : K=4 branches that continue the trunk’s context). Scores are on each task’s own scale, so rows are not comparable with each other; the last two rows min–max normalize every run of a task over all four methods before averaging. Best per row in bold , second best in italics . The ablation is run at 6,000 credits on Borden-inv , and every column there is read at that budget.
task
source
problem
score
credits
hours
polyomino
FCS algo 0
Pack a given set of polyominoes into a bounding rectangle of minimal area (C++ heuristic).
Ratio to a reference solution.
25k
29
job-shop
FCS algo 46
Job-shop scheduling: order operations on machines to minimize the makespan.
Ratio to known bounds.
10k
15
AHC014
FCS algo 159
Port of AtCoder Heuristic Contest 014.
Ratio to a reference solution.
10k
68
p150
FCS algo 150
Port of an AtCoder Heuristic Contest problem.
Ratio to a reference solution.
10k
70
p171
FCS algo 171
Port of an AtCoder Heuristic Contest problem.
Ratio to a reference solution.
10k
80
kernel
Anthropic take-home ( Anthropic, 2026 )
Anthropic’s performance-engineering take-home: optimize a kernel for a simulated VLIW machine; a deterministic simulator counts cycles.
Simulated-cycle score, higher is better.
100k
64
Appendix
Table 4: The thirteen tasks. “Credits” is the per-run budget, set per task to where single runs plateau. On tinyCIFAR it is the length of the shortest single-run baseline. “Hours” is the median time a single-run baseline takes to reach the budget, counting only time the agent is working and excluding pauses between resumed legs. Most of it is the agent’s own solvers and training jobs, which consume no credits; on LWE-rec and order-add single episodes of this kind run for up to two days.
Figure 5: Independent single runs across four embedded tasks: polyomino , p150 , AHC014 , and job-shop . Left: Best score achieved versus compute spent. Right: Run artifacts embedded by judged functional similarity, with k -means basins shaded.
task
single-run
plateau trigger
one round at 50%
rounds at 33/66%
polyomino
87.54
93.84 (1)
94.43 (1)
93.82 (6)
job-shop
9.23
11.44 (1)
10.24 (2)
10.99 (2)
AHC014
21.32
21.30 (3)
26.98 (4)
40.97 (5)
p150
118.24
126.72 (4)
145.13 (2)
135.52 (7)
p171
83.70
80.16 (4)
85.83 (5)
89.11 (11)
kernel
120.36
–
132.11 (5)
136.77 (2)
Appendix
Table 5: Fork-and-flush under three schedules, development runs ( K=4 ). We report the mean final score at the task’s budget, with the number of runs in parentheses. The single-run column is the mean of the eight baselines from the same period.
Figure 6: Figure 3 under two embeddings of the same judged triplets. Top: t-STE, as in Figure 3 . Bottom: SOE, aligned to the same frame, with basins recomputed on its coordinates; 99% of single-run artifacts keep their basin. In both, the rejected branches ( × ) stay in the branch point’s basin and within about one median single-run step of it, while the selected branch lands in a different basin: 43× the median single-run step under t-STE and 10× under SOE, in each case more than twice the largest step any single run takes.
t-STE
SOE
task
judged triplets
train
held out
train
held out
polyomino
6,375
0.90
0.87
0.90
0.85
AHC014
6,375
0.90
0.87
0.91
0.86
p171
6,375
0.88
0.85
0.88
0.84
p150
5,610
0.87
0.83
0.86
0.80
job-shop
6,375
0.86
0.81
0.85
0.79
Appendix
Table 6: Held-out validation of the idea-space maps. Agreement is the fraction of LLM triplet judgments the fitted two-dimensional map reproduces; chance is 0.50 . t-STE is the embedding used in the figures; SOE refits the same triplets with soft ordinal embedding.
Figure 7: Fork-and-flush trajectories against the single-run baselines of each task.
Figure 8: Nine fork-and-flush runs on their tasks’ idea maps (task-map basins shaded). Blue: the selected path ( ∘ opening round, ⋄ fork with a flush, ⋆ final artifact); vermillion × : rejected branches.
Department of Computer Science, Columbia University · 2AI, Analytics and Future of Work Initiative, Georgetown University · Department of Applied Mathematics, Columbia University +3