Organizations: Chengdu University of Information Technology · Eindhoven University of Technology · Nanyang University of Technology · Shandong University · Singapore Management University
Despite recent progress in autoresearch, applying it to practical operations research problems, typically formulated as NP-hard mixed-integer linear or nonlinear programs (MILPs or MINLPs), remains challenging because effective research requires systematically managing competing ideas and long-horizon experimental trajectories. We introduce AutoMIP, a reusable agent skill for organizing long-horizon autoresearch in mixed-integer programming through idea pooling and algorithm tree search. AutoMIP maintains a persistent pool of complementary candidate ideas while organizing executable experiments into an algorithm tree, enabling the agent to preserve unexplored hypotheses, refine promising algorithms, and switch to alternative methodological directions based on historical states. On MILP and MINLP benchmark cohorts, AutoMIP achieves the highest final success rates among the evaluated autoresearch frameworks. On MIPLib, AutoMIP discovers new best solutions for 31 of 60 instances, surpassing existing autoresearch frameworks. On MINLPLib, it achieves new best solutions for 52 of 60 instances. Ablation studies further demonstrate the complementary contributions of idea pooling and algorithm tree search, highlighting the importance of jointly maintaining diverse research ideas and structured experimental trajectories for long-horizon autoresearch.
Figures & tables
Figure 1: AutoMIP workflow. AutoMIP preserves candidate hypotheses in idea pool, selects refinements or new branches (research directions) from reusable experimental states in algorithm tree, and updates idea pool and algorithm tree from experimental feedback. The workflow supports MILP and MINLP optimization, with MPS taken as an example input format.
Figure 2: Coupling a direction to its starting state. After n0 → n1 → n4, selecting P5 and reusing n0 creates the R /P5 branch n3. Solid edges record ancestry; dashed arc restores historical artifacts.
Input: Exploration state Mt=(Tt,Ft,Pt,ct,At) .
Output: Registered experiment and restored experimental state.
Read the current node ct , frontier nodes Ft , candidate pool Pt , and artifacts At .
if An admissible refinement exists:
Set dt=L and pt=ct .
else if an unused candidate is available:
Sample qt∼Uniform(Ut) .
Algorithm 1 Algorithm Tree Search
Method
0–1 h
1–3 h
3–6 h
6–12 h
Total
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
AutoMIP
14/24
1,211
320,244
4/5
5,545
912,789
0/2
11,256
1,960,932
0/0
43,200
961,095
18/31
22,202
734,057
Codex
1/13
1,822
393,785
3/11
6,426
1,070,770
0/4
13,980
2,857,174
1/1
42,049
1,137,639
5/29
24,931
1,078,847
Loop
3/15
1,888
323,983
1/6
7,123
648,990
0/3
11,665
1,685,786
0/1
42,679
1,700,813
4/25
27,375
1,310,672
AutoEoH
0/12
2,157
606,353
3/9
7,142
1,035,190
0/4
14,175
2,556,814
0/1
43,162
2,660,954
3/26
27,626
1,999,227
EvoX
0/6
1,302
306,763
0/5
7,032
720,454
0/3
13,390
1,586,613
1/1
42,975
1,244,128
1/15
34,333
2,297,150
Table 1: Open cohort, 60 instances.
Method
0–10 min
10–20 min
20–30 min
30–60 min
Total
solved
time (s)
token
solved
time (s)
token
solved
time (s)
token
solved
time (s)
token
solved
time (s)
token
AutoMIP
14
399
96,926
8
869
120,307
5
1,511
154,571
3
3,454
171,106
30
2,035
145,646
Codex
6
322
196,416
6
909
116,999
5
1,471
153,526
12
2,912
206,999
29
2,333
192,485
Loop
4
518
80,309
8
873
102,209
4
1,364
135,665
6
3,218
228,227
22
2,602
195,393
AutoEoH
2
478
97,654
6
923
110,134
4
1,471
142,942
7
2,943
182,870
19
2,561
170,094
EvoX
3
574
117,775
6
961
101,781
4
1,546
203,375
2
2,894
254,562
15
2,495
229,032
Table 2: Hard cohort, 60 instances and a one-hour horizon.
Figure 3: Cumulative task success from per-instance timestamps. Each step adds a successful instance; the denominator is 60 in every panel. Curves extend to 43,200 s for Open and MINLPLib and 3,600 s for Hard. The cutoffs agree with Tables 1 , 2 , and 4 , respectively.
Method
0–1 h
1–3 h
3–6 h
6–12 h
Total
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
AutoMIP
5/8
1,440
317,315
3/5
5,545
912,789
1/2
11,256
1,960,932
0/0
43,200
2,076,973
9/15
13,888
1,070,460
NoPool
1/5
2,019
233,353
2/4
5,643
450,314
1/2
11,460
1,168,343
0/1
42,334
2,874,266
4/12
21,829
1,558,655
NoTree
1/3
1,304
202,388
1/6
5,841
518,799
0/1
13,560
604,118
0/0
43,200
2,935,203
2/10
24,226
1,683,805
Table 3: Ablations on 20 shared instance identities.
Method
0–1 h
1–3 h
3–6 h
6–12 h
Total
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
AutoMIP
13/40
1,076
203,122
3/10
5,895
623,950
0/1
19,140
1,051,147
0/1
41,089
3,376,242
16/52
8,182
763,362
Codex
9/35
1,219
200,220
2/10
5,501
409,595
0/0
-
-
0/0
43,200
2,927,059
11/45
12,428
916,826
Loop
8/32
1,553
262,392
1/6
6,437
465,120
0/2
18,703
1,779,411
0/0
43,200
3,537,847
9/40
16,495
1,425,050
AutoEoH
7/25
933
193,259
1/4
6,538
546,097
0/0
-
-
0/0
43,200
3,513,110
8/29
23,145
1,932,038
EvoX
8/31
1,111
218,111
0/0
-
-
0/0
-
-
0/0
43,200
3,241,573
8/31
21,454
1,679,451
Table 4: MINLPLib, the supplied common 60-instance subset.
Method
0–1 h
1–3 h
3–6 h
6–12 h
Total
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
w/s
time (s)
token
Codex
6/9
1,385
312,966
2/5
5,089
944,082
1/1
11,085
1,688,996
0/0
43,200
4,641,312
9/15
13,250
1,621,633
Claude Code
2/7
1,819
305,213
2/3
6,326
1,182,345
2/2
13,433
1,679,399
0/0
43,200
4,680,348
6/12
20,209
2,324,256
Table 5: Cross-agent evaluation on 20 instances. Both agents use AutoMIP.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Mechanism
NoPool
NoTree
PoolThink and persistent candidate pool
Removed
Retained
On-demand direction generation
Retained
Not used
Random selection from unused candidates
Removed
Retained
Multiway tree and parent links
Retained
Removed
L/R labels and frontier
Retained
Removed
Historical-state selection and backtracking
Retained
Removed
Appendix
Table 6: Components retained and removed in the two complementary ablations.
Figure 4: Instance-level convergence over ten trajectories per instance: five random seeds for each of the two methods. Lines show each method’s pointwise median; shaded bands show its interquartile range (25th–75th percentiles) across the five seeds.
Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations. To address these challenges, we introduce the Agentic Idea Manager (AIM), a fully autonomous framework for managing and exploring research directions in idea-driven automated research. Inspired by Bayesian optimization, AIM uses an Agentic Surrogate and an Agentic Acquisition mechanism to organize discovered ideas and guide their selection. A Solution Auditor maintains idea-solution integrity, while a Resource Planner adaptively allocates the remaining experimental budget across parallel search branches. Experiments on 10 AutoLab benchmark tasks show that AIM surpasses the strongest baseline by 1.6 percentage points on System Optimization tasks and 4.9 percentage points on long-horizon Model Development & CUDA tasks. Notably, AIM reaches the best baseline performance up to 3.1x faster in wall-clock time. We further provide a theoretical analysis of when searching over ideas becomes beneficial. Our analysis shows that explicit idea-level allocation makes semantic coverage directly controllable, and that broader coverage becomes increasingly valuable when competitive research directions are sparse among many plausible alternatives. Project Page: https://imhgchoi.github.io/agentic-idea-manager/
Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged by its terminal pipeline. A terminal score cannot reveal which technical decision produced a gain or distinguish a reusable discovery from a change adapted to development feedback. We introduce intervention-centered Auto Research, which validates research decisions rather than only final artifacts and makes their reliability measurable. Feature, Model, Representation, and Data axes are searched independently with inner five-fold feedback. Each axis winner is frozen before an outer-holdout matrix compares all alternatives on evidence the loop never sees. Across 701 agent-executed attempts spanning ten Matbench endpoints, outer evidence confirms the selected intervention on nine of ten endpoints and preserves 89.3% of non-tied intervention orderings. It also rejects an aggregate Representation gain that inner feedback endorsed. The resulting matrix reveals an information-dependent hierarchy. Composition-only tasks support several routes to improvement, whereas structure-informed tasks favor local geometry features and complementary tree ensembles. A subsequent compatibility test combines already frozen Feature and Model code without further search or tuning and raises mean outer-holdout improvement from 19.0% to 26.3%. By validating decisions rather than only artifacts, this design turns adaptive search into reusable evidence wherever agents propose executable alternatives against a fixed evaluator.
Jingjie Ning, Xiaochuan Li, Shanshan Zhong +2
School of Computer Science, Carnegie Mellon University · 2DP Technology
Mixed-integer programming (MIP) research is both mathematically sophisticated and engineering-intensive: testing an algorithmic hypothesis within a branch-and-cut solver requires substantial implementation, debugging, tuning, and large-scale benchmarking. We propose an agentic MIP research framework that shortens this feedback loop by embedding LLM agents into a solver-aware harness for generating, verifying, and evaluating plugins for the open-source solver SCIP. Propagation methods play a central role in accelerating MIP solving by exploiting global constraints. We instantiate our framework on the semantic lifting of MIP formulations into global constraints and the automatic construction of propagation-only SCIP constraint handlers. On the MIPLIB 2017 benchmark set, the framework successfully recovers global constraint structures from constraint programming and generates executable constraint detectors and propagation-only constraint handlers. Furthermore, the framework naturally extends to in-context learning within a sandboxed environment, enabling agents not only to tune and debug generated constraint handlers on real instances, but also to explore global constraint patterns in MIP problems and discover novel propagation strategies not yet implemented in SCIP. This framework allows us to systematically distinguish meaningful algorithmic improvements from low-value or overly costly candidates: the novel propagation methods successfully solved five additional instances within the explored benchmark. Overall, this framework demonstrates that LLM agents can autonomously navigate the complex MIP research loop, paving the way for a more automated solver development process.
Liding Xu, Yugeng Zhou, Sebastian Pokutta
Zuse Institute Berlin, Berlin, Germany. · Independent Researcher. · Zuse Institute Berlin +1