Organizations: East China Normal University · Alibaba Group · Fudan University · The University of Hong Kong · Tokentide AI · Shanghai Jiao Tong University · Stanford University
Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.
Figures & tables
Figure 1: Motivating evidence for strategy complementarity. Left: Pass@8 success sets of DeepSeek-V4-Pro on MIPLIB-NL under three strategies. The partial overlap indicates that different strategies solve complementary subsets of instances. Right: a single TSP instance solved successfully by all three strategies, illustrating that one optimization problem can admit multiple valid computational solution paths.
Figure 2: Overview of the file-grounded data construction pipeline.
Figure 3: Overview of Strategy-Diverse Reinforcement Learning (SDRL).
Method
Self-contained
File-grounded
Avg.
NL4Opt
MAMO-E
MAMO-C
Ind-OR
OptMath
OptiBench
MIPLIB-NL
Base Models
Qwen3-4B-Instruct-2507
72.7
68.9
42.4
40.0
17.5
54.1
7.7
43.3
Qwen3-32B
85.3
88.8
66.0
39.0
17.5
60.0
10.5
52.4
Fine-tuned Models
ORLM-Llama3-8B*
85.7
82.3
37.4
24.0
2.6
51.1
0.6
40.5
Table 1: Pass@1 accuracy across seven optimization benchmarks, with MIPLIB-NL representing a file-grounded, industrial-scale setting that requires agentic optimization modeling capabilities.
Method
MAMO-C
IndustryOR
OptMath
MIPLIB-NL
Average
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
Base model
42.4
76.9
40.0
62.0
17.5
36.8
7.7
19.1
26.9
48.7
SDRL w/o Rdiv
70.9
87.7
51.0
69.0
35.5
48.2
13.1
29.1
42.6
58.5
SDRL w/o Rmicro
73.9
88.7
50.0
72.0
37.4
53.6
14.6
31.8
44.0
61.5
SDRL w/o Rmacro
75.9
88.2
52.0
72.0
34.9
50.0
14.6
25.9
44.4
59.0
Full SDRL
75.9
90.2
53.0
71.0
36.7
54.8
16.8
33.2
45.6
62.3
Table 2: Ablation of the hierarchical diversity reward using text-only data.
Training Data
Text Avg.
MIPLIB-NL
P@1
P@8
P@1
P@8
Text-only
69.6
80.3
16.8
33.2
File-grounded
64.3
77.6
21.8
38.6
Mixed-format
71.4
79.5
23.6
35.9
Table 3: Ablation of the mixed-format training scheme.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Check
Pass rate
1
Structural and numerical integrity (deterministic)
88.2%
2
Source–conversion consistency (Claude-Sonnet-4.6)
86.2%
3
Solve-back reproduces the optimal value (Claude-Opus-4.8)
82.1%
Appendix
Table 4: Verification of converted file-grounded training instances. Each stage is applied to all instances that pass the preceding stage, and pass rates are conditional on the preceding stage. Only instances passing all three stages are used for training.
Benchmark
# Instances
NL4Opt
245
MAMO-EasyLP
642
MAMO-ComplexLP
203
IndustryOR
100
OptMATH-Bench
166
OptiBench
605
Appendix
Table 5: Number of validated problem instances used in each evaluation benchmark. For MIPLIB-NL, three instances with empty data files are excluded from the original 223.
Type
Parameter
Value
Model
Backbone (small)
Qwen3-4B-Instruct-2507
Backbone (large)
Qwen3-32B
Algorithm
Advantage estimator
GRPO
Training steps
400
Data
Batch size
64
Learning rate
1×10−6
Appendix
Table 6: Training parameters.
Method
Metric
NL4Opt
MAMO-E
MAMO-C
IndustryOR
OptMATH
OptiBench
MIPLIB-NL
Avg.
Qwen3-4B-Instruct
P@1
72.7
68.9
42.4
40.0
17.5
54.1
7.7
43.3
P@8
93.5
89.6
76.9
62.0
36.8
70.1
19.1
64.0
SDRL-Qwen3-4B
P@1
93.6
92.4
79.3
55.0
41.0
67.1
23.6
64.6
P@8
95.1
96.3
88.2
70.0
54.2
73.2
35.9
73.3
Qwen3-32B
P@1
85.3
88.8
66.0
39.0
17.5
60.0
10.5
52.4
P@8
95.5
96.3
86.2
67.0
42.2
70.3
27.7
69.3
Appendix
Table 7: Detailed Pass@1 and Pass@8 results (%) of the Base Model and SDRL using Qwen3-4B-Instruct-2507 and Qwen3-32B across seven benchmarks.
MAMO-C
IndustryOR
OptMath
MIPLIB-NL
Average
λ
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
0.0
75.9
88.2
52.0
72.0
34.9
50.0
14.6
25.9
44.4
59.0
0.1
72.9
88.2
52.0
72.0
36.8
52.4
15.9
32.3
44.4
61.2
0.3
74.9
88.2
53.0
69.0
34.3
53.0
14.6
33.2
44.2
60.9
0.5
74.4
88.2
50.0
70.0
33.1
53.4
14.6
32.7
43.0
61.1
0.7
75.9
90.2
53.0
71.0
36.7
54.8
16.8
33.2
45.6
62.3
Appendix
Table 8: Effect of the macro/micro weight λ on Qwen3-4B-Instruct-2507 (text-only training). λ=0 uses only Rmicro , λ=1 only Rmacro ; λ=0.7 is our default. Bold marks the best value in each column.
NL4Opt
MAMO-E
MAMO-C
IndustryOR
OptMath
OptiBench
MIPLIB-NL
Avg.
Method
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
Base (SIR)
91.8
95.5
89.1
94.9
41.9
66.0
48.0
63.0
12.7
32.5
62.5
70.1
8.6
25.8
50.7
64.0
Base (Meta)
72.7
93.5
68.9
89.6
42.4
76.9
40.0
62.0
17.5
36.8
54.1
70.1
7.7
19.1
43.3
64.0
SIR-RL
94.7
95.9
92.8
96.6
70.4
85.7
48.0
65.0
36.1
48.2
66.8
69.6
14.1
28.2
60.4
69.9
SDRL (Ours)
93.9
95.5
92.2
97.4
75.9
90.2
53.0
71.0
36.7
54.8
65.8
72.6
16.8
33.2
62.0
73.5
Appendix
Table 9: Pass@1 and Pass@8 accuracy (%) for disentangling the effects of strategy routing and reinforcement learning. Base (SIR) and Base (Meta) evaluate Qwen3-4B-Instruct-2507 under the rigid SIR Prompt and the Meta Prompt, respectively. SIR-RL applies reinforcement learning while remaining restricted to the SIR paradigm, whereas SDRL (Ours) enables strategy-diverse routing under the Meta Prompt.
Base model
SDRL
Benchmark
n
Median
P90
n
Median
P90
NL4Opt
1443
0.013
0.05
1743
0.020
0.10
MAMO-E
3494
0.020
0.10
4627
0.061
0.14
MAMO-C
689
0.022
0.12
1270
0.083
0.16
IndustryOR
314
0.065
0.19
445
0.090
0.21
OptMath
212
0.018
0.10
498
0.032
0.11
Appendix
Table 10: Runtime (s) of correct programs on the Pass@8 rollouts. n is the number of correct programs; Median and P90 are the median and 90th percentile of their wall-clock execution time.
Strategy
Representative procedure
Objective
SIR
MILP with MTZ subtour elimination, solved with Gurobi
310.6
Exact Combinatorial Algorithm
Exhaustive enumeration of the 4! tours from a fixed start
310.6
Heuristic Search
Simulated annealing with 2-opt moves
310.6
Appendix
Table 11: Three complementary solution strategies for the same five-exhibit traveling salesperson instance.
Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating complicated descriptions of these problems into solver-ready formulations requires specialized operations research (OR) expertise, making it hard to scale. We present AutoOR, a scalable synthetic data generation and reinforcement learning pipeline that trains LLMs to autoformalize optimization problems specified in natural language across linear, mixed-integer, and non-linear categories. AutoOR generates verified training data from standard optimization forms and uses solver execution feedback as the reward signal for RL post-training. AutoOR applied to an 8B model achieves state-of-the-art or competitive results across six established OR benchmarks, matching significantly larger frontier models. For a non-linear problem class involving physical dynamics, where frontier models score near 0%, we introduce a curriculum RL strategy that bootstraps from limited initial training data to make this class tractable for post-training. We believe that methods such as AutoOR can significantly accelerate industrial decision-making with AI.
Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov +4
Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotated or teacher-generated datasets. This dependence introduces a Generalization Ceiling, where models overfit to narrow data distributions, and Capability Anchoring, where models' reasoning is bounded by annotator proficiency and teacher model capability. In response, we propose OPT-Zero, the first fully self-play training framework for optimization modeling that requires zero external training data. OPT-Zero employs a single LLM in a dual-role closed loop: a Proposer that synthesizes increasingly challenging optimization problems alongside their mathematical formulations and solving code, and a Solver that attempts to resolve the problems given only natural-language problem descriptions. Grounded in execution feedback from external optimization solvers, we alternately train both roles using reinforcement learning. This process fosters an auto-curriculum in which the Proposer and Solver co-evolve: generating harder valid problems by the Proposer seamlessly enhances the structural reasoning ability of the Solver. Extensive results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.
Xia Jiang, Yaoxin Wu, Chenyu Zhou +3
Eindhoven University of Technology · Shanghai Jiaotong University
Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a challenging problem for optimization-oriented large language models (LLMs). Existing approaches typically rely on large-scale supervised datasets, costly reasoning annotations, and expensive intermediate step verification, resulting in substantial training overhead. To address these challenges, we propose MiniOpt, a reinforcement learning framework that learns to solve optimization problems through an "reasoning-to-model-and-solve" paradigm. MiniOpt decomposes optimization reasoning into structured optimization modeling and executable solver generation. Building upon this paradigm, we introduce OptReward, a reward function with hierarchical score structure that jointly evaluates formulation and solution, enabling effective policy learning without expert demonstrations. We further develop an optimization-oriented policy optimization strategy that improves exploration efficiency and stabilizes reinforcement learning for compact models. Extensive experiments show that MiniOpt-3B exhibits strong optimization generalization across various optimization types, problem scenarios, and task domains. For models with fewer than 10B parameters, MiniOpt series achieves the highest average solving accuracy (SA). For models with more than 10B parameters, MiniOpt still shows competitive performance. These results suggest that optimization-oriented reward design and reinforcement learning provide an effective pathway for developing compact optimization-specialized language models with strong optimization generalization capabilities. The code is available at https://github.com/Hsiang-1/MiniOpt.
Ke Zhao, Zixiang Di, Hong Qian +9
East China Normal University, Shanghai 200062, China · AntGroup, Hangzhou 310000, China · Southern University of Science and Technology, Shenzhen 518055, China +1