Organizations: East China Normal University · Alibaba Group · Fudan University · The University of Hong Kong · Tokentide AI · Shanghai Jiao Tong University · Stanford University
Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.
Figures & tables
Figure 1: Motivating evidence for strategy complementarity. Left: Pass@8 success sets of DeepSeek-V4-Pro on MIPLIB-NL under three strategies. The partial overlap indicates that different strategies solve complementary subsets of instances. Right: a single TSP instance solved successfully by all three strategies, illustrating that one optimization problem can admit multiple valid computational solution paths.
Figure 2: Overview of the file-grounded data construction pipeline.
Figure 3: Overview of Strategy-Diverse Reinforcement Learning (SDRL).
Method
Self-contained
File-grounded
Avg.
NL4Opt
MAMO-E
MAMO-C
Ind-OR
OptMath
OptiBench
MIPLIB-NL
Base Models
Qwen3-4B-Instruct-2507
72.7
68.9
42.4
40.0
17.5
54.1
7.7
43.3
Qwen3-32B
85.3
88.8
66.0
39.0
17.5
60.0
10.5
52.4
Fine-tuned Models
ORLM-Llama3-8B*
85.7
82.3
37.4
24.0
2.6
51.1
0.6
40.5
Table 1: Pass@1 accuracy across seven optimization benchmarks, with MIPLIB-NL representing a file-grounded, industrial-scale setting that requires agentic optimization modeling capabilities.
Method
MAMO-C
IndustryOR
OptMath
MIPLIB-NL
Average
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
Base model
42.4
76.9
40.0
62.0
17.5
36.8
7.7
19.1
26.9
48.7
SDRL w/o Rdiv
70.9
87.7
51.0
69.0
35.5
48.2
13.1
29.1
42.6
58.5
SDRL w/o Rmicro
73.9
88.7
50.0
72.0
37.4
53.6
14.6
31.8
44.0
61.5
SDRL w/o Rmacro
75.9
88.2
52.0
72.0
34.9
50.0
14.6
25.9
44.4
59.0
Full SDRL
75.9
90.2
53.0
71.0
36.7
54.8
16.8
33.2
45.6
62.3
Table 2: Ablation of the hierarchical diversity reward using text-only data.
Training Data
Text Avg.
MIPLIB-NL
P@1
P@8
P@1
P@8
Text-only
69.6
80.3
16.8
33.2
File-grounded
64.3
77.6
21.8
38.6
Mixed-format
71.4
79.5
23.6
35.9
Table 3: Ablation of the mixed-format training scheme.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Check
Pass rate
1
Structural and numerical integrity (deterministic)
88.2%
2
Source–conversion consistency (Claude-Sonnet-4.6)
86.2%
3
Solve-back reproduces the optimal value (Claude-Opus-4.8)
82.1%
Appendix
Table 4: Verification of converted file-grounded training instances. Each stage is applied to all instances that pass the preceding stage, and pass rates are conditional on the preceding stage. Only instances passing all three stages are used for training.
Benchmark
# Instances
NL4Opt
245
MAMO-EasyLP
642
MAMO-ComplexLP
203
IndustryOR
100
OptMATH-Bench
166
OptiBench
605
Appendix
Table 5: Number of validated problem instances used in each evaluation benchmark. For MIPLIB-NL, three instances with empty data files are excluded from the original 223.
Type
Parameter
Value
Model
Backbone (small)
Qwen3-4B-Instruct-2507
Backbone (large)
Qwen3-32B
Algorithm
Advantage estimator
GRPO
Training steps
400
Data
Batch size
64
Learning rate
1×10−6
Appendix
Table 6: Training parameters.
Method
Metric
NL4Opt
MAMO-E
MAMO-C
IndustryOR
OptMATH
OptiBench
MIPLIB-NL
Avg.
Qwen3-4B-Instruct
P@1
72.7
68.9
42.4
40.0
17.5
54.1
7.7
43.3
P@8
93.5
89.6
76.9
62.0
36.8
70.1
19.1
64.0
SDRL-Qwen3-4B
P@1
93.6
92.4
79.3
55.0
41.0
67.1
23.6
64.6
P@8
95.1
96.3
88.2
70.0
54.2
73.2
35.9
73.3
Qwen3-32B
P@1
85.3
88.8
66.0
39.0
17.5
60.0
10.5
52.4
P@8
95.5
96.3
86.2
67.0
42.2
70.3
27.7
69.3
Appendix
Table 7: Detailed Pass@1 and Pass@8 results (%) of the Base Model and SDRL using Qwen3-4B-Instruct-2507 and Qwen3-32B across seven benchmarks.
MAMO-C
IndustryOR
OptMath
MIPLIB-NL
Average
λ
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
0.0
75.9
88.2
52.0
72.0
34.9
50.0
14.6
25.9
44.4
59.0
0.1
72.9
88.2
52.0
72.0
36.8
52.4
15.9
32.3
44.4
61.2
0.3
74.9
88.2
53.0
69.0
34.3
53.0
14.6
33.2
44.2
60.9
0.5
74.4
88.2
50.0
70.0
33.1
53.4
14.6
32.7
43.0
61.1
0.7
75.9
90.2
53.0
71.0
36.7
54.8
16.8
33.2
45.6
62.3
Appendix
Table 8: Effect of the macro/micro weight λ on Qwen3-4B-Instruct-2507 (text-only training). λ=0 uses only Rmicro , λ=1 only Rmacro ; λ=0.7 is our default. Bold marks the best value in each column.
NL4Opt
MAMO-E
MAMO-C
IndustryOR
OptMath
OptiBench
MIPLIB-NL
Avg.
Method
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
P@1
P@8
Base (SIR)
91.8
95.5
89.1
94.9
41.9
66.0
48.0
63.0
12.7
32.5
62.5
70.1
8.6
25.8
50.7
64.0
Base (Meta)
72.7
93.5
68.9
89.6
42.4
76.9
40.0
62.0
17.5
36.8
54.1
70.1
7.7
19.1
43.3
64.0
SIR-RL
94.7
95.9
92.8
96.6
70.4
85.7
48.0
65.0
36.1
48.2
66.8
69.6
14.1
28.2
60.4
69.9
SDRL (Ours)
93.9
95.5
92.2
97.4
75.9
90.2
53.0
71.0
36.7
54.8
65.8
72.6
16.8
33.2
62.0
73.5
Appendix
Table 9: Pass@1 and Pass@8 accuracy (%) for disentangling the effects of strategy routing and reinforcement learning. Base (SIR) and Base (Meta) evaluate Qwen3-4B-Instruct-2507 under the rigid SIR Prompt and the Meta Prompt, respectively. SIR-RL applies reinforcement learning while remaining restricted to the SIR paradigm, whereas SDRL (Ours) enables strategy-diverse routing under the Meta Prompt.
Base model
SDRL
Benchmark
n
Median
P90
n
Median
P90
NL4Opt
1443
0.013
0.05
1743
0.020
0.10
MAMO-E
3494
0.020
0.10
4627
0.061
0.14
MAMO-C
689
0.022
0.12
1270
0.083
0.16
IndustryOR
314
0.065
0.19
445
0.090
0.21
OptMath
212
0.018
0.10
498
0.032
0.11
Appendix
Table 10: Runtime (s) of correct programs on the Pass@8 rollouts. n is the number of correct programs; Median and P90 are the median and 90th percentile of their wall-clock execution time.
Strategy
Representative procedure
Objective
SIR
MILP with MTZ subtour elimination, solved with Gurobi
310.6
Exact Combinatorial Algorithm
Exhaustive enumeration of the 4! tours from a fixed start
310.6
Heuristic Search
Simulated annealing with 2-opt moves
310.6
Appendix
Table 11: Three complementary solution strategies for the same five-exhibit traveling salesperson instance.
East China Normal University, Shanghai 200062, China · AntGroup, Hangzhou 310000, China · Southern University of Science and Technology, Shenzhen 518055, China +1