Organizations: School of Computer Science and Engineering, Beihang University, Beijing, China · MIIT Key Laboratory of Data and Decision Intelligence, Beihang University, Beijing, China · School of Data Science, The Chinese University of Hong Kong, Shenzhen, China · Cardinal Operations Technology Co., Shanghai, China
Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, but they remain prone to errors. In practice, these errors can be divided into two categories: syntactic errors refer to solver code that fails to run successfully or is judged infeasible by the solver; semantic errors refer to solver code that successfully returns an objective value but violates the intent of the original problem. Since semantic errors do not trigger runtime failures, they are difficult to detect and rectify. To address this problem, we introduce SemOPT, a semantic-guided framework for correcting LLM-based optimization models. SemOPT combines a semantic reward model that distinguishes faithful math models from plausible but incorrect ones with an adaptive correction system that applies hierarchical reward-guided search over the modeling space. Experiments on seven optimization modeling benchmarks show that SemOPT establishes a new state of the art and achieves an average 7.6% accuracy improvement over the strongest baseline on complex datasets.
Figures & tables
Figure 1: (a) Distribution of syntactic versus semantic errors. (b)(c) Correction rates of syntactic and semantic errors via self-reflection.
Figure 2: Overview of the three stages in optimization modeling process.
Figure 3: Overview of the SemOPT framework. (a) Training pipeline of the semantic reward model. (b) Inference architecture of the adaptive correction system.
Category
Method
Standard Datasets
Complex Datasets
NL4Opt
NL4LP
EasyLP
ReSocratic
IndustryOR
ComplexLP
ComplexOR
Standard
Prompting
65.4%
75.8%
87.2%
67.5%
47.6%
41.4%
44.4%
Reflection
65.0%
76.4%
87.2%
68.0%
54.8%
39.6%
38.9%
Workflow
OptiMUS
79.0%
90.5%
93.6%
78.7%
57.1%
45.1%
59.3%
Chain-of-Experts
74.8%
90.5%
93.2%
79.4%
59.5%
46.0%
50.0%
OptiTree
73.8%
82.0%
93.4%
72.5%
47.6%
73.8%
63.0%
Table 1: Main results on seven optimization benchmarks. The best results are highlighted in bold , and the second-best results are underlined . Red arrows indicate absolute performance gains over the strongest baseline.
Dataset
Difficulty
SIRL
SemOPT
NL4Opt
Easy
93.6%
94.4% ( ↑ 0.8%)
Medium
82.0%
84.0% ( ↑ 2.0%)
Hard
69.2%
71.8% ( ↑ 2.6%)
ComplexLP
Easy
88.6%
88.6% (+0.0%)
Medium
94.1%
97.1% ( ↑ 3.0%)
Hard
54.5%
63.6% ( ↑ 9.1%)
Table 2: Accuracy (Best-of-N) across difficulty levels. We compare SemOPT against the single strongest baseline (SIRL) on Easy, Medium, and Hard subsets.
Backbone LLM
Method
NL4Opt
ComplexLP
ComplexOR
Gemini-2.5 Flash-Lite
Prompting (BoN)
70.1%
64.0%
55.6%
SemOPT (BoN)
79.0%
75.7%
63.0%
GPT-4.1 mini
Prompting (BoN)
74.8%
64.8%
55.6%
SemOPT (BoN)
86.9%
79.3%
66.7%
Qwen3-Max
Prompting (BoN)
76.2%
66.7%
55.6%
SemOPT (BoN)
88.8%
82.0%
66.7%
Table 3: Performance across LLM backbones. We compare SemOPT with standard prompting in BoN settings.
Variant
IndustryOR
ComplexLP
ComplexOR
ORLM
42.9%
59.5%
50.0%
SemOPT w/ ORLM generator
61.9%
74.8%
66.7%
SIRL
45.2%
80.2%
44.4%
SemOPT w/ SIRL generator
66.7%
85.6%
66.7%
Table 4: Fine-tuned generator control on complex datasets. We replace only SemOPT ’s generator with the fine-tuned models of ORLM and SIRL.
Table 6: Factorizing the semantic evaluator and the correction hierarchy on complex datasets (BoN accuracy).
Figure 4: (a) Breakdown of model outputs under different critic models. (b) Efficiency evaluation of the adaptive gating mechanism.
Method
Acc.
Time (s)
API Calls
SR (25 iters)
55.7%
61.2
26.0
BS (beam 2, branch 5)
63.3%
38.9
30.0
SemOPT
67.1 %
32.3
25.1
Table 7: Efficiency comparison with iterative and search baselines on the combined hard subsets of NL4Opt, ComplexLP, and ComplexOR.
Method
NL4Opt
ComplexLP
ComplexOR
NS
89.7% / 84.0
83.8% / 84.0
66.7% / 84.0
SemOPT
87.9% / 10.2
82.9% / 13.6
66.7% / 15.9
Table 8: Comparison with exhaustive Naive Search (NS). Each cell reports accuracy / API calls.
Figure 5: Comparison between SemOPT and the standard prompting baseline across optimization benchmarks.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Operator
Implementation Logic
Variable
Type Relax
Relax variable types (e.g., BINARY → CONTINUOUS)
Bound Remove
Remove lower/upper bounds
Objective
Sense Flip
Swap direction (e.g., MIN → MAX)
Coeff Noise
Flip sign of variable coefficient in objective
Term Drop
Remove a term from objective expression
Constraint
Comparator Flip
Swap inequality signs (e.g., ≥→≤ )
Appendix
Table 9: Semantic mutation operators for hard negative construction.
Parameter
Value
Generator (LLM) Settings
Temperature
0.9
Top-p
0.9
Max Tokens
2048
MCTS Settings
Exploration Constant ( ω )
1.414
Appendix
Table 10: Hyperparameter configurations for SemOPT .
Base Model
Train / Val Problems
Pairs
Epochs
Batch
LR
Max Len.
Qwen3-4B-Instruct-2507
2700 / 300
8254
3
32
2e−5
2048
Appendix
Table 11: Detailed training configuration for the semantic reward model.
Method
GED Acc.
Oracle Acc.
Obj Acc.
SIRL
36%
36%
38%
SemOPT
44%
47%
49%
Appendix
Table 12: Formulation-level validation on OptiVerse-LP-100.
Dataset
τ=0.2
τ=0.4
τ=0.6
τ=0.8
τ=1.0
NL4Opt
72.9% / 4.1
79.0% / 8.4
87.9% / 10.2
88.8% / 19.6
88.8% / 26.8
ComplexLP
53.2% / 4.6
73.8% / 10.2
82.9% / 13.6
83.8% / 20.1
81.1% / 27.4
ComplexOR
44.4% / 5.8
63.0% / 13.1
66.7% / 15.9
66.7% / 22.3
66.7% / 29.5
Appendix
Table 13: Sensitivity to the gating threshold τ . Each cell reports accuracy / API calls.
Figure 6: Impact of node expansion size H on SemOPT performance across three benchmarks.
Hard Subset
SIRL
SemOPT
NL4Opt
69.2% ± 2.56%
71.8% ± 4.44%
ComplexLP
54.5% ± 5.25%
63.6% ± 3.03%
ComplexOR
28.6% ± 0.00%
57.1% ± 0.00%
Appendix
Table 14: Hard-subset accuracy with variation.
Subset
SIRL
SemOPT
SemOPT -only / SIRL-only
Gain
90% CI / p
NL4Opt-hard
27/39
28/39
1/0
+2.6%
[−2.3%,2.6%] / 1.000
ComplexLP-hard
18/33
21/33
4/1
+9.1%
[−4.8%,14.8%] / 0.375
ComplexOR-hard
2/7
4/7
2/0
+28.5%
[−15.8%,28.6%] / 0.500
Combined hard
47/79
53/79
7/1
+7.6%
[0.6%,10.0%] / 0.070
Appendix
Table 15: Paired analysis on hard subsets. Counts are the number of solved instances over the subset size.
Dataset
Syntax-error Rate
Correctly Repaired
Still Erroneous (Score 0)
IndustryOR
33.1%
64.1%
35.9%
ComplexLP
34.2%
74.7%
25.3%
ComplexOR
46.8%
81.8%
18.2%
Appendix
Table 16: Syntactic error handling during search. The error rate is computed over all evaluated nodes; repair outcomes are computed over the nodes that initially triggered repair.
Error Category
Count
Share
By hierarchy layer (40 errors)
Modeling Strategy
9
22.5%
Math Model
28
70.0%
Solver Code
3
7.5%
By formulation component (28 math-model errors)
Parameter
2
7.1%
Appendix
Table 17: Distribution of the 40 remaining SemOPT errors by hierarchy layer, and of the 28 math-model errors by formulation component.
Figure 7: Case study on IndustryOR-37.
Figure 8: Full prompt for the modeling-strategy layer of MCTS.
Figure 9: Full prompt for the math model layer of MCTS.
Figure 10: Full prompt for the solver-code layer of MCTS.
Building mathematical optimization models is critical in operations research (OR), while it requires substantial human expertise. Recent advancements have utilized large language models (LLMs) to automate this modeling process. However, existing works often struggle to verify the correctness of the generated optimization models, without checking the rationality of the constraints and variables or the validity of solutions to the generated models. This hampers the subsequent verification and correction steps, and thus it severely hurts the modeling accuracy. To address this challenge, we propose a novel LLM-based framework with Dual-side Verification (Opt-Verifier) from both structure and solution perspectives, thereby improving the modeling accuracy. The structure-side verification ensures that the modeling structure of the generated optimization models aligns with the original problem description, accurately capturing the problem's constraints and requirements. Meanwhile, the solution-side verification interprets and evaluates the solutions' validity, confirming that the optimization models are logically and mathematically sound. Experiments on popular benchmarks demonstrate that our approach achieves over 20% improvement in accuracy.
Haoyang Liu, Jie Wang, Boxuan Niu +8
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China · Noah’s Ark Lab, Huawei Technologies · Tianjin University
Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiting practical scalability. Therefore, we systematically investigate whether LLMs can identify problem structure from natural-language descriptions and apply suitable optimization modeling techniques to generate mathematical models and solver code that solve the problems correctly and efficiently. To this end, we first curate OptTips, a knowledge base of 50 expert modeling techniques in eight families. Using this knowledge, we develop OptDachshund, a multi-agent framework that transforms problems from existing optimization benchmarks into new tasks for evaluating LLMs' use of modeling techniques. It constructs conventional and expert mathematical models with solver code for the same task and data, providing baselines for correctness and computational cost. The resulting EfficientOpt benchmark contains 561 expert-reviewed tasks with paired reference implementations. Evaluation of 11 representative LLMs reveals an efficiency gap on correctly solved tasks with comparable measurements: for every LLM, most generated programs take longer to solve than their expert counterparts. Within the comparable reference-size subset, 57% of programs with correct objective values and fewer variables and linear constraints have longer recorded solver times. Case studies show that different modeling techniques can achieve the same optimal value at similar recorded cost. Faster solving may not reduce execution time if the code takes longer to prepare data and build the model. LLM optimization modeling should therefore be evaluated for both correctness and computational efficiency.
Zhong Li, Xin Huang, Jinhui Wan +6
Great Bay University · Beijing Jiaotong University · Beihang University +1
Large language models (LLMs) can generate syntactically valid optimization programs, yet often struggle to reliably choose an effective modeling strategy, leading to incorrect formulations and inefficient solver behavior. We propose SAGE, a strategy-aware framework that makes Modeling Strategy explicit in both data construction and post-training. SAGE builds a solver-verified multi-strategy dataset and trains a student model with supervised fine-tuning followed by Segment-Weighted GRPO using a composite reward over format compliance, correctness, and solver efficiency. Across eight benchmarks spanning synthetic and real-world settings, SAGE improves average pass@1 from 72.7 to 80.3 over the strongest open-source baseline. With multiple generations, SAGE discovers more distinct correct formulations and improves component-level diversity at pass@16 by 19-29%. At the largest scale, SAGE produces more compact constraint systems with 14.2% fewer constraints than the baseline, consistent with solver-efficient modeling. Overall, these results show that making Modeling Strategy explicit improves automated optimization modeling. Code is available at https://github.com/rachhhhing/SAGE.
Ruiqing Zhao, Fengzhi Li, Yuan Zuo +5
School of Computer Science and Engineering, Beihang University, Beijing, China · JIUTIAN Research, Beijing, China · MIIT Key Laboratory of Data and Decision Intelligence, Beihang University, Beijing, China