Organizations: School of Computer Science and Engineering, Beihang University, Beijing, China · MIIT Key Laboratory of Data and Decision Intelligence, Beihang University, Beijing, China · School of Data Science, The Chinese University of Hong Kong, Shenzhen, China · Cardinal Operations Technology Co., Shanghai, China
Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, but they remain prone to errors. In practice, these errors can be divided into two categories: syntactic errors refer to solver code that fails to run successfully or is judged infeasible by the solver; semantic errors refer to solver code that successfully returns an objective value but violates the intent of the original problem. Since semantic errors do not trigger runtime failures, they are difficult to detect and rectify. To address this problem, we introduce SemOPT, a semantic-guided framework for correcting LLM-based optimization models. SemOPT combines a semantic reward model that distinguishes faithful math models from plausible but incorrect ones with an adaptive correction system that applies hierarchical reward-guided search over the modeling space. Experiments on seven optimization modeling benchmarks show that SemOPT establishes a new state of the art and achieves an average 7.6% accuracy improvement over the strongest baseline on complex datasets.
Figures & tables
Figure 1: (a) Distribution of syntactic versus semantic errors. (b)(c) Correction rates of syntactic and semantic errors via self-reflection.
Figure 2: Overview of the three stages in optimization modeling process.
Figure 3: Overview of the SemOPT framework. (a) Training pipeline of the semantic reward model. (b) Inference architecture of the adaptive correction system.
Category
Method
Standard Datasets
Complex Datasets
NL4Opt
NL4LP
EasyLP
ReSocratic
IndustryOR
ComplexLP
ComplexOR
Standard
Prompting
65.4%
75.8%
87.2%
67.5%
47.6%
41.4%
44.4%
Reflection
65.0%
76.4%
87.2%
68.0%
54.8%
39.6%
38.9%
Workflow
OptiMUS
79.0%
90.5%
93.6%
78.7%
57.1%
45.1%
59.3%
Chain-of-Experts
74.8%
90.5%
93.2%
79.4%
59.5%
46.0%
50.0%
OptiTree
73.8%
82.0%
93.4%
72.5%
47.6%
73.8%
63.0%
Table 1: Main results on seven optimization benchmarks. The best results are highlighted in bold , and the second-best results are underlined . Red arrows indicate absolute performance gains over the strongest baseline.
Dataset
Difficulty
SIRL
SemOPT
NL4Opt
Easy
93.6%
94.4% ( ↑ 0.8%)
Medium
82.0%
84.0% ( ↑ 2.0%)
Hard
69.2%
71.8% ( ↑ 2.6%)
ComplexLP
Easy
88.6%
88.6% (+0.0%)
Medium
94.1%
97.1% ( ↑ 3.0%)
Hard
54.5%
63.6% ( ↑ 9.1%)
Table 2: Accuracy (Best-of-N) across difficulty levels. We compare SemOPT against the single strongest baseline (SIRL) on Easy, Medium, and Hard subsets.
Backbone LLM
Method
NL4Opt
ComplexLP
ComplexOR
Gemini-2.5 Flash-Lite
Prompting (BoN)
70.1%
64.0%
55.6%
SemOPT (BoN)
79.0%
75.7%
63.0%
GPT-4.1 mini
Prompting (BoN)
74.8%
64.8%
55.6%
SemOPT (BoN)
86.9%
79.3%
66.7%
Qwen3-Max
Prompting (BoN)
76.2%
66.7%
55.6%
SemOPT (BoN)
88.8%
82.0%
66.7%
Table 3: Performance across LLM backbones. We compare SemOPT with standard prompting in BoN settings.
Variant
IndustryOR
ComplexLP
ComplexOR
ORLM
42.9%
59.5%
50.0%
SemOPT w/ ORLM generator
61.9%
74.8%
66.7%
SIRL
45.2%
80.2%
44.4%
SemOPT w/ SIRL generator
66.7%
85.6%
66.7%
Table 4: Fine-tuned generator control on complex datasets. We replace only SemOPT ’s generator with the fine-tuned models of ORLM and SIRL.
Table 6: Factorizing the semantic evaluator and the correction hierarchy on complex datasets (BoN accuracy).
Figure 4: (a) Breakdown of model outputs under different critic models. (b) Efficiency evaluation of the adaptive gating mechanism.
Method
Acc.
Time (s)
API Calls
SR (25 iters)
55.7%
61.2
26.0
BS (beam 2, branch 5)
63.3%
38.9
30.0
SemOPT
67.1 %
32.3
25.1
Table 7: Efficiency comparison with iterative and search baselines on the combined hard subsets of NL4Opt, ComplexLP, and ComplexOR.
Method
NL4Opt
ComplexLP
ComplexOR
NS
89.7% / 84.0
83.8% / 84.0
66.7% / 84.0
SemOPT
87.9% / 10.2
82.9% / 13.6
66.7% / 15.9
Table 8: Comparison with exhaustive Naive Search (NS). Each cell reports accuracy / API calls.
Figure 5: Comparison between SemOPT and the standard prompting baseline across optimization benchmarks.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Operator
Implementation Logic
Variable
Type Relax
Relax variable types (e.g., BINARY → CONTINUOUS)
Bound Remove
Remove lower/upper bounds
Objective
Sense Flip
Swap direction (e.g., MIN → MAX)
Coeff Noise
Flip sign of variable coefficient in objective
Term Drop
Remove a term from objective expression
Constraint
Comparator Flip
Swap inequality signs (e.g., ≥→≤ )
Appendix
Table 9: Semantic mutation operators for hard negative construction.
Parameter
Value
Generator (LLM) Settings
Temperature
0.9
Top-p
0.9
Max Tokens
2048
MCTS Settings
Exploration Constant ( ω )
1.414
Appendix
Table 10: Hyperparameter configurations for SemOPT .
Base Model
Train / Val Problems
Pairs
Epochs
Batch
LR
Max Len.
Qwen3-4B-Instruct-2507
2700 / 300
8254
3
32
2e−5
2048
Appendix
Table 11: Detailed training configuration for the semantic reward model.
Method
GED Acc.
Oracle Acc.
Obj Acc.
SIRL
36%
36%
38%
SemOPT
44%
47%
49%
Appendix
Table 12: Formulation-level validation on OptiVerse-LP-100.
Dataset
τ=0.2
τ=0.4
τ=0.6
τ=0.8
τ=1.0
NL4Opt
72.9% / 4.1
79.0% / 8.4
87.9% / 10.2
88.8% / 19.6
88.8% / 26.8
ComplexLP
53.2% / 4.6
73.8% / 10.2
82.9% / 13.6
83.8% / 20.1
81.1% / 27.4
ComplexOR
44.4% / 5.8
63.0% / 13.1
66.7% / 15.9
66.7% / 22.3
66.7% / 29.5
Appendix
Table 13: Sensitivity to the gating threshold τ . Each cell reports accuracy / API calls.
Figure 6: Impact of node expansion size H on SemOPT performance across three benchmarks.
Hard Subset
SIRL
SemOPT
NL4Opt
69.2% ± 2.56%
71.8% ± 4.44%
ComplexLP
54.5% ± 5.25%
63.6% ± 3.03%
ComplexOR
28.6% ± 0.00%
57.1% ± 0.00%
Appendix
Table 14: Hard-subset accuracy with variation.
Subset
SIRL
SemOPT
SemOPT -only / SIRL-only
Gain
90% CI / p
NL4Opt-hard
27/39
28/39
1/0
+2.6%
[−2.3%,2.6%] / 1.000
ComplexLP-hard
18/33
21/33
4/1
+9.1%
[−4.8%,14.8%] / 0.375
ComplexOR-hard
2/7
4/7
2/0
+28.5%
[−15.8%,28.6%] / 0.500
Combined hard
47/79
53/79
7/1
+7.6%
[0.6%,10.0%] / 0.070
Appendix
Table 15: Paired analysis on hard subsets. Counts are the number of solved instances over the subset size.
Dataset
Syntax-error Rate
Correctly Repaired
Still Erroneous (Score 0)
IndustryOR
33.1%
64.1%
35.9%
ComplexLP
34.2%
74.7%
25.3%
ComplexOR
46.8%
81.8%
18.2%
Appendix
Table 16: Syntactic error handling during search. The error rate is computed over all evaluated nodes; repair outcomes are computed over the nodes that initially triggered repair.
Error Category
Count
Share
By hierarchy layer (40 errors)
Modeling Strategy
9
22.5%
Math Model
28
70.0%
Solver Code
3
7.5%
By formulation component (28 math-model errors)
Parameter
2
7.1%
Appendix
Table 17: Distribution of the 40 remaining SemOPT errors by hierarchy layer, and of the 28 math-model errors by formulation component.
Figure 7: Case study on IndustryOR-37.
Figure 8: Full prompt for the modeling-strategy layer of MCTS.
Figure 9: Full prompt for the math model layer of MCTS.
Figure 10: Full prompt for the solver-code layer of MCTS.
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China · Noah’s Ark Lab, Huawei Technologies · Tianjin University
School of Computer Science and Engineering, Beihang University, Beijing, China · JIUTIAN Research, Beijing, China · MIIT Key Laboratory of Data and Decision Intelligence, Beihang University, Beijing, China