Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling
Organizations: Great Bay University · Beijing Jiaotong University · Beihang University · Peking University
Abstract
Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiting practical scalability. Therefore, we systematically investigate whether LLMs can identify problem structure from natural-language descriptions and apply suitable optimization modeling techniques to generate mathematical models and solver code that solve the problems correctly and efficiently. To this end, we first curate OptTips, a knowledge base of 50 expert modeling techniques in eight families. Using this knowledge, we develop OptDachshund, a multi-agent framework that transforms problems from existing optimization benchmarks into new tasks for evaluating LLMs' use of modeling techniques. It constructs conventional and expert mathematical models with solver code for the same task and data, providing baselines for correctness and computational cost. The resulting EfficientOpt benchmark contains 561 expert-reviewed tasks with paired reference implementations. Evaluation of 11 representative LLMs reveals an efficiency gap on correctly solved tasks with comparable measurements: for every LLM, most generated programs take longer to solve than their expert counterparts. Within the comparable reference-size subset, 57% of programs with correct objective values and fewer variables and linear constraints have longer recorded solver times. Case studies show that different modeling techniques can achieve the same optimal value at similar recorded cost. Faster solving may not reduce execution time if the code takes longer to prepare data and build the model. LLM optimization modeling should therefore be evaluated for both correctness and computational efficiency.
Figures & tables
| Correct answers | Cost ratio (LLM/reference) | |||||
|---|---|---|---|---|---|---|
| Runtime | Work | Build+opt | ||||
| Model | Count | Rate (%) | Ordinary | Expert | Expert | Expert |
| Gemini 3.1 Pro | 512 | 91.3 | 0.52 | 1.68 | 1.98 | 1.43 |
| GPT-5.5 | 447 | 79.7 | 0.58 | 1.79 | 2.08 | 1.64 |
| Claude Opus 4.6 | 442 | 78.8 | 0.46 | 1.49 | 1.77 | 1.42 |
| DeepSeek-V4 Flash | 441 | 78.6 | 0.49 | 1.62 | 1.78 | 1.47 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Primary problem category | Application field | ||
|---|---|---|---|
| Resource allocation & production | 177 | Mobility & logistics | 114 |
| Routing & network flow | 121 | Manufacturing & industry | 106 |
| Packing & selection | 88 | Unspecified | 76 |
| Covering & facility planning | 68 | Public services & research | 61 |
| Assignment & matching | 45 | Business & resource planning | 47 |
| Scheduling & control | 37 | Health & life sciences | 46 |
| Mutually exclusive outcomes | Solver status | |||
| Model | Correct | Objective mismatch | Other outcome | OPTIMAL |
| Gemini 3.1 Pro | 512 | 10 | 39 | 522 |
| GPT-5.5 | 447 | 23 | 91 | 470 |
| Claude Opus 4.6 | 442 | 36 | 83 | 478 |
| DeepSeek-V4 Flash | 441 | 25 | 95 | 466 |
| Kimi K2.6 | 433 | 23 | 105 | 456 |
| Runtime (s) | Work | Build | Build+opt | Shared Runtime | |||
|---|---|---|---|---|---|---|---|
| Model | Median | Mean | P90 | Median | Median | Median | Mean ( ) |
| Gemini 3.1 Pro | 32.51 | 197.66 | 414.53 | 22.66 | 0.63 | 37.80 | 77.04 |
| GPT-5.5 | 28.84 | 208.50 | 671.57 | 21.57 | 0.97 | 37.77 | 197.44 |
| Claude Opus 4.6 | 23.52 | 177.59 | 369.99 | 18.84 | 0.87 | 37.38 | 130.45 |
| DeepSeek-V4 Flash | 23.11 | 190.92 | 430.92 | 18.19 | 0.82 | 32.71 | 109.45 |
| Kimi K2.6 | 29.70 | 210.28 | 599.96 | 20.36 | 0.76 | 36.10 | 166.83 |
| Models with correct answers | Tasks | Runtime pairs | /Expert | Slower (%) |
|---|---|---|---|---|
| 1–3 | 26 | 67 | 1.69 | 58.2 |
| 4–7 | 192 | 1,169 | 1.58 | 57.5 |
| 8–10 | 276 | 2,449 | 1.67 | 66.8 |
| 11 (all models) | 47 | 517 | 2.82 | 82.2 |
| No larger on three counts | Fewer variables and rows | ||||||
| Model | Eligible | Slower | % | Slower | % | ||
| Gemini 3.1 Pro | 465 | 100 | 55 | 55.0 | 45 | 24 | 53.3 |
| GPT-5.5 | 401 | 45 | 26 | 57.8 | 19 | 11 | 57.9 |
| Claude Opus 4.6 | 395 | 82 | 44 | 53.7 | 41 | 19 | 46.3 |
| DeepSeek-V4 Flash | 395 | 59 | 31 | 52.5 | 29 | 10 | 34.5 |
| Kimi K2.6 | 390 | 70 | 40 | 57.1 | 28 | 13 | 46.4 |
| Model | LLM runs | Build longer (%) | LLM–expert pairs | Faster to slower | Slower to faster |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 504 | 58 (11.5) | 499 | 0 | 49 |
| GPT-5.5 | 441 | 73 (16.6) | 435 | 8 | 33 |
| Claude Opus 4.6 | 435 | 76 (17.5) | 430 | 21 | 29 |
| DeepSeek-V4 Flash | 433 | 88 (20.3) | 428 | 14 | 18 |
| Kimi K2.6 | 427 | 71 (16.6) | 423 | 14 | 22 |
| Qwen 3.6 27B | 391 | 88 (22.5) | 386 | 24 | 15 |
| Model | API latency (s) | Prompt | Completion | Reasoning | ||
|---|---|---|---|---|---|---|
| Gemini 3.1 Pro | 559/558 | 55.4 | 743.5 | 6,714.5 | 5,827 | 558 |
| GPT-5.5 | 535/534 | 61.2 | 715.5 | 1,431 | 516 | 522 |
| Claude Opus 4.6 | 561/561 | 17.2 | 937 | 748 | 0 | 561 |
| DeepSeek-V4 Flash | 561/561 | 81.9 | 714 | 11,327 | 10,562 | 561 |
| Kimi K2.6 | 561/561 | 82.2 | 714 | 11,680 | 10,714 | 561 |
| Qwen 3.6 27B | 558/558 | 194.9 | 724 | 8,319 | 8,193 | 190 |
| All correct cost-analysis runs | Matched subset with process RSS | |||||||
| Gurobi peak | RSS snapshot max | Gurobi peak | ||||||
| Model | Median | P90 | Median | P90 | Median | P90 | ||
| Gemini 3.1 Pro | 504 | 124.7 | 2,397.9 | 210 | 79.2 | 2,011.6 | 72.0 | 2,403.3 |
| GPT-5.5 | 441 | 136.3 | 2,198.6 | 186 | 70.8 | 2,069.4 | 58.7 | 2,440.1 |
| Claude Opus 4.6 | 435 | 118.4 | 2,510.0 | 176 | 78.8 | 2,082.5 | 57.8 | 2,507.5 |
| DeepSeek-V4 Flash | 433 | 130.4 | 1,938.5 | 184 | 75.9 | 1,234.1 | 69.6 | 1,660.2 |
| Model | Correct (%) | OPTIMAL | Static failure | Code error | Time limit | Memory limit | Other |
|---|---|---|---|---|---|---|---|
| OptMATH (32B) | 137 (24.42) | 189 | 11 | 212 | 51 | 57 | 41 |
| SIRL (32B) | 118 (21.03) | 188 | 34 | 164 | 77 | 72 | 26 |
| LLMOPT (14B) | 36 (6.42) | 77 | 20 | 377 | 28 | 35 | 24 |
| OptMATH (7B) | 1 (0.18) | 4 | 401 | 146 | 3 | 6 | 1 |
| SIRL (7B) | 0 (0.00) | 0 | 504 | 57 | 0 | 0 | 0 |
| All eligible runs | Slowest 10% | ||||||
|---|---|---|---|---|---|---|---|
| Model | Median (s) | Total (h) | Time share (%) | Slower | LLM/expert | ||
| Gemini 3.1 Pro | 504 | 32.5 | 27.67 | 51 | 72.4 | 48/51 | 5.43 |
| GPT-5.5 | 441 | 28.8 | 25.54 | 45 | 70.2 | 43/45 | 7.71 |
| Claude Opus 4.6 | 435 | 23.5 | 21.46 | 44 | 72.7 | 42/44 | 5.31 |
| DeepSeek-V4 Flash | 433 | 23.1 | 22.96 | 44 | 72.9 | 43/44 | 5.80 |
| Kimi K2.6 | 427 | 29.7 | 24.94 | 43 | 68.4 | 42/43 | 4.97 |
| Model | Task | LLM (s) | Expert (s) | Share (%) |
|---|---|---|---|---|
| Gemini 3.1 Pro | T24_002 | 4,677.80 | 303.66 | 4.7 |
| GPT-5.5 | T25_019 | 4,071.28 | 16.38 | 4.4 |
| Claude Opus 4.6 | T48_009 | 4,718.84 | 22.63 | 6.1 |
| DeepSeek-V4 Flash | T15_012 | 5,431.07 | 18.95 | 6.6 |
| Kimi K2.6 | T06_001 | 4,324.49 | 2,905.46 | 4.8 |
| Qwen 3.6 27B | T07_003 | 7,065.35 | 15.21 | 8.4 |
| Quantity | Recorded measurement | What it captures |
|---|---|---|
| Numerical check | OPTIMAL status and a finite objective matching the verified reference value | Numerical agreement on the tested instance; full task semantics and required decisions need separate verification. |
| Technique use | Automated labels from code and task context | Identifies target, alternative, none, or N/A; does not establish correctness or efficiency (Appendix F.1 ). |
| Solver cost | Gurobi Runtime (s), Work (units) | Time and solver effort for the recorded optimization calls. |
| Build+opt | Preparation time + separately timed solver calls (s) | An execution subtotal; excludes work outside the recorded intervals. |
| Generation | API latency (s); provider-reported Tokens | Waiting time for the final recorded call; token usage reported for its response. Earlier calls are excluded. |
| Model size | Variables, constraints, linear-matrix nonzeros | Dimensions before presolve. Figure 5 and Appendix C.4 use different counting scopes, explained below. |
| LLM technique label | Combined diagnostic category | ||||||||
| Model | Target | Alt. | None | N/A | C1 | C2 | C3 | C4 | C5 |
| Claude Opus 4.6 | 311 | 131 | 98 | 21 | 253 | 52 | 66 | 68 | 122 |
| Gemini 3.1 Pro | 297 | 173 | 79 | 12 | 272 | 67 | 99 | 73 | 50 |
| DeepSeek-V4 Flash | 296 | 139 | 100 | 26 | 247 | 62 | 62 | 71 | 119 |
| Qwen 3.6 27B | 286 | 132 | 128 | 15 | 196 | 44 | 75 | 79 | 167 |
| GPT-5.5 | 284 | 147 | 106 | 24 | 223 | 68 | 71 | 82 | 117 |
| ID | Target OptTips group | Target (%) | Alt. | None | N/A | |
|---|---|---|---|---|---|---|
| T01 | Equality/inequality transformations | 352 | 283 (80.4) | 10 | 39 | 20 |
| T02 | Auxiliary variables and epigraphs | 165 | 156 (94.5) | 0 | 1 | 8 |
| T03 | Bound tightening and scaling | 209 | 12 (5.7) | 115 | 79 | 3 |
| T04 | LP relaxation | 275 | 23 (8.4) | 94 | 147 | 11 |
| T05 | Big- modeling | 209 | 171 (81.8) | 7 | 10 | 21 |
| T06 | Indicators, semicontinuous variables, SOS | 231 | 58 (25.1) | 149 | 15 | 9 |
| Task and output | Implemented choice | What the evidence supports |
|---|---|---|
| T10_011, GPT-5.5 | Binary pairwise products with separate product inequalities; the expert uses continuous interactions with row–column marginals. | Both formulations enforce the same products at integer assignments, but their LP relaxations differ. The GPT model does not use the expert’s marginal formulation. |
| T13_006, Gemini | Primal resource-allocation LP with explicit implied upper bounds; the expert uses a dual LP. | The approaches attain the same objective with similar recorded Build+opt times. |
| T20_011, Kimi / Qwen 3.5 122B | Kimi uses cumulative auxiliary variables; Qwen uses direct coverage constraints. | Eliminating Kimi’s auxiliaries preserves feasible integer shift-start decisions and the objective. Qwen’s longer construction time largely offsets its solver-time advantage. |
| Code | Methodological family | Cards |
|---|---|---|
| A | Formulation standardization and dimension reduction | 6 |
| B | Relaxation strengthening and convexification | 5 |
| C | Duality, optimality, and equilibrium | 4 |
| D | Specialized combinatorial and structured models | 8 |
| E | Discrete and logical modeling | 4 |
| F | Convex, nonlinear, and structure-preserving modeling | 9 |
| ID | Technique | Family | ID | Technique | Family |
|---|---|---|---|---|---|
| T01 | Equality/inequality equivalent transformation | A | T26 | Complementarity and KKT/MPEC modeling | E |
| T02 | Auxiliary-variable epigraph/hypograph modeling | A | T27 | Multi-objective and lexicographic optimization | H |
| T03 | Bound tightening and scaling | A | T28 | Penalty, slack, and feasibility-relaxation modeling | H |
| T04 | LP relaxation of IP/MIP | E | T29 | Solver-native modeling constructs | H |
| T05 | Big- modeling | E | T30 | Warm starts and incumbent injection | H |
| T06 | Indicator, semicontinuous, and SOS constraints | E | T31 | Algorithm–structure matching | H |