OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework
Authors: Jinzhi Bu, Haixin Tang, Huanan Zhang
Organizations: Department of Logistics and Maritime Studies, The Hong Kong Polytechnic University, Hong Kong · Leeds School of Business, University of Colorado Boulder, Boulder, CO, 80309, USA
Large language models can translate business descriptions into optimization models, but executable code may misrepresent constraints or objectives. A solver can then return an optimal solution to the wrong problem. Even when the solution satisfies the intended operating rules, a better plan may exist. For organizations that repeatedly use optimization modeling, an LLM-based framework should produce accurate formulations at low cost and, ideally, run locally. We study how to verify improvements and allocate attempts across LLMs that differ in price and capability. We develop OSCAR (Optimization modeling by Simulator, Coder, And Reviewer), which uses an offline Simulator certified against labeled decision examples to compare candidates and continues searching beyond feasibility. We model the search for the next certified improvement as sequential decisions under unobserved difficulty: which LLMs to call and when to stop. In a simplified known-prior setting, we give conditions under which cost-ordered escalation is optimal. For general menus, we derive a prior-free competitive guarantee. On five benchmark problems, OSCAR achieves 95% to 100% accuracy at the reported settings using two small open-weight LLMs, each deployable locally on a single GPU. Their single-attempt accuracies average 29% and 48%. In five runs per problem, Codex and Claude Code incur average token costs 3.1 and 5.8 times OSCAR's, respectively. OSCAR supports open-weight models locally or in the cloud, depending on budget and confidentiality requirements. Firms should maintain labeled decision examples of feasible and infeasible decisions to clarify plain-language operating rules. OSCAR follows these labels when an LLM's interpretation conflicts with them. As LLM capabilities and prices change, OSCAR's simple operating rules and adjustable settings help firms adapt their model choices and benefit from these advances.
Figures & tables
Figure 1Figure 2
Problem class
FrontierOR class (problems)
Problem ID
Formulation
Routing
Vehicle routing & TSP (30)
colombi2017
MIP
Graph
Graph optimization (19)
mehrotra1996
IP
Packing
Packing / cutting stock (17)
letelier2022
Binary IP
Scheduling
Scheduling (17)
elci2022
Stochastic
Quadratic
Quadratic optimization (3)
buchheim2018
QP / MIQP
Table 1: The five problems.
One-shot
OSCAR
Frontier coding agents
Problem class
Qwen3.5
Qwen3.6
n=1
n=(1,1,3,3)
Codex with GPT-6 Astra
Claude Code with Fable 5.1
Routing
7 ($0.02)
18 ($0.04)
72 ($0.17)
95 ($0.28)
5/5 ($0.65)
5/5 ($1.23)
Graph
77 ($0.01)
89 ($0.01)
100 ($0.05)
–
5/5 ($0.50)
5/5 ($0.65)
Packing
7 ($0.01)
12 ($0.02)
97 ( $0.27 )
–
5/5 ($0.64)
5/5 ($1.43)
Scheduling
4 ($0.01)
35 ($0.04)
95 ($0.24)
–
5/5 ($0.46)
5/5 ($1.13)
Quadratic
49 ($0.01)
84 ($0.01)
100 ($0.13)
–
5/5 ($0.74)
5/5 ($1.20)
Table 2: Runs reaching the reference optimum (%), with cost per run in US dollars in parentheses. For each problem, we run each one-shot model and tested OSCAR setting 100 times and each frontier agent five times.
OSCAR
Resampling, Qwen3.5
Resampling, Qwen3.6
Problem class
Correct (%)
Cost per run
k
Cost per run
k
Cost per run
Routing
95
$0.28
27
0.53(1.9\times$ )
10
0.58(2.1\times$ )
Graph
100
$0.05
3
0.03(0.7\times$ )
3
0.04(0.8\times$ )
Packing
97
$0.27
39
0.55(2.0\times$ )
24
0.66(2.5\times$ )
Scheduling
95
$0.24
50
0.80(3.3\times$ )
6
0.39(1.6\times$ )
Quadratic
100
$0.13
7
0.14(1.0\times$ )
3
0.12(0.9\times$ )
Table 3: Estimated cost of resampling to within 1% of OSCAR’s accuracy. Parentheses show cost ratios relative to OSCAR.
Component
Model
Routing n=1
Routing n=(1,1,3,3)
Graph
Packing
Scheduling
Quadratic
Simulator build
both
0.025 (9.0)
0.025 (9.0)
0.007 (3.0)
0.084 ( 16.0 )
0.075 (11.0)
0.062 (14.0)
Structure
Qwen3.5
0.002 (1.0)
0.002 (1.0)
0.001 (1.0)
0.002 (1.0)
0.002 (1.0)
0.001 (1.0)
Initial formulation
both
0.052 (4.4)
0.043 (4.0)
0.007 (1.9)
0.033 (4.2)
0.051 (3.8)
0.020 (3.6)
Reviewer
Qwen3.5
0.010 (3.5)
0.021 (7.1)
0.005 (2.1)
0.011 (3.8)
0.009 (3.4)
0.007 (2.2)
Reviewer
Qwen3.6
0.023 (2.7)
0.046 (5.1)
0.008 (2.0)
0.048 (6.0)
0.025 (3.4)
0.014 (2.1)
Coder
Qwen3.5
0.007 (3.8)
0.008 (4.7)
0.002 (2.1)
0.006 (3.9)
0.005 (3.5)
0.003 (2.2)
Table 4: Mean cost in dollars (model calls) per OSCAR run at n=1 unless indicated otherwise. Initial formulation includes its syntax repair. Loop costs include the Type 1 extension. Totals may differ from component sums because of rounding.
Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiting practical scalability. Therefore, we systematically investigate whether LLMs can identify problem structure from natural-language descriptions and apply suitable optimization modeling techniques to generate mathematical models and solver code that solve the problems correctly and efficiently. To this end, we first curate OptTips, a knowledge base of 50 expert modeling techniques in eight families. Using this knowledge, we develop OptDachshund, a multi-agent framework that transforms problems from existing optimization benchmarks into new tasks for evaluating LLMs' use of modeling techniques. It constructs conventional and expert mathematical models with solver code for the same task and data, providing baselines for correctness and computational cost. The resulting EfficientOpt benchmark contains 561 expert-reviewed tasks with paired reference implementations. Evaluation of 11 representative LLMs reveals an efficiency gap on correctly solved tasks with comparable measurements: for every LLM, most generated programs take longer to solve than their expert counterparts. Within the comparable reference-size subset, 57% of programs with correct objective values and fewer variables and linear constraints have longer recorded solver times. Case studies show that different modeling techniques can achieve the same optimal value at similar recorded cost. Faster solving may not reduce execution time if the code takes longer to prepare data and build the model. LLM optimization modeling should therefore be evaluated for both correctness and computational efficiency.
Zhong Li, Xin Huang, Jinhui Wan +6
Great Bay University · Beijing Jiaotong University · Beihang University +1
Operations Research (OR) provides a rigorous framework for high-stakes decision-making, but effective OR modeling requires substantial domain knowledge, mathematical abstraction, and solver expertise. Recent LLM-based systems automate parts of this pipeline, yet remain limited by low accuracy on complex problems, opaque outputs, and narrow solver support. We propose COOPA (COoperative OPerations Agent), a modular LLM-agent architecture for interpretable and scalable OR decision support. It combines three components: iterative confidence-based modeling, which generates multiple candidate formulations, self-evaluates them across modeling dimensions, and selects one using a max-min confidence criterion; element-level provenance and confidence explanations, which link variables, parameters, constraints, and objectives to quoted source text and provide an audit trail for human verification; and multi-solver routing to specialized optimizer agents for different OR problem classes. Across three OR benchmarks, eight LLM backbones, and four baselines under identical conditions, COOPA achieves the best macro-average accuracy on six of eight backbones and improves over the strongest baseline by up to 6.7 percentage points. A within-system ablation isolates the contribution of iterative confidence-based modeling, while additional analyses and case studies illustrate the value of source traceability and multi-solver dispatch.
Chuanhao Li, Xiaoan Xu, Dirk Bergemann +3
1Tsinghua University · 2Duke University · 3Yale University
LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct models and solve efficiently. We introduce Opti-Agent-Bench, an end-to-end benchmark that evaluates Large Language Models (LLMs) across the complete optimization R&D pipeline, from understanding business-language descriptions through mathematical modeling, algorithm selection, and code implementation, to solution report generation. Our design rests on three pillars: (1) businesssemantic authenticity with anti-template traps that defeat pattern matching; (2) modular evaluation with cross-module consistency checking across Problem Understanding, Formal Modeling, Implementation, and Reporting; and (3) the ORAC bi-level validity framework that simultaneously ensures task quality and scoring integrity. Across several industrialscale tasks spanning integer programming, robust optimization, stochastic programming, and non-convex optimization, we expose critical failure modes of current models, including constraint omission, model-code inconsistency, and report-implementation divergence, that remain invisible under conventional single-metric evaluation.
Yongchang Fu, Xinjie Huang, Chengjun Dai +3
Ding Talk, Alibaba Group, Hangzhou, China · Zhejiang University, Hangzhou, China