OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework
Authors: Jinzhi Bu, Haixin Tang, Huanan Zhang
Organizations: Department of Logistics and Maritime Studies, The Hong Kong Polytechnic University, Hong Kong · Leeds School of Business, University of Colorado Boulder, Boulder, CO, 80309, USA
Large language models can translate business descriptions into optimization models, but executable code may misrepresent constraints or objectives. A solver can then return an optimal solution to the wrong problem. Even when the solution satisfies the intended operating rules, a better plan may exist. For organizations that repeatedly use optimization modeling, an LLM-based framework should produce accurate formulations at low cost and, ideally, run locally. We study how to verify improvements and allocate attempts across LLMs that differ in price and capability. We develop OSCAR (Optimization modeling by Simulator, Coder, And Reviewer), which uses an offline Simulator certified against labeled decision examples to compare candidates and continues searching beyond feasibility. We model the search for the next certified improvement as sequential decisions under unobserved difficulty: which LLMs to call and when to stop. In a simplified known-prior setting, we give conditions under which cost-ordered escalation is optimal. For general menus, we derive a prior-free competitive guarantee. On five benchmark problems, OSCAR achieves 95% to 100% accuracy at the reported settings using two small open-weight LLMs, each deployable locally on a single GPU. Their single-attempt accuracies average 29% and 48%. In five runs per problem, Codex and Claude Code incur average token costs 3.1 and 5.8 times OSCAR's, respectively. OSCAR supports open-weight models locally or in the cloud, depending on budget and confidentiality requirements. Firms should maintain labeled decision examples of feasible and infeasible decisions to clarify plain-language operating rules. OSCAR follows these labels when an LLM's interpretation conflicts with them. As LLM capabilities and prices change, OSCAR's simple operating rules and adjustable settings help firms adapt their model choices and benefit from these advances.
Figures & tables
Figure 1Figure 2
Problem class
FrontierOR class (problems)
Problem ID
Formulation
Routing
Vehicle routing & TSP (30)
colombi2017
MIP
Graph
Graph optimization (19)
mehrotra1996
IP
Packing
Packing / cutting stock (17)
letelier2022
Binary IP
Scheduling
Scheduling (17)
elci2022
Stochastic
Quadratic
Quadratic optimization (3)
buchheim2018
QP / MIQP
Table 1: The five problems.
One-shot
OSCAR
Frontier coding agents
Problem class
Qwen3.5
Qwen3.6
n=1
n=(1,1,3,3)
Codex with GPT-6 Astra
Claude Code with Fable 5.1
Routing
7 ($0.02)
18 ($0.04)
72 ($0.17)
95 ($0.28)
5/5 ($0.65)
5/5 ($1.23)
Graph
77 ($0.01)
89 ($0.01)
100 ($0.05)
–
5/5 ($0.50)
5/5 ($0.65)
Packing
7 ($0.01)
12 ($0.02)
97 ( $0.27 )
–
5/5 ($0.64)
5/5 ($1.43)
Scheduling
4 ($0.01)
35 ($0.04)
95 ($0.24)
–
5/5 ($0.46)
5/5 ($1.13)
Quadratic
49 ($0.01)
84 ($0.01)
100 ($0.13)
–
5/5 ($0.74)
5/5 ($1.20)
Table 2: Runs reaching the reference optimum (%), with cost per run in US dollars in parentheses. For each problem, we run each one-shot model and tested OSCAR setting 100 times and each frontier agent five times.
OSCAR
Resampling, Qwen3.5
Resampling, Qwen3.6
Problem class
Correct (%)
Cost per run
k
Cost per run
k
Cost per run
Routing
95
$0.28
27
0.53(1.9\times$ )
10
0.58(2.1\times$ )
Graph
100
$0.05
3
0.03(0.7\times$ )
3
0.04(0.8\times$ )
Packing
97
$0.27
39
0.55(2.0\times$ )
24
0.66(2.5\times$ )
Scheduling
95
$0.24
50
0.80(3.3\times$ )
6
0.39(1.6\times$ )
Quadratic
100
$0.13
7
0.14(1.0\times$ )
3
0.12(0.9\times$ )
Table 3: Estimated cost of resampling to within 1% of OSCAR’s accuracy. Parentheses show cost ratios relative to OSCAR.
Component
Model
Routing n=1
Routing n=(1,1,3,3)
Graph
Packing
Scheduling
Quadratic
Simulator build
both
0.025 (9.0)
0.025 (9.0)
0.007 (3.0)
0.084 ( 16.0 )
0.075 (11.0)
0.062 (14.0)
Structure
Qwen3.5
0.002 (1.0)
0.002 (1.0)
0.001 (1.0)
0.002 (1.0)
0.002 (1.0)
0.001 (1.0)
Initial formulation
both
0.052 (4.4)
0.043 (4.0)
0.007 (1.9)
0.033 (4.2)
0.051 (3.8)
0.020 (3.6)
Reviewer
Qwen3.5
0.010 (3.5)
0.021 (7.1)
0.005 (2.1)
0.011 (3.8)
0.009 (3.4)
0.007 (2.2)
Reviewer
Qwen3.6
0.023 (2.7)
0.046 (5.1)
0.008 (2.0)
0.048 (6.0)
0.025 (3.4)
0.014 (2.1)
Coder
Qwen3.5
0.007 (3.8)
0.008 (4.7)
0.002 (2.1)
0.006 (3.9)
0.005 (3.5)
0.003 (2.2)
Table 4: Mean cost in dollars (model calls) per OSCAR run at n=1 unless indicated otherwise. Initial formulation includes its syntax repair. Loop costs include the Type 1 extension. Totals may differ from component sums because of rounding.