Developing optimization models for production scheduling requires substantial expert effort. Research on large language models (LLMs) has followed two directions: specialized approaches for automated modeling, mostly for mixed-integer linear programming, which often rely on dedicated training or problem-specific architectures that limit industrial deployment; and agentic artificial intelligence for operational decision support, which generally assumes that the optimization model already exists. This study bridges both directions by assessing whether general-purpose LLMs, orchestrated as agents without task-specific training, can formulate and implement constraint programming models from natural-language problem descriptions. Singleagent and multi-agent architectures are integrated with a Model Context Protocol server that provides context-aware retrieval of solver documentation to mitigate hallucinations during implementation. Both are compared with a direct LLM baseline on six industry-oriented problems covering flow-shop, job-shop, flexible job-shop and resource-constrained warehouse scheduling, using three LLMs and assessing modeling accuracy, execution success, latency and token consumption. Formulation proves largely within reach of current LLMs, whereas implementation is the main barrier. The multi-agent workflow raises the share of scripts that run correctly as generated from 14.8% with a direct LLM call to 59.3%, reaching 80.6% on the four less complex problems, while tightly coupled intralogistics models remain an open challenge.
Figures & tables
Work
Approach Summary
Method
Train
Scalability (Sol ∣ Prob ∣ LLM)
OptiMind
SLM fine-tuned and hint injection
MILP
✓
⊖∣∙∣∘
OptiMUS-0.3
Modular agent with constraint graph & reflective correction
MILP
×
∙∣∙∣∙
SMILO
Modelling graph-guided tasks & template assembly
MILP
×
⊖∣⊖∣∙
Chen et al.
Domain specialization via knowledge-augmented fine-tuning
MILP
✓
⊖∣∘∣∘
El Baz et al.
Agentic model with human-in-the-loop OR expert
MILP
×
⊖∣∘∣∙
RideAgent
Small-sample guided heuristic agent for variable fixing
MILP
×
⊖∣∘∣∙
Table 1: Recent contributions for automated mathematical modelling with LLMs.
Workflow
Component
Task
MCP tools
Max. iterations a
(a) Multi-agent
Formulator Agent
Formulation and implementation plan
–
10
Executor Agent
Coding of the given formulation
6
50
(b) Single agent
Executor Agent
Formulation and coding
6
50
(c) Direct LLM
LLM call
Formulation and coding
–
1
Table 2: Langflow components of each workflow. All components share the same LLM settings: temperature 0.15 and a limit of 15 000 output tokens per call.
Figure 1: Three workflows compared in this study. (a) Two-stage multi-agent pipeline: a tool-free Formulator Agent hands a formulation to an Executor Agent connected to the MCP documentation server. (b) Single Executor Agent connected to the same MCP server. (c) Direct LLM baseline without tool access.
Figure 2: Structure of the prompts. The task message that enters every workflow, shown for problem PC1: four sections in a fixed order, with the process described in natural language only, the identity of the SQLite instance, the schema of every table generated from the database file, and the conventions left open by the description. The complete prompts and the task messages of PC1–PC6 are provided as supplementary material and in the project repository ( Sánchez-Fernández et al., 2026 ) .
Term
Definition
Example
Module
A docplex.cp sub-package, documented in one reference file (four in the corpus)
docplex.cp.model
Symbol
A documented function, class or method (186, 15 and 132 in the corpus)
method solve of class CpoModel
Qualified name
Unique path of a symbol: module, class (if any) and name. Short names may be ambiguous
docplex.cp.model.CpoModel.solve short name solve ; add matches two classes
Chunk
Retrieval unit: the complete entry of one symbol, or one whole example script, stored under a deterministic identifier
Table 3: Nomenclature of the documentation corpus, illustrated with the method solve of class CpoModel .
Figure 3: Layered architecture of the documentation MCP server. Token counts are estimated at four characters per token. Every tool response ends with a footer stating the next required step of the retrieval protocol.
Family
Tool
Input → Output a
Key features
Discovery capped
search_docs
[concepts] → candidate names
Merges all queries and groups hits by module; ends with the ready-made get_signature call
list_functions
module → symbol index
Browsing fallback when keyword search finds nothing
list_examples
problem text → ranked examples
Ranked by similarity to the description; alphabetical without one
Precision complete
get_signature
[names] → complete entries
Order kept; unknown names form a gap list, never approximated
get_example
[ids] → full source
Complete working models, up to ≈ 5k tokens each
get_module_reference
module (+ [names]) → reference
For many symbols of one module; ≈ 28k tokens for modeler
Table 4: Tools exposed by the MCP server. Discovery responses are non-authoritative and capped at 400 characters per summary line and 40 candidates per call; precision responses are authoritative and never truncated.
Figure 4: Retrieval protocol of the Executor Agent. Top: the four steps, the tools used at each step and the in-server enforcement. Bottom: example of the tool-call trace; calls 3–4 form a repair cycle triggered by the gap list returned in call 2.
Factor
Levels
Codes a
n
Workflow
(a) multi-agent + MCP; (b) single agent + MCP; (c) direct LLM (Figure 1 )
WA, WB, WC
3
LLM b
Gemini 2.5 Pro; GPT-5; DeepSeek V4 Pro
Gem, GPT, DS
3
Problem
PC1–PC6 (Table 6 )
PC1, …, PC6
6
Replicate
Independent generations
r1, r2, r3
3
Runs ( 3×3×6×3 )
162
Table 5: Full-factorial experimental design and run nomenclature.
PC1
PC2
PC3
PC4
PC5
PC6
Problem
Type a
FSP
JSP
FJSP
RCS
RCS
RCS
Sets
Jobs / orders
11 lots
10 orders
6 batches
29 orders
29 orders
20 orders
Activities
55 operations
44 operations
18 operations
139 transports
348 transports b
360 transports b
Resources
5 machines
6 work centres
8 machines, 3 operators
1 AGV, 1 station
1 AGV, 1 station
2 AGVs, 2 robots, buffer of 5
Other
4 product families
–
38 eligible pairs
5 pallets, 5 kit types
12 pallets, 5 kit types
18 pallets, 5 kit types
Variables
Interval
55
44
18
168
29
20
Table 6: Benchmark problems: structure, size of the reference CP model and ground-truth result.
Dimension
Metric
Definition
Effectiveness
AccF+C (%)
Percentage of replications of the same experimental setting for which the generated script executes successfully without modification and returns an objective value satisfying z=zref
AccF (%)
Percentage of replications of the same experimental setting that reach zref after implementation errors are corrected, without modifying the underlying mathematical formulation
Efficiency
Tokens
Input and output tokens of all LLM calls of the run
Time (s)
Wall-clock time of the complete workflow
Tool calls
Number of calls to the MCP documentation server (workflows (a) and (b))
Diagnostics
Model outcome
Result of the generated model, as generated and after correction: reaches zref ; feasible but worse than zref ; invalid (objective below a proven bound); no solution within 60 s or proven infeasible; no objective reported; execution error; no complete script
Table 7: Evaluation metrics reported for each run.
Accuracy (%)
Tokens (k)
Time
Tool
Workflow
Problem
LLM
F
F+C
Input
Output
(s)
calls
WA
PC1 (FSP)
Gem
100.0
100.0
152.1 ± 3.1
21.8 ± 1.5
202 ± 16
6.0 ± 0.0
GPT
100.0
66.7
414.2 ± 173.7
24.4 ± 6.2
333 ± 84
10.3 ± 3.1
DS
100.0
66.7
1108.6 ± 179.7
19.6 ± 3.2
516 ± 135
21.7 ± 2.5
PC2 (JSP)
Gem
100.0
100.0
140.3 ± 25.0
20.1 ± 5.0
178 ± 42
6.3 ± 1.2
GPT
100.0
100.0
606.7 ± 121.8
26.6 ± 4.0
372 ± 72
13.7 ± 2.3
Table 8: Summary of results per experimental setup (workflow, problem and LLM): accuracy at both levels, token consumption, inference time and MCP tool calls.
Figure 5: Accuracy per problem and workflow across the three LLMs: (a) formulation + coding accuracy, AccF+C ; (b) formulation accuracy, AccF . Mean ± standard deviation over the nine runs of each workflow and problem (three LLMs × three replicates).
Figure 6: Formulation + coding accuracy, AccF+C , per problem, LLM and workflow: mean ± standard deviation over three replicates.
Figure 7: Formulation accuracy, AccF , per problem, LLM and workflow: mean ± standard deviation over three replicates.
Figure 8: Inference time (bars, left axis) and tokens per run (diamonds, right axis, logarithmic scale) per LLM and workflow: mean ± standard deviation over the 18 runs of each workflow–LLM pair.
Workflow
Problem
Outcome of the generated model
(a)
(b)
(c)
PC1–PC4
PC5
PC6
All
Runs
54
54
54
108
27
27
162
Script as generated (single execution, no edits)
Reaches zref (correct model)
32
18
8
53
5
–
58
Invalid: objective below a proven bound
–
–
1
–
1
–
1
No solution within 60 s or proven infeasible
7
5
2
–
4
10
14
Table 9: Outcome of the 162 generated models, as generated and after correcting their implementation errors, aggregated per workflow and per problem.
Figure 9: Error analysis of the generated scripts. (a) Outcome of the 54 scripts of each workflow: correct as generated, or first point at which the script fails; result reporting includes the 14 scripts that fail after printing zref , and the last class groups the scripts that run to completion without a valid solution. (b) The twelve most frequent implementation errors among the 159 distinct errors of the 104 incorrect runs, grouped by nature: hallucinated symbol (a name that does not exist in docplex.cp ), misuse of an existing element of the API, silent semantic error, and format, environment or generic Python defects.
Antai College of Economics and Management, Shanghai Jiao Tong University · University of Chicago Booth School of Business · School of Economics and Management, Tsinghua University +1