Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming
Organizations: Universidade da Coru˜na, Campus Industrial de Ferrol, CITENI, Grupo Integrado de Ingenier´ıa, Esteiro s/n, Ferrol, 15403.
Abstract
Developing optimization models for production scheduling requires substantial expert effort. Research on large language models (LLMs) has followed two directions: specialized approaches for automated modeling, mostly for mixed-integer linear programming, which often rely on dedicated training or problem-specific architectures that limit industrial deployment; and agentic artificial intelligence for operational decision support, which generally assumes that the optimization model already exists. This study bridges both directions by assessing whether general-purpose LLMs, orchestrated as agents without task-specific training, can formulate and implement constraint programming models from natural-language problem descriptions. Singleagent and multi-agent architectures are integrated with a Model Context Protocol server that provides context-aware retrieval of solver documentation to mitigate hallucinations during implementation. Both are compared with a direct LLM baseline on six industry-oriented problems covering flow-shop, job-shop, flexible job-shop and resource-constrained warehouse scheduling, using three LLMs and assessing modeling accuracy, execution success, latency and token consumption. Formulation proves largely within reach of current LLMs, whereas implementation is the main barrier. The multi-agent workflow raises the share of scripts that run correctly as generated from 14.8% with a direct LLM call to 59.3%, reaching 80.6% on the four less complex problems, while tightly coupled intralogistics models remain an open challenge.
Figures & tables
| Work | Approach Summary | Method | Train | Scalability (Sol Prob LLM) |
| OptiMind | SLM fine-tuned and hint injection | MILP | ||
| OptiMUS-0.3 | Modular agent with constraint graph & reflective correction | MILP | ||
| SMILO | Modelling graph-guided tasks & template assembly | MILP | ||
| Chen et al. | Domain specialization via knowledge-augmented fine-tuning | MILP | ||
| El Baz et al. | Agentic model with human-in-the-loop OR expert | MILP | ||
| RideAgent | Small-sample guided heuristic agent for variable fixing | MILP |
| Workflow | Component | Task | MCP tools | Max. iterations a |
|---|---|---|---|---|
| (a) Multi-agent | Formulator Agent | Formulation and implementation plan | – | 10 |
| Executor Agent | Coding of the given formulation | 6 | 50 | |
| (b) Single agent | Executor Agent | Formulation and coding | 6 | 50 |
| (c) Direct LLM | LLM call | Formulation and coding | – | 1 |
| Term | Definition | Example |
|---|---|---|
| Module | A docplex.cp sub-package, documented in one reference file (four in the corpus) | docplex.cp.model |
| Symbol | A documented function, class or method (186, 15 and 132 in the corpus) | method solve of class CpoModel |
| Qualified name | Unique path of a symbol: module, class (if any) and name. Short names may be ambiguous | docplex.cp.model.CpoModel.solve short name solve ; add matches two classes |
| Chunk | Retrieval unit: the complete entry of one symbol, or one whole example script, stored under a deterministic identifier | api:docplex.cp.model: docplex.cp.model.CpoModel.solve ( 535 tokens) example:flow_shop_permutation |
| Family | Tool | Input Output a | Key features |
|---|---|---|---|
| Discovery capped | search_docs | [concepts] candidate names | Merges all queries and groups hits by module; ends with the ready-made get_signature call |
| list_functions | module symbol index | Browsing fallback when keyword search finds nothing | |
| list_examples | problem text ranked examples | Ranked by similarity to the description; alphabetical without one | |
| Precision complete | get_signature | [names] complete entries | Order kept; unknown names form a gap list, never approximated |
| get_example | [ids] full source | Complete working models, up to 5k tokens each | |
| get_module_reference | module (+ [names]) reference | For many symbols of one module; 28k tokens for modeler |
| Factor | Levels | Codes a | |
|---|---|---|---|
| Workflow | (a) multi-agent + MCP; (b) single agent + MCP; (c) direct LLM (Figure 1 ) | WA, WB, WC | 3 |
| LLM b | Gemini 2.5 Pro; GPT-5; DeepSeek V4 Pro | Gem, GPT, DS | 3 |
| Problem | PC1–PC6 (Table 6 ) | PC1, …, PC6 | 6 |
| Replicate | Independent generations | r1, r2, r3 | 3 |
| Runs ( ) | 162 | ||
| PC1 | PC2 | PC3 | PC4 | PC5 | PC6 | ||
|---|---|---|---|---|---|---|---|
| Problem | Type a | FSP | JSP | FJSP | RCS | RCS | RCS |
| Sets | Jobs / orders | 11 lots | 10 orders | 6 batches | 29 orders | 29 orders | 20 orders |
| Activities | 55 operations | 44 operations | 18 operations | 139 transports | 348 transports b | 360 transports b | |
| Resources | 5 machines | 6 work centres | 8 machines, 3 operators | 1 AGV, 1 station | 1 AGV, 1 station | 2 AGVs, 2 robots, buffer of 5 | |
| Other | 4 product families | – | 38 eligible pairs | 5 pallets, 5 kit types | 12 pallets, 5 kit types | 18 pallets, 5 kit types | |
| Variables | Interval | 55 | 44 | 18 | 168 | 29 | 20 |
| Dimension | Metric | Definition |
|---|---|---|
| Effectiveness | (%) | Percentage of replications of the same experimental setting for which the generated script executes successfully without modification and returns an objective value satisfying |
| (%) | Percentage of replications of the same experimental setting that reach after implementation errors are corrected, without modifying the underlying mathematical formulation | |
| Efficiency | Tokens | Input and output tokens of all LLM calls of the run |
| Time (s) | Wall-clock time of the complete workflow | |
| Tool calls | Number of calls to the MCP documentation server (workflows (a) and (b)) | |
| Diagnostics | Model outcome | Result of the generated model, as generated and after correction: reaches ; feasible but worse than ; invalid (objective below a proven bound); no solution within 60 s or proven infeasible; no objective reported; execution error; no complete script |
| Accuracy (%) | Tokens (k) | Time | Tool | |||||
| Workflow | Problem | LLM | F | F+C | Input | Output | (s) | calls |
| WA | PC1 (FSP) | Gem | 100.0 | 100.0 | 152.1 3.1 | 21.8 1.5 | 202 16 | 6.0 0.0 |
| GPT | 100.0 | 66.7 | 414.2 173.7 | 24.4 6.2 | 333 84 | 10.3 3.1 | ||
| DS | 100.0 | 66.7 | 1108.6 179.7 | 19.6 3.2 | 516 135 | 21.7 2.5 | ||
| PC2 (JSP) | Gem | 100.0 | 100.0 | 140.3 25.0 | 20.1 5.0 | 178 42 | 6.3 1.2 | |
| GPT | 100.0 | 100.0 | 606.7 121.8 | 26.6 4.0 | 372 72 | 13.7 2.3 | ||
| Workflow | Problem | ||||||
| Outcome of the generated model | (a) | (b) | (c) | PC1–PC4 | PC5 | PC6 | All |
| Runs | 54 | 54 | 54 | 108 | 27 | 27 | 162 |
| Script as generated (single execution, no edits) | |||||||
| Reaches (correct model) | 32 | 18 | 8 | 53 | 5 | – | 58 |
| Invalid: objective below a proven bound | – | – | 1 | – | 1 | – | 1 |
| No solution within 60 s or proven infeasible | 7 | 5 | 2 | – | 4 | 10 | 14 |