TACIT: Optimization Models that Learn from Their Mistakes
Organizations: Massachusetts Institute of Technology, Cambridge, MA, USA · Microsoft Research, Redmond, WA, USA
Abstract
Real-world optimization problems are difficult to model accurately because many objectives and constraints reside in domain experts' tacit knowledge, making them hard to formalize. As a result, optimization models often contain miscalibrated objectives, missing constraints, or omitted decision variables, leading to solutions that fail to reflect operational realities. We address this challenge by automatically repairing misspecified formulations using historical data consisting of past solutions and subsequent user overrides. Traditional approaches such as inverse optimization and constraint learning tend to overfit sparse data and produce complex formulations. Our central idea is to combine the reasoning capabilities and prior knowledge of LLMs with the formal grounding provided by optimization. We realize this idea through two complementary paradigms. Top-down, an LLM proposes structural repairs, including new constraints and variables, whose numerical parameters are calibrated and validated through optimization. Bottom-up, optimization infers cuts from observed decisions, which the LLM contextualizes into interpretable, generalizable modeling constraints. We evaluate our approach on 38 misspecification scenarios spanning nine classes of optimization problems, several drawn from real-world applications, and show that TACIT can repair 78.9% of them (vs. 60.5% for the best baseline).
Figures & tables
| Training | Test | ||||||||
| Model | sep. | feas. | opt. | feas. | opt. | recovery | recoverable | ||
| AOP | AOP linear ( ) | 0.94 | 0.0% | 0.57 | 0.07 | 0.57 | 0.04 | 0.0% | 0.0% |
| AOP tree ( ) | 0.93 | 0.0% | 0.58 | 0.10 | 0.57 | 0.07 | 0.0% | 0.0% | |
| LLM-only | DeepSeek-V4-F | 0.64 | 18.9% | 0.57 | 0.37 | 0.57 | 0.37 | 19.5% | 31.6% |
| Kimi-K2.6 | 0.56 | 29.5% | 0.66 | 0.46 | 0.66 | 0.47 | 30.0% | 36.8% | |
| GPT-5.6-sol | 0.40 | 47.9% | 0.74 | 0.59 | 0.75 | 0.59 | 46.8% | 60.5% | |
| Training | Test | |||||||
| sep. | feas. | opt. | feas. | opt. | recovery | recoverable | ||
| Obj. calibration only | 0.76 | 5.8% | 0.41 | 0.23 | 0.41 | 0.23 | 5.3% | 5.3% |
| Cuts | 0.35 | 33.2% | 0.57 | 0.33 | 0.54 | 0.27 | 14.7% | 21.1% |
| Cuts + contextualize | 0.35 | 33.2% | 0.59 | 0.35 | 0.57 | 0.30 | 17.9% | 26.3% |
| Template | 0.60 | 22.6% | 0.56 | 0.39 | 0.56 | 0.39 | 21.1% | 31.6% |
| Template + complete | 0.31 | 47.9% | 0.77 | 0.59 | 0.77 | 0.58 | 41.1% | 57.9% |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Objective | Constraints | Variables | Constants | ||||||
| Miscal. | Missing | Local | Global | Extra | Missing | Domain | Missing | All | |
| Instances | 7 | 6 | 10 | 15 | 5 | 10 | 2 | 16 | 38 |
| Initial model | 0.24 | 0.22 | 0.20 | 0.11 | 0.17 | 0.18 | 0.01 | 0.13 | 0.17 |
| Calib. only | 0.49 | 0.26 | 0.20 | 0.11 | 0.28 | 0.21 | 0.01 | 0.13 | 0.23 |
| Cuts | 0.52 | 0.20 | 0.41 | 0.07 | 0.40 | 0.12 | 0.00 | 0.06 | 0.27 |
| Cuts + contextualize | 0.52 | 0.22 | 0.43 | 0.11 | 0.41 | 0.13 | 0.00 | 0.08 | 0.30 |
| Training instances | 32 | 64 | 128 | 256 |
| Calibration | 4.7% | 5.8% | 5.3% | 6.3% |
| Cuts + contextualize | 7.9% | 13.7% | 17.9% | 20.5% |
| Template + complete | 36.8% | 43.2% | 41.1% | 40.5% |
| All | 36.3% | 45.8% | 45.3% | 43.2% |
| Training | Test | ||||||
| Dataset | feas. | opt. | feas. | opt. | rec. % | recoverable | |
| Optimal targets | 0.19 | 0.79 | 0.64 | 0.78 | 0.61 | 45.3% | 63.2% |
| Suboptimal targets | 0.19 | 0.81 | 0.58 | 0.79 | 0.53 | 40.5% | 52.6% |
| Training | Test | |||||||
| Model | sep. | feas. | opt. | feas. | opt. | rec. % | recoverable | |
| Initial | 0.85 | 0.0% | 0.46 | 0.17 | 0.46 | 0.17 | - | - |
| AOP linear ( ) | 0.90 | 0.0% | 0.50 | 0.15 | 0.51 | 0.13 | 0.0% | 0.0% |
| AOP linear ( ) | 0.94 | 0.0% | 0.57 | 0.07 | 0.57 | 0.04 | 0.0% | 0.0% |
| AOP linear ( ) | 0.98 | 0.0% | 0.61 | 0.02 | 0.62 | 0.01 | 0.0% | 0.0% |
| AOP tree ( ) | 0.91 | 0.0% | 0.52 | 0.17 | 0.52 | 0.15 | 0.0% | 0.0% |
| GPU Scenario 3: demand 0 can only be satisfied whenever all other demands are satisfied. | |
| Missing constraint: | |
| Raw cut from ( P cut ): | |
| LLM refined cut: | |
| WA Scenario 3: employees 3 and 5 can never work on the same day. | |
| Missing constraint: | |
| Raw cut from ( P cut ): | |
| Template (raw) | Template (fitted) | |
| CVRP Scenario 2: the ground truth minimizes the longest route instead of the total distance. | ||
| Variables: | ||
| Constants: | – | – |
| Constraints: | ||
| TP Scenario 3: goods are shipped in batches of a fixed size. | ||
| Variables: | ||
| Objective | Constraints | Variables | Constants | ||||||
| Scenario | Miscalibrated | Missing | Local | Global | Extra | Missing | Domain | Missing | |
| CFL | S1 | ||||||||
| S2 | |||||||||
| S3 | |||||||||
| S4 | |||||||||
| S5 | |||||||||