cs.SEFeb 17, 2026

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

Authors: Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo

Organizations: Institute of Operations Research and Analytics, National University of Singapore · McCormick School of Engineering, Northwestern University · Wenzhou Buyi Pharmacy Chain Co., Ltd. · College of Computer Science and Artificial Intelligence, Wenzhou University · Department of Decision Analytics and Operations, City University of Hong Kong

Abstract

Large language models (LLMs) can translate natural-language problem descriptions into optimization code, but the code is prone to silent failures: it executes and returns a solver-feasible solution while encoding a semantically incorrect formulation. On compositional problems, the resulting feasibility-correctness gap reaches 90 percentage points. We introduce ReLoop, which combines two mechanisms. Structured generation decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify) to reduce formulation errors during generation. Behavioral verification detects the errors that remain by testing whether the formulation responds correctly to solver-based parameter perturbation, a signal that comes from the solver rather than from LLM self-review and requires no ground truth. The two mechanisms address different error structures: structured generation gives the largest gain on compositional problems (+8.5pp accuracy on RetailOpt-190 with Claude Opus 4.6), and behavioral verification gives its largest gain on localized defects (+4.4pp on MAMO-ComplexLP). With diagnostic execution recovery, ReLoop reaches 100% executable code on Claude Opus 4.6, and relative to direct generation it raises or preserves every reported metric of the three chat-tuned foundation models on all three benchmarks. For the narrowly fine-tuned SFT model we test, the chain-of-thought prompt conflicts with its learned output format and lowers its accuracy on MAMO-ComplexLP; we document and analyze this interaction. We release RetailOpt-190, 190 compositional retail optimization scenarios in which several constraints interact.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SemOPT: Fixing Semantic Errors in LLM-based Optimization Modeling via Reward-Guided Search

    Sep 29, 2026Zetong Zhou, Wentao Zhang, Jingyuan Wang +3Optimization ModelingFunctional Code Solvers

  2. Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling

    Sep 30, 2026Zhong Li, Xin Huang, Jinhui Wan +6Optimization ModelingLarge Language Model Benchmarks