cs.LGSep 28, 2026

LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization

Authors: Shihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen, Jianfeng Feng, Dongdong Ge, Yinyu Ye

Organizations: East China Normal University · Alibaba Group · Fudan University · The University of Hong Kong · Tokentide AI · Shanghai Jiao Tong University · Stanford University

Abstract

Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems

    Apr 18, 2026Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov +4Operations ResearchAuto Research

  2. Learning to Optimize through Solver-Grounded Self-Play

    Sep 28, 2026Xia Jiang, Yaoxin Wu, Chenyu Zhou +3Optimization ModelingSelf-Play

  3. MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

    Jun 24, 2026Ke Zhao, Zixiang Di, Hong Qian +9Optimization ModelingLarge Reasoning Models