cs.CLSep 29, 2026

VLM Fine-Tuning for End-to-End Combinatorial Optimization

Authors: Qingsong Yan, Xia Jiang, Yaoxin Wu, Wen Song, Lu Zhang, Yingjie Zhou

Organizations: College of Computer Science, Sichuan University · Department of Industrial Engineering&Innovation Sciences, Eindhoven University of Technology · Institute of Marine Science and Technology, Shandong University · School of Cybersecurity, Chengdu University of Information Technology

Abstract

Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.

Figures & tables

Explore similar work

CardsList
  1. Formalize, Don't Optimize: The Heuristic Trap in LLM-Generated Combinatorial Solvers

    May 12, 2026Haoyu Wang, Yuliang Song, Tao Li +5Functional Code SolversSolvers

  2. Beyond Objective Equivalence: Constraint Injection for LLM-Based Optimization Modeling on Vehicle Routing Problems

    Jun 3, 2026Xizi Luo, Changhong He, Dongdong Geng +2Optimization ModelingVehicle Routing Problem

  3. Beyond Inference-Time Search: Reinforcement Learning Synthesizes Reusable Solvers

    May 18, 2026Soheyl Massoudi, Gabriel Apaza, Milad Habibi +1Functional Code SolversOffline Reinforcement Learning