VLM Fine-Tuning for End-to-End Combinatorial Optimization
Organizations: College of Computer Science, Sichuan University · Department of Industrial Engineering&Innovation Sciences, Eindhoven University of Technology · Institute of Marine Science and Technology, Shandong University · School of Cybersecurity, Chengdu University of Information Technology
Abstract
Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.
Figures & tables
| Method | TSP | OP | CVRP | MIS | MVC | PFSP | JSSP | Time | |||||||
| Feas. | Gap | Feas. | Gap | Feas. | Gap | Feas. | Gap | Feas. | Gap | Feas. | Gap | Feas. | Gap | ||
| General-purpose Language Models | |||||||||||||||
| GPT-4o | 39% | 33.79% ±16.6 | 59% | 55.19% ±15.7 | 15% | 76.62% ±7.9 | 8% | 11.70% ±11.8 | 6% | 16.67% ±7.0 | 88% | 20.57% ±9.2 | 7% | 97.85% ±23.7 | 5.3s |
| Claude-Sonnet | 66% | 24.53% ±10.7 | 49% | 34.62% ±14.1 | 30% | 38.34% ±15.9 | 13% | 12.51% ±12.5 | 2% | 6.25% ±6.3 | 100% | 18.42% ±8.9 | 10% | 90.00% ±21.6 | 5.4s |
| DeepSeek-V3 | 73% | 35.75% ±15.4 | 50% | 46.10% ±13.4 | 21% | 58.22% ±26.8 | 5% | 12.05% ±12.9 | 15% | 37.15% ±24.8 | 58% | 20.81% ±9.4 | 52% | 103.19% ±26.9 | 26.4s |
| Llama3.3-70B | 50% | 69.08% ±31.4 | 27% | 48.98% ±14.6 | 31% | 97.31% ±69.3 | 8% | 37.12% ±29.5 | 20% | 22.86% ±13.6 | 98% | 21.97% ±8.4 | 29% | 105.01% ±24.5 | 2.1s |
| TSP | Method | Small instances | Medium instances | Large instances | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gap | Gap@1 | Gap@5 | Gap@10 | Gap | Gap@1 | Gap@5 | Gap@10 | Gap | Gap@1 | Gap@5 | Gap@10 | ||
| OR-Tools | 0.82% | 76% | 96% | 99% | 2.59% | 28% | 86% | 99% | 3.59% | 12% | 80% | 99% | |
| ACO | 1.98% | 48% | 88% | 100% | 17.98% | 0% | 1% | 6% | 36.69% | 0% | 0% | 0% | |
| LLM | 0.14 % | 96% | 100% | 100% | 0.70% | 74% | 100% | 100% | 1.34% | 44% | 100% | 100% | |
| Ours | 0.20% | 95% | 100% | 100% | 0.65% | 81% | 100% | 100% | 1.21% | 55% | 100% | 100% | |
| OP | Tsili | 3.85% | 21% | 68% | 96% | 9.54% | 0% | 2% | 55% | 13.80% | 0% | 0% | 8% |