cs.LGJul 31, 2026

Stabilized Best-of-K Training for Neural Combinatorial Optimization

Authors: Melveena JollyMidhun Xavier

Organizations: 1Independent Researcher

Abstract

Leader Reward modifies POMO training to emphasize the best trajectory produced by repeated inference. We test a narrow extension: replace its binary leader/non-leader distinction with a stabilized rank signal indexed by a sampling budget KK. With the POMO architecture, 3,050-epoch schedule, and TSP-100 test set held fixed, the Leader Reward reimplementation obtains 7.76627.7662 under 100-start, 8-augmentation greedy decoding, matching the reported 7.7667.766 at its displayed precision. Under independent sampling, the stabilized K=8K=8 recipe lowers realized Best-of-8 cost in all three paired training seeds: 7.79447.7944 versus 7.81367.8136. This observation is estimation-only and decoder-specific: three seeds are below the six-seed testing floor, Leader Reward is better at sampled K=1K=1, and it remains slightly better under its original augmented-greedy protocol. We make no unbiased-estimator, universal superiority, or state-of-the-art claim.

Explore similar work

CardsList