cs.LGOct 5, 2026

Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits

Authors: Vikram Kakaria, Anish Kataria, Anany Kotawala

Organizations: Princeton University

Abstract

Combinatorial Thompson sampling (CTS) draws independent posterior samples for every arm, so its exploration dynamics ignore any relation among arms. We study a minimal change to those dynamics: an LLM is queried once for a partition of the arms, the partition becomes a positive-definite correlation matrix ΣΣ through an RBF kernel on cluster ranks, and the per-round posterior sample is drawn with covariance ΣΣ while the Beta posteriors are updated from real rewards only, so the LLM shapes how the sampler moves, not what it believes. We give a self-contained Bayesian regret bound for the idealized Gaussian sampler whose information gain splits into a Klog⁡TK\log T term from the KK-cluster structure and a ridge term that grows to dlog⁡Td\log T: the d/K\sqrt{d/K} improvement over independent sampling is a finite-horizon transient, exact only as the within-cluster correlation tends to one. The correlated sampler reduces regret by 19% over CTS on 16 synthetic Bernoulli families at T=2,500T=2{,}500 (6-7% at T=25,000T=25{,}000 with data-adaptive kernels) and by 41% on the Microsoft MIND-small news benchmark (d=200d=200 real articles), while pseudo-observation warm starts give nothing. An LLM-free ablation with a simulated oracle of controlled quality shows that on unstructured instances the gain is a property of the kernel shape (a random partition, or a plain tempering of the sampling noise, reproduces it), while belief injection at matched oracle quality never helps.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Worst-Case Regret Bounds for Combinatorial Thompson Sampling in Sleeping Semi-Bandits

    May 10, 2026Zhiming Huang, Bingshan Hu, Jianping PanThompson SamplingStochastic Multi-Armed Bandits

  2. Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

    Jan 5, 2026Yifan Zhu, John C. Duchi, Benjamin Van RoyThompson SamplingRegret

  3. MINTS: Minimalist Thompson Sampling

    Jun 1, 2026Kaizheng WangThompson SamplingBayesian