stat.MLApr 28, 2026

Spectral bandits

Authors: Tomáš KocákRémi MunosBranislav KvetonShipra AgrawalMichal Valko

Organizations: ENS de Lyon, 15 Parvis Ren´e Descartes, 69342 Lyon, France · Inria Lille – Nord Europe, SequeL team · DeepMind Paris, 14 Rue de Londres, 75009 Paris, France · Google Research, 1600 Amphitheatre Parkway, Mountain View, CA 94043, United States · Columbia University, West 120th Street, New York, NY, 10027 United States

Abstract

Smooth functions on graphs have wide applications in manifold and semi-supervised learning. In this work, we study a bandit problem where the payoffs of arms are smooth on a graph. This framework is suitable for solving online learning problems that involve graphs, such as content-based recommendation. In this problem, each item we can recommend is a node of an undirected graph and its expected rating is similar to the one of its neighbors. The goal is to recommend items that have high expected ratings. We aim for the algorithms where the cumulative regret with respect to the optimal policy would not scale poorly with the number of nodes. In particular, we introduce the notion of an effective dimension, which is small in real-world graphs, and propose three algorithms for solving our problem that scale linearly and sublinearly in this dimension. Our experiments on content recommendation problem show that a good estimator of user preferences for thousands of items can be learned from just tens of node evaluations.

Explore similar work

CardsList
  1. Spectral bandits for smooth graph functions

    Apr 20, 2026Michal Valko, Rémi Munos, Branislav Kveton +1Multi-Armed BanditsBandits