stat.MLApr 24, 2026

Concave Statistical Utility Maximization Bandits via Influence-Function Gradients

Authors: Matías CarrascoAlejandro Cholaquidis

Abstract

We study stochastic multi-armed bandits in which the objective is a statistical functional of the long-run reward distribution, rather than expected reward alone. Under mild continuity assumptions, we show that the infinite-horizon problem reduces to optimizing over stationary mixed policies: each weight vector ww on the simplex induces a mixture law PwP^w, and performance is measured by the concave utility U(w)=U(Pw)U(w)=\mathfrak U(P^w). For differentiable statistical utilities, we use influence-function calculus to derive stochastic gradient estimators from bandit feedback. This leads to an entropic mirror-ascent algorithm on a truncated simplex, implemented through multiplicative-weights updates and plug-in estimates of the influence function. We establish regret bounds that separate the mirror-ascent optimization error from the bias caused by estimating the influence function. The framework is developed for general concave distributional utilities and illustrated through variance and Wasserstein objectives, with numerical experiments comparing exact and plug-in influence-function implementations.

Explore similar work

CardsList
  1. Trading off rewards and errors in multi-armed bandits

    May 1, 2026Akram Erraqabi, Alessandro Lazaric, Michal Valko +2Multi-Armed BanditsRegret

  2. Replicable Bandits with UCB based Exploration

    Apr 21, 2026Rohan Deb, Udaya Ghai, Karan Singh +1BanditsReplication