stat.MLSep 28, 2026

Elicitation and Decision Geometry in Single-Index Bandits

Authors: Sakshi Arya, Cheng Soon Ong

Organizations: Case Western Reserve University, USA · CSIRO, Australia

Abstract

We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (NBL), a greedy procedure that uses a sequential Stein contrast to learn the optimal boundary directly, without estimating the reward functions or the common link. We characterize the local Riemannian dynamics of NBL through a decision stability coefficient balancing arm separation, link geometry, and the context distribution. We show that this stability is connected to the elicitation geometry of the underlying convex potential. Under local decision stability, NBL contracts toward the optimal boundary and achieves O(log⁡n)O(\log n) expected regret. Numerical experiments illustrate the predicted stability regimes and compare NBL with a parametric greedy benchmark under link misspecification.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Optimal Regret for Single Index Bandits

    May 10, 2026Devdan Dey, Sujoy Bhore, Avishek GhoshStochastic Multi-Armed BanditsOptimal Regret

  2. Active Context Selection Improves Simple Regret in Contextual Bandits

    May 19, 2026Mohammad Shahverdikondori, Jalal Etesami, Negar KiyavashAdaptive Sampling

  3. Trading off rewards and errors in multi-armed bandits

    May 1, 2026Akram Erraqabi, Alessandro Lazaric, Michal Valko +2Stochastic Multi-Armed BanditsRegret