cs.LGOct 7, 2026
Savem-Set Adversarial Bandits with Winner Feedback
Organizations: Universit`a degli Studi di Milano · Politecnico di Milano
Abstract
We show upper and lower bounds on the regret of -set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations). By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates. Our main technical contributions are the information-theoretic lower bounds on the regret. Experiments on synthetic data confirm our theoretical analyses.
Figures & tables
| Utility | Feedback | Name | Regret | Reference |
| Sum of | ||||
| rewards | Sum of | |||
| rewards | Full-bandit | Bubeck et al. (2012) | ||
| Ito et al. (2019) | ||||
| \SetRow bg=rowcolor Sum of | ||||
| rewards | Reward of |
Table 1: Old and new results for -set bandits.
Figure 1: Results of no-regret algorithms with their respective feedback on the stationary correlated reward instance. On the left (a), total regret per round, averaged over 10 seeds; on the right (b), final total regret after rounds for increasing action size , averaged over 5 seeds; 95% bootstrap confidence intervals are shown.
Figure 2: Results on corrupted rewards. Total regret per round, averaged over 10 seeds, with 95% bootstrap confidence intervals.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Results on stationary correlated rewards. Total regret per round, averaged over 100 independent reward sequences, with 99% bootstrap confidence intervals.
Figure 4: Results on stationary uniform rewards. Total regret per round, averaged over 10 seeds, with 95% bootstrap confidence intervals.
Figure 5: Results on nonstationary rewards. Total dynamic regret per round, averaged over 10 seeds, with 95% bootstrap confidence intervals.
Figure 6: Results on binary rewards. Total regret per round, averaged over 10 seeds, with 95% bootstrap confidence intervals.