cs.LGSep 9, 2026

Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits

Authors: Hao LiJie XuZheng Xie

Organizations: College of Science, National University of Defense Technology, Changsha 410073, China

Abstract

Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action sets remains largely unexplored. To address this problem, we propose Meta-LinEXP3, an online-within-online algorithm that constructs a predictable task-level prior from completed tasks to guide the inner LinEXP3 learner. For known context distributions, we develop a policy-centered estimator that achieves an intrinsic-dimension O(n)\mathcal{O}(\sqrt{n}) per-task regret bound. For unknown distributions, we introduce a past-only regularized moment estimator with an O(n2/3)\mathcal{O}(n^{2/3}) leading regret term and explicit finite-sample error. We further establish a direct connection between prior accuracy and transfer regret, showing that increasingly accurate priors yield sublinear transfer-dependent regret across tasks. Experiments demonstrate the effectiveness of Meta-LinEXP3, including its application to structured hyperspectral tensor sampling.

Explore similar work

CardsList
  1. Offline-to-Online Learning in Linear Bandits

    Jun 3, 2026Kushagra Chandak, Toshinori Kitamura, Xiaoqi TanLinear BanditsOffline Reinforcement Learning