cs.GTOct 7, 2022

Stackelberg POMDP: Learning to Lead via Reinforcement Learning

Authors: Matthias Gerstgrasser, Gianluca Brero, Alon Eden, Darshan Chakrabarti, Nicolas Lepore, Vincent Li, Eric Mibuari, Amy Greenwald, +1 more

Organizations: John A. Paulson School of Engineering and Applied Sciences, Harvard University · Data Science Initiative, Brown University · School of Computer Science and Engineering, Hebrew University of Jerusalem · Department of Industrial Engineering and Operations Research, Columbia University · Department of Computer Science, Brown University

Abstract

Many real-world domains--including e-commerce platform design, security planning, and multi-agent coordination--feature leader-follower problems where one decision-maker commits to a policy and others react strategically. We develop a reinforcement learning framework for such interactions in sequential environments with partial observations and multiple followers. Followers may adapt through no-regret learning or reinforcement learning, potentially departing from equilibrium behavior. The framework embeds follower adaptation into the leader's environment to construct a single-agent partially observable Markov decision process--the Stackelberg POMDP. For policy-interactive response algorithms, which access the leader's policy through queries, we prove that an optimal policy based only on the leader's game history yields an optimal commitment under the specified response procedure. We use proximal policy optimization with a centralized critic and train contextual meta-followers to respond across leader policies. In indirect mechanism design, mechanisms using buyer messages achieve higher social welfare than optimal standard sequential price mechanisms across all tested type counts, with responses certified as approximate Bayesian coarse correlated equilibria. In platform design, learned display rules increase mean consumer surplus by 8.4% over an optimized fixed price cap while accommodating hidden seller costs. In Atari bilateral trade, meta-learned follower responses support joint learning of visual gameplay and economic decisions; assigning leadership to the seller or buyer shifts transaction prices and payoffs in that agent's favor. Controlled ablations examine how response credit, policy consistency, and reward timing affect learning.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning in Structured Stackelberg Games

    Apr 11, 2025Maria-Florina Balcan, Kiriaki Fragkia, Keegan HarrisGame TheoryStackelberg Games

  2. Learnable Randomization as Commitment Against Adaptive Optimizers

    Sep 26, 2026Zihan Deng, Chuanzhi Xu, Xiaozhen Zhong +2Policy OptimizationStackelberg Games

  3. Minimax-Optimal Policy Regret in Partially Observable Markov Games

    Jun 1, 2026Raman AroraRegret Minimization in RLImperfect-Information Games