cs.LGNov 10, 2025

Scalable Decision Making for Games of Imperfect Information

Authors: Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter, Gabriele Farina

Organizations: Carnegie Mellon University · NYU Tandon School of Engineering · Stanford University · Massachusetts Institute of Technology

Abstract

Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts, top-human-level play at Stratego---a board wargame with hidden information on a massive scale---has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin---achieving, to our knowledge, the first superhuman result in the game's history---while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.

Figures & tables

Explore similar work

CardsList
  1. Superhuman AI for Generals.io Using Self-Play Reinforcement Learning

    Jun 22, 2026Matej Straka, Viliam Lisý, Martin SchmidSelf-PlaySimulation-Based Reinforcement Learning

  2. SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing

    Sep 28, 2026Zhiwei Chen, Tianchun Wang, Zhongtao Rao +3Strategic ReasoningLarge Language Model Agents

  3. Study and Improvement of Search Algorithms in Multi-Player Perfect-Information Games

    Apr 19, 2026Quentin Cohen-SolalImperfect-Information GamesMinimax