cs.LGAug 29, 2026

PokaiTrainer: Scaling Equilibrium Search to Competitive Pokémon VGC

Authors: Max Yu

Organizations: Independent Researcher

Abstract

Decision-time equilibrium search carried poker to superhuman play, but it has so far relied on tractable subgames: a handful of actions per decision, chance confined to card deals, one player moving at a time. Competitive Pokémon in its official doubles format (VGC) breaks all three assumptions at once. Both players act simultaneously from joint menus in the hundreds, each joint action resolves to hundreds of stochastic outcomes, and the opponent's reserves and stat allocations are hidden. No prior Pokémon agent performs equilibrium search, and whether it scales to this regime was open; we show that it does, and report what it took. PokaiEngine, our Rust battle engine, enumerates a joint action's full weighted outcome distribution in one pass, at ∼99%{\sim}99\% parity with Pokémon Showdown and a fraction of the cost of sampling it. PokaiTrainer adapts Student of Games to this scale and trains it by self-play over hundreds of human teams. Each decision is solved by counterfactual regret minimization as a Bayesian matrix game over public belief states, subgames grow under an explicit compute budget, and value targets are harvested from the interior of every solve and grounded by realized outcomes. The strength is in the search. The network's policy alone loses even to a shallow heuristic search. PokaiTrainer is, to our knowledge, the first VGC agent rated on the live Showdown ladder. Under open team sheets it wins 59% of 150 best-of-three sets against a human field averaging ∼1320{\sim}1320 Elo, holds a 1350-1400 Elo band, and at its peak reached 1492 Elo, entering the format's top 500.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?

    May 28, 2026Dongdong Hua, Yifei Sun, Renhong Huang +3Self-Evolving AgentsAgentic Benchmarks

  2. PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers

    May 28, 2026Boning Li, Baoxiang Wang, Longbo HuangPoker

  3. AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play

    May 9, 2026Vlad Murgoci, Matthijs Spaan, Yaniv OrenPokerImperfect-Information Games