Decision-time equilibrium search carried poker to superhuman play, but it has so far relied on tractable subgames: a handful of actions per decision, chance confined to card deals, one player moving at a time. Competitive Pokémon in its official doubles format (VGC) breaks all three assumptions at once. Both players act simultaneously from joint menus in the hundreds, each joint action resolves to hundreds of stochastic outcomes, and the opponent's reserves and stat allocations are hidden. No prior Pokémon agent performs equilibrium search, and whether it scales to this regime was open; we show that it does, and report what it took. PokaiEngine, our Rust battle engine, enumerates a joint action's full weighted outcome distribution in one pass, at ∼99% parity with Pokémon Showdown and a fraction of the cost of sampling it. PokaiTrainer adapts Student of Games to this scale and trains it by self-play over hundreds of human teams. Each decision is solved by counterfactual regret minimization as a Bayesian matrix game over public belief states, subgames grow under an explicit compute budget, and value targets are harvested from the interior of every solve and grounded by realized outcomes. The strength is in the search. The network's policy alone loses even to a shallow heuristic search. PokaiTrainer is, to our knowledge, the first VGC agent rated on the live Showdown ladder. Under open team sheets it wins 59% of 150 best-of-three sets against a human field averaging ∼1320 Elo, holds a 1350-1400 Elo band, and at its peak reached 1492 Elo, entering the format's top 500.
Figures & tables
Chess / Go
HUNL poker
Scotland Yard
No-press Diplomacy
VGC doubles
Simultaneous actors
—
—
—
7 powers
2 players × 2 slots
Hidden information
—
∼103 hands
≤ 199 stations
—
15 brings ×∼107 spreads per Pokémon
Chance
—
card deals only
—
—
102 – 103 outcomes per joint action
Actions per decision
∼ 35 / ∼ 250
a few bet sizes
∼ 10
up to ∼1024 per power; a few dozen searched
∼102 per player, ∼104 joint
Horizon
∼ 80 / ∼ 150 plies
≤ 4 bet rounds
24 rounds
open-ended
5–20 turns
Table 1: VGC compared to the domains decision-time equilibrium search has mastered: the four of Student of Games ( Schmid et al., 2023 ) and no-press Diplomacy ( Gray et al., 2021 ) . The ∼107 spread grid matters less than its size suggests ( Section 4.4 ).
Figure 1: One attack’s three chance events under the two engine modes. Compute rolls each event and continues a single timeline. Search forks the state and carries every branch, holding the 16 damage rolls inside one state as an HP distribution split only where it straddles a knock-out.
Method
Mean time (ms)
Mass covered
Mass misallocated
Search , exhaustive
9.9
100%
0% (reference)
Search , 32-branch cap
1.0
99.5%
0.7%
Compute / Showdown sampling, N=16
1.3 / 36.0
94.3% / 93.0%
10.0% / 10.2%
Compute / Showdown sampling, N=64
4.9 / 132.0
98.4% / 98.2%
4.8% / 4.8%
Compute / Showdown sampling, N=256
19.4 / 509.0
99.7% / 99.4%
2.8% / 2.6%
Table 2: Recovering the outcome distribution of one joint action: Search -mode enumeration versus sampling, averaged over 500 random openings with random joint move actions. Outcomes are distinct when they differ in any end-of-turn discrete features. Coverage is the reference outcome mass the method observed, and misallocation is the total-variation distance from the reference weights. A 32-branch cap is simultaneously faster and more accurate than sampling at any tested budget.
Figure 4
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: System architecture. The engine runs one rules code in Compute or Search mode, the latter exposed as the branch() API returning weighted outcome sets; the Showdown adapter reconstructs the same player-view BattleState the engine emits from protocol lines; and BattleEnv is one seat-view surface over all three backends. Python agents implement one Agent interface and run unchanged on any backend; the search player additionally calls branch() on hypothesized states.
walk over a matrix game; admissions priced by compute cost, budget in cost units ( Section 4.3 )
Value targets
SoG (solver values + TD(1) rows)
per-row λ -mix, grounded on the realized world’s slot only ( Section 4.6 )
Extra supervision
SoG (re-solved leaf queries)
interior nodes of the same solve, reach-sampled and bootstrap-weighted
Off-line coverage
SoG (recursive queries)
hypothetical augmentation, grounded by play
Appendix
Table 4: Where PokaiTrainer follows its lineage and where it departs.
Figure 4: Anatomy of one budgeted solve. Left: the subgame rooted at the public belief state β=(s,b) ; every decision node holds a matrix game, chance nodes are the engine’s exactly enumerated outcome batches T , leaves are priced by vθ , and the red path is one PUCT expansion walk ( Equation 3 ) admitting a new continuation. Right, zooming into a turn node: one payoff matrix Uw per world, weighted by the belief b and solved jointly by CFR ( Equations 1 and 2 ); each cell resolves through one Search -mode engine pass of T , and its successors either recurse or read the leaf value ⟨b,vθ⟩ .
Decision
Chooses
A1 / A2
Hidden
Solve size
Team preview
both
bring-4 × lead pair
none ( ∣W∣=1 )
90×90
Turn
both
move-or-switch per slot
w∼b
∣A1∣×k2×∣W∣ root; k1×k2×∣W∣ interior
Forced switch (after faints)
both
replacement
w∼b
≤2×2×∣W∣
Mid-turn switch (pivot)
one seat
replacement (one slot)
w , opponent’s uncommitted actions H
≤2×∣H∣×∣W∣
Appendix
Table 5: The decision types of a VGC battle as instances of the Bayesian matrix game ( Section 4.2 ), with solve sizes as rows × columns × worlds. ki is seat i ’s menu after shortlisting by the policy prior (root solves keep our full menu; values in Section 5.1 ); replacement menus are bounded by the two-Pokémon bench; ∣W∣≤15 before the leads are revealed and ≤6 after.
Table 6: The eval16 evaluation pool, sorted by strength. Each team is a public paste from the VGCPastes repository ( VGCPastes, 2026 ) (linked; author credit on the paste page), identified by its corpus index. Elo: Bradley–Terry fit over 2,940 internal head-to-head games among the 16 (mean 1500). Italics mark Pokémon holding a Mega Stone.
Table 7: Full standings of the 32-team elite round-robin at the deepened deployment shape. Win% is over 93 games per team; cRk is the team’s rank within the 32 by the corpus-wide fit. Links and italics as in Table 6 .
Arm
Change
Result
Verdict
no grounding
λ=0 , λint=1
50.0% (256/512) h2h, same deploy
wash at 2 rounds; calibration drifts
decoupled + zero-sum
θ′ re-priced targets, moment loss, λ=0
42.6% (109/256) vs sibling r3
loses; dropped
hypothetical forks, 3/game
augmentation only
48.8% (125/256) compute-matched
wash per unit compute
hypothetical forks, 1/game + PUCT
52.0% (133/256) on half the rows
kept
full- z on hypo rows
λ=1 on augmented rows
40.6% (104/256) vs control
loses
entry continuations
replacement + preview roots
52.0% r5, 57.0% r7
kept
Appendix
Table 8: Target and expansion ablations on the v7 deep line; each row is a head-to-head against its round-matched sibling unless noted. Most levers are within a 256-game cell’s ±6 pp of parity: the binding constraint is the value network, not the target construction.
Agent’s solver
Agent’s play
Best-response win% [95% CI]
Gain / decision
CFR
argmax (eval cells)
77.7
[72.2, 82.4]
0.097
CFR
T=0.5 (ladder)
52.3
[46.2, 58.4]
0.027
CFR
T=1 (solved mix)
49.6
[43.5, 55.7]
0.0003
decoupled UCT
argmax
74.6
[68.9, 79.6]
0.193
decoupled UCT
T=0.5
72.7
[66.9, 77.8]
0.110
decoupled UCT
T=1
83.6
[78.6, 87.6]
0.058
Appendix
Table 9: An informed local best response, seated second, against the round-32 agent under each solver and play mode (eval16, 256 games per cell). Gain is the best response’s in-model advantage per decision over the agent’s expectation, in payoff units on [−1,1] . The two solvers are at parity against fixed opponents (51.7%, 2,048 games) and 34pp apart here.