Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5× fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.
Figures & tables
Figure 1: Exploitability in one-shot games with pure Nash equilibria using different gradient descent with or without magnet.
Figure 2: Exploitability in one-shot games with pure Nash equilibria, using PPO with and without magnet regularization
Figure 3: Exploitability with 95% confidence intervals of different baselines in one-shot games
Figure 4: Approximate exploitability with 95% confidence intervals of different baselines in sequential games
Figure 5: Head-to-head performance with 95% confidence intervals against different methods during the training.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
One-shot games
Sequential games
PPO
MMD
MMPO
PPO
MMD
MMPO
BR-PPO
PPO epochs
2
2
2
1
1
1
1
Entropy weight
0.05
0.05
0.05
0.05
0.02
0.02
0.01
Magnet weight η
-
0.2
0.2
-
0.2
0.2
-
Magnet update
-
500
500
-
1000
1000
-
Minimal std σmin
-
-
0.001
-
-
0.1
0.001
Appendix
Table 1: Hyperparameters used in experiments for different policy-gradient algorithms and also for computing the approximate exploitability. The amount of components in one-shot games depend on game, in matching pennies it is 1, and in other games it is 4.
NFSP
PSRO
NFSP
PSRO
BR train steps
103
103
104
104
BR batch size
64
64
64
64
PSRO payoff samples
-
256
-
20000
NFSP average steps
400
-
1000
-
Reservoir capacity
106
-
106
-
NFSP reservoir batch
256
-
256
-
Appendix
Table 2: Hyperparameters used in experiments for different best-response based algorithms. Training of the best response uses PPO hyperparameters from Table 1 , besides a different batch size.
Optimizer
Optimistic GD
Optimism
0.333
Learning rate
10−3
Max grad norm
5
Batch size
256
Noise dimension
16
Activation
Mish
Appendix
Table 3: Randomized policy networks hyperparameters used in one-shot games
Figure 6: Exploitability in one-shot games with pure Nash equilibria using gradient descent with or without magnet, using only the gradient estimate.
Figure 7: The evolution of a strategy in continuous matching pennies during training for each of the experimental setting
Figure 8: Exploitability with 95% confidence intervals of different baselines in one-shot games
Figure 9: Approximate exploitability with 95% confidence intervals of different baselines in sequential games
Figure 10: Final exploitability of MMPO and MMD on discretized action space based on the amount of Gaussians or discrete components used.
Figure 11: Final exploitability of MMPO and MMD on discretized action space based on the amount of Gaussians or discrete components used.
PPO epochs
1
Categorical magnet weight ηw
0.2
Gaussian magnet weight ηN
0.02
Magnet update
20000
Minimal std σmin
0.05
Components K
5
Exploration ϵ
0.01
Appendix
Table 4: Hyperparameters used in training of heads-up no limit Texas hold’em
Figure 12: Small blind initial bet with different cards averaged over all suits.
Figure 13: Big blind initial bet after check from the small blind with different cards averaged over all suits.
Figure 14: (left) Average win-rate of different checkpoints against all other checkpoints. (right) Weighted outcomes based on different results of the hand against Slumbot.
Figure 15: All bets made by the final trained checkpoint in self-play over million hands
Figure 16: All pots, which were encountered by the final trained checkpoint in self-play over million hands