Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5× fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.
Figures & tables
Figure 1: Exploitability in one-shot games with pure Nash equilibria using different gradient descent with or without magnet.
Figure 2: Exploitability in one-shot games with pure Nash equilibria, using PPO with and without magnet regularization
Figure 3: Exploitability with 95% confidence intervals of different baselines in one-shot games
Figure 4: Approximate exploitability with 95% confidence intervals of different baselines in sequential games
Figure 5: Head-to-head performance with 95% confidence intervals against different methods during the training.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
One-shot games
Sequential games
PPO
MMD
MMPO
PPO
MMD
MMPO
BR-PPO
PPO epochs
2
2
2
1
1
1
1
Entropy weight
0.05
0.05
0.05
0.05
0.02
0.02
0.01
Magnet weight η
-
0.2
0.2
-
0.2
0.2
-
Magnet update
-
500
500
-
1000
1000
-
Minimal std σmin
-
-
0.001
-
-
0.1
0.001
Appendix
Table 1: Hyperparameters used in experiments for different policy-gradient algorithms and also for computing the approximate exploitability. The amount of components in one-shot games depend on game, in matching pennies it is 1, and in other games it is 4.
NFSP
PSRO
NFSP
PSRO
BR train steps
103
103
104
104
BR batch size
64
64
64
64
PSRO payoff samples
-
256
-
20000
NFSP average steps
400
-
1000
-
Reservoir capacity
106
-
106
-
NFSP reservoir batch
256
-
256
-
Appendix
Table 2: Hyperparameters used in experiments for different best-response based algorithms. Training of the best response uses PPO hyperparameters from Table 1 , besides a different batch size.
Optimizer
Optimistic GD
Optimism
0.333
Learning rate
10−3
Max grad norm
5
Batch size
256
Noise dimension
16
Activation
Mish
Appendix
Table 3: Randomized policy networks hyperparameters used in one-shot games
Figure 6: Exploitability in one-shot games with pure Nash equilibria using gradient descent with or without magnet, using only the gradient estimate.
Figure 7: The evolution of a strategy in continuous matching pennies during training for each of the experimental setting
Figure 8: Exploitability with 95% confidence intervals of different baselines in one-shot games
Figure 9: Approximate exploitability with 95% confidence intervals of different baselines in sequential games
Figure 10: Final exploitability of MMPO and MMD on discretized action space based on the amount of Gaussians or discrete components used.
Figure 11: Final exploitability of MMPO and MMD on discretized action space based on the amount of Gaussians or discrete components used.
PPO epochs
1
Categorical magnet weight ηw
0.2
Gaussian magnet weight ηN
0.02
Magnet update
20000
Minimal std σmin
0.05
Components K
5
Exploration ϵ
0.01
Appendix
Table 4: Hyperparameters used in training of heads-up no limit Texas hold’em
Figure 12: Small blind initial bet with different cards averaged over all suits.
Figure 13: Big blind initial bet after check from the small blind with different cards averaged over all suits.
Figure 14: (left) Average win-rate of different checkpoints against all other checkpoints. (right) Weighted outcomes based on different results of the hand against Slumbot.
Figure 15: All bets made by the final trained checkpoint in self-play over million hands
Figure 16: All pots, which were encountered by the final trained checkpoint in self-play over million hands
Mixture policies theoretically offer greater flexibility than unimodal policies in continuous action reinforcement learning, but the practical benefits of this complexity remain elusive. Mixture policies are notably absent from most state-of-the-art algorithms, raising a fundamental question: Is the added representational overhead useful? We show that increased flexibility can theoretically enhance solution quality and entropy robustness. Yet standard algorithms like SAC do not leverage these advantages. A core issue is the lack of a low-variance reparameterization trick for mixtures, a luxury Gaussian policies enjoy. We propose a marginalized reparameterization (MRP) estimator to address this, proving it offers lower variance than the standard likelihood-ratio (LR) approach. Our experiments across Gym MuJoCo, DeepMind Control Suite, and MetaWorld show that MRP mixture policies significantly outperform their LR ones, and reach parity (sometimes better) with Gaussian counterparts. In addition, we do find several cases where MRP mixture policies exhibit clear empirical advantages. In this paper, we provide a clearer understanding of the trade-offs involved, elevating MRP mixture policies from theoretical curiosity to a practical tool.
Jiamin He, Samuel Neumann, Jincheng Mei +2
University of Alberta & Amii · Google DeepMind · Canada CIFAR AI Chair
Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-player zero-sum imperfect-information games. The uniform distribution has emerged as a strong policy regularization target for this purpose, but it regularizes equally toward all actions regardless of their viability. We introduce EMAgnet, which instead regularizes toward an exponential moving average (EMA) of the last-iterate policy's parameters, providing an adaptive regularization target that evolves with the agent's improving strategy. We evaluate EMAgnet on both standard two-player zero-sum benchmarks and modified benchmarks with exploration challenges and large numbers of strictly dominated strategies. Relative to PPO self-play with uniform-magnet regularization under both linear and power-law annealing schedules, EMAgnet achieves lower exploitability in the majority of tested environments, with consistent performance gains across games containing strictly dominated strategies.
Tristan Maidment, JB Lanier, Chase McDonald +5
1Riot Games · 3New York University · University of California, Irvine
We study reinforcement learning in hybrid discrete-continuous action spaces, such as settings where the discrete component selects a regime (or index) and the continuous component optimizes within it -- a structure common in robotics, control, and operations problems. Standard model-free policy gradient methods rely on score-function (SF) estimators and suffer from severe credit-assignment issues in high-dimensional settings, leading to poor gradient quality. On the other hand, differentiable simulation largely sidesteps these issues by backpropagating through a simulator, but the presence of discrete actions or non-smooth dynamics yields biased or uninformative gradients. To address this, we propose Hybrid Policy Optimization (HPO), which backpropagates through the simulator wherever smoothness permits, using a mixed gradient estimator that combines pathwise and SF gradients while maintaining unbiasedness. We also show how problems with action discontinuities can be reformulated in hybrid form, further broadening its applicability. Empirically, HPO substantially outperforms PPO on inventory control and switched linear-quadratic regulator problems, with performance gaps increasing as the continuous action dimension grows. Finally, we characterize the structure of the mixed gradient, showing that its cross term -- which captures how continuous actions influence future discrete decisions -- becomes negligible near a discrete best response, thereby enabling approximate decentralized updates of the continuous and discrete components and reducing variance near optimality. All resources are available at github.com/MatiasAlvo/hybrid-rl.