Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games
Organizations: NLR Royal Netherlands Aerospace Centre Amsterdam, The Netherlands · Data Science Center of Excellence Faculty of Military Sciences, Breda, NL · Departement Intelligent Systems Tilburg University, Tilburg, NL
Abstract
Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution. Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy. This requires an executor that can follow different commands and measurable criteria for assessing whether it does so. We introduce a constrained command-conditioned Proximal Policy Optimization (PPO) policy, the executor, for MicroRTS, a real-time strategy environment. Discrete commands specify strategic objectives and behavioral requirements for economy, army composition, military posture, and worker policy over multiple environment steps; the executor determines the unit-level actions used to fulfill them. A Thompson-sampling bandit acts as the strategist, selecting command tuples from an estimate of the opponent's strategy built from in-game observations rather than opponent identity. In a controlled comparison with a flat PPO baseline trained with the same architecture, budget, curriculum and self-play league, the strategist-executor system wins significantly more often against three of the four strongest opponents on a training map, including the two strongest held-out ones (0.55 to 0.97 and 0.01 to 0.34), with no significant difference against the others.
Figures & tables
| Economy | Composition | Military posture | Worker policy |
|---|---|---|---|
| STABILIZE | LIGHT | OFFENSE | ECON_ONLY |
| GROW | HEAVY | DEFENSE | HARASS |
| HOLD | RANGED | SPLIT | |
| DOUBLE_BARRACKS | LIGHT_HEAVY | FLANK | |
| LIGHT_RANGED | |||
| HEAVY_RANGED |
| Command | Physical metric | Train rate | Reference | Intervention | Response [95% interval] |
|---|---|---|---|---|---|
| GROW | new workers | 0.82 | 0.91 | 2.23 | [ , ] |
| HEAVY | Heavy share of new army | 0.75 | 0.00 | 1.00 | [ , ] |
| RANGED | Ranged share of new army | 0.80 | 0.00 | 1.00 | [ , ] |
| OFFENSE | army position | 0.79 | 0.469 | 0.463 | [ , ] |
| HARASS | share workers in enemy half | 0.79 | 0.109 | 0.088 | [ , ] |
| Win rate | Flat vs Bandit | |||||
|---|---|---|---|---|---|---|
| Rank | Opponent | Split | Flat | Bandit | Static | |
| 1 | RAISocketAI | holdout | 0.010 | 0.344 | 0.470 | |
| 2 | TMA | train | 1.000 | 0.950 | 0.949 | 0.53 |
| 3 | agent_sota | train | 0.612 | 0.929 | 0.898 | |
| 4 | obiBotKenobi | holdout | 0.545 | 0.970 | 1.000 | |
| 5 | mayari | train | 1.000 | 0.980 | 1.000 | 1.00 |
| Most common | First | Games with a change | |||
|---|---|---|---|---|---|
| Opponent | economy sequence (games) | change | Composition | Posture | Worker |
| RAISocketAI | STABILIZE GROW (86) | 440 | 18 (H, L+H) | 0 | 0 |
| TMA | STABILIZE GROW (97) | 360 | 2 (H) | 0 | 0 |
| agent_sota | STABILIZE GROW (94) | 360 | 9 (H, L+H) | 0 | 0 |
| obiBotKenobi | STABILIZE GROW HOLD (50) | 360 | 2 (H) | 0 | 0 |
| mayari | STABILIZE DOUBLE_BARRACKS (46) | 441 | 3 (R) | 0 | 0 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Command | Success | Failure | |
| STABILIZE | Base, four workers, and a barracks; recovery exceptions below. | Unauthorized second barracks. | 450 |
| GROW | Living worker count increases by at least one from window start. | Unauthorized second barracks. | 300 |
| HOLD | No prohibited worker production for 140 ticks; a base or barracks survives. | New worker production without a worker deficit, or unauthorized second barracks. | 250 |
| DOUBLE_BARRACKS | A base and at least two barracks survive. | — | 1200 |
| LIGHT, HEAVY, RANGED | At least three new military units; requested type accounts for . | At least three new units and non-target types. | 700, 800, 700 |
| LIGHT_HEAVY, LIGHT_RANGED, HEAVY_RANGED | Both requested types produced; each accounts for , together . | Both types produced, at least four new units, and non-target types. | 1300, 1200, 1300 |
| Opponent | A | B | C | D | E |
|---|---|---|---|---|---|
| randomBiasedAI | 24 | 18 | 8 | 2 | 2 |
| workerRushAI | 16 | 8 | 4 | 4 | 4 |
| lightRushAI | – | 10 | 8 | 4 | 4 |
| rojo | – | 4 | 4 | – | – |
| mayari | – | – | 16 | 14 | 8 |
| coacAI | – | – | – | 8 | 14 |
| Command | Command | ||||
|---|---|---|---|---|---|
| STABILIZE | 0.05 | 0.94 | LIGHT_RANGED | 0.05 | 0.92 |
| GROW | 0.20 | 0.82 | HEAVY_RANGED | 0.05 | 0.91 |
| HOLD | 0.45 | 0.79 | OFFENSE | 0.24 | 0.79 |
| DOUBLE_BARRACKS | 0.05 | 1.00 | DEFENSE | 0.05 | 0.96 |
| LIGHT | 0.83 | 0.80 | SPLIT | 0.05 | 1.00 |
| HEAVY | 0.47 | 0.75 | FLANK | 0.29 | 0.92 |
| Contract success | Wins/losses | |||||
|---|---|---|---|---|---|---|
| Setting | coacAI | droplet | obiBotKenobi | coacAI | droplet | obiBotKenobi |
| reference | — | — | — | 8/0 | 8/0 | 7/0 |
| GROW | 4/8 | 3/7 | 8/8 | 7/1 | 7/0 | 8/0 |
| HEAVY | 8/8 | 3/6 | 3/6 | 0/8 | 6/0 | 4/2 |
| RANGED | 8/8 | 4/7 | 7/8 | 8/0 | 7/0 | 1/7 |
| OFFENSE | 0/8 | 3/8 | 0/7 | 8/0 | 8/0 | 7/0 |
| Opponent | Split | Pooled | 95% CI | A | D | I | as P1 | Win |
|---|---|---|---|---|---|---|---|---|
| RAISocketAI | holdout | |||||||
| TMA | train | |||||||
| agent_sota (map A) | train | – | – | – | ||||
| obiBotKenobi | holdout | |||||||
| mayari | train | |||||||
| coacAI | train |
| Flat | Bandit | Static | |||
|---|---|---|---|---|---|
| Rank | Opponent | W/L/D | W/L/D | Selected tuple | W/L/D |
| 1 | RAISocketAI | 1/98/1 | 33/63/4 | GROW+L | 47/53/0 |
| 2 | TMA | 100/0/0 | 95/5/0 | DOUBLE+L+R | 94/5/1 |
| 3 | agent_sota (map A) | 60/38/2 | 92/7/1 | STABILIZE+L | 88/10/2 |
| 4 | obiBotKenobi | 54/45/1 | 97/3/0 | GROW+L | 100/0/0 |
| 5 | mayari | 100/0/0 | 98/2/0 | STABILIZE+L | 100/0/0 |
| Rank | Opponent | L | H | R | L+H | L+R | H+R | HOLD |
|---|---|---|---|---|---|---|---|---|
| 1 | RAISocketAI | 0.50 | 0.17 | 0.03 | 0.33 | 0.14 | 0.07 | 0.02 |
| 2 | TMA | 0.95 | 0.97 | 0.97 | 0.94 | 0.95 | 0.95 | 0.28 |
| 3 | agent_sota | 0.83 | 0.42 | 0.25 | 0.78 | 0.71 | 0.28 | 0.30 |
| 4 | obiBotKenobi | 0.97 | 0.58 | 0.11 | 0.82 | 0.77 | 0.36 | 0.38 |
| 5 | mayari | 0.98 | 0.99 | 1.00 | 1.00 | 1.00 | 1.00 | 0.95 |
| 6 | coacAI | 0.73 | 0.00 | 1.00 | 0.26 | 0.92 | 0.69 | 0.45 |