PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning
Organizations: Nanyang Technological University, Singapore. · Cornell University, USA. · University of Bristol, UK.
Abstract
Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: https://github.com/powerzoojax/PowerZooJax.
Figures & tables
| Tasks | Power System Problems | RL Challenges |
|---|---|---|
| GenCos | Decentralized electricity market bidding with a 3-dim action space and a 12-dim observation space. | Multi-agent RL with partial observability (each agent’s local information). |
| TSO | Centralized unit commitment with a 108-dim action space and a 410-dim observation space. | Single-agent RL with a large, mixed action space of continuous and binary variables. |
| DSO | Centralized control of flexible loads with a 12-dim action space and a 195-dim observation space. | Single-agent RL with an easy-to-start setup to verify initial design ideas. |
| DERs | Decentralized control of three types of DERs (battery, PV, flexible loads) with a 2-dim action space and a 15-dim observation space. | Multi-agent RL with partial observability (each agent’s local and neighbor information) and heterogeneous agents. |
| DCMG | Centralized long-horizon scheduling with a 5-dim action space and a 24-dim observation space. | Single-agent RL with a long-horizon sequential decision making process and multi-objectives. |
| Wall-clock to matched budget (s) | Throughput at envs (steps/s) | |||||||||
| Task | Budget | PowerZooJax | SBX | SB3 | speedup | PowerZooJax | SBX | SB3 | speedup | |
| GenCos | 5M | 256 | 67 | 9,602 | 10,868 | 161 | 506,134 | 442 | 244 | 2,077 |
| TSO | 20M | 256 | 515 | 10,799 | 10,801 | 21 | 53,724 | 1,649 | 411 | 131 |
| DSO | 3M | 128 | 32 | 1,667 | 3,801 | 117 | 134,660 | 3,435 | 2,207 | 61 |
| DERs | 10M | 128 | 79 | 61,883 | 59,778 | 780 | 414,506 | 203 | 233 | 2,040 |
| DCMG | 1M | 64 | 33 | 221 | 257 | 8 | 25,132 | 1,167 | 1,298 | 22 |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| CMDP and policy | |
| State and action spaces of one task. | |
| Transition kernel; deterministic given exogenous time-series . | |
| Scalar reward and per-step cost vector with task-specific entries. | |
| Discount factor and episode horizon ( or ). | |
| Cost thresholds (Form 1) or zero-violation set (Form 2). | |
| Kernel | Used by | Output |
|---|---|---|
| Radial AC PF (DistFlow recurrence) | DSO, DERs | Branch flows, bus voltages, network loss |
| DC OPF (ADMM) | TSO | Generator dispatch, line flows, operating cost |
| Bid-based market clearing (PDIPM) | GenCos | Cleared dispatch, locational marginal prices, profit |
| Resource device dynamics | DSO, DERs, DCMG | SoC, PQ injections, flexible-load shifting, diesel output |
| DCMG balance and thermal coupling | DCMG | IT power, cooling demand, grid import, SLA tracking, carbon emission |
| Task | Decision problem | Primary metric | Cost channels |
|---|---|---|---|
| GenCos | 5-agent strategic bidding on 5-bus market, 48 half-hour steps | Total per-agent profit | Thermal overload (diagnostic) |
| TSO | 1-agent SCUC on 118-bus system, 48 half-hour steps | Operating cost | Reserve, thermal (Form 2) |
| DSO | 1-agent demand response on 33-bus feeder, 48 half-hour steps | Network loss (MWh) | Voltage band |
| DERs | 12-agent cooperative DER control on 141-bus feeder, 48 half-hour steps | Active power loss (MW) | Voltage band, thermal, resource |
| DCMG | 1-agent data center microgrid scheduling, 288 five-minute steps | Episode return | SLA, over-temperature, power balance |
| Task | System / case | Time-series source(s) | Splits | Seed budget |
|---|---|---|---|---|
| GenCos | 5-bus market | GB demand | train, in-distribution , demand shift, renewable shock | 5 seeds, 30 episodes/seed |
| TSO | 118-bus transmission | GB demand, wind, and solar | train, in-distribution , load stress, line tightening | 5 seeds, 50 episodes/seed |
| DSO | 33-bus feeder | Ausgrid distribution demand | in-distribution | 5 seeds, 50 episodes/seed |
| DERs | 141-bus feeder | Ausgrid distribution demand and GB solar | train, in-distribution , voltage tightening, PV shift, load stress | 5 seeds, 30 episodes/seed |
| DCMG | Behind-the-meter microgrid | Google workload (Azure and Alibaba for workload-OOD splits); GB solar; GB MID price | train, in-distribution , cooling stress, renewable drought; appendix: workload swap, workload shock, dg derating, SLA tightening | 5 seeds, 10 episodes/seed |
| Split type | Tasks / split names | Physical mechanism | Role |
|---|---|---|---|
| Routine in-distribution | TSO, DSO, DERs, GenCos, DCMG / in-distribution | Held-out episodes from the routine evaluation regime | Main task comparison for the primary metric |
| Load / demand shift | TSO / load stress; DERs / load stress; GenCos / demand shift | Higher or shifted demand changes reserve pressure, feeder loading, or market clearing | Whether an in-distribution policy remains useful when demand changes |
| Renewable / PV shift | DERs / PV shift; GenCos / renewable shock; DCMG / renewable drought | Renewable generation changes the controllable/exogenous balance | Sensitivity to renewable availability |
| Network / security tightening | TSO / line tightening; DERs / voltage tightening | Tighter physical operating limits (line capacity or voltage band) | Whether lower cost or lower loss survives under tighter operating limits |
| Cooling / workload stress | DCMG / cooling stress (main); workload swap, workload shock, SLA tightening (appendix) | Data center thermal load, workload mix, or service deadlines change | Long-horizon resource scheduling under thermal, workload, or service stress |
| Resource-availability stress | DCMG / dg derating | Backup-generation capacity is reduced | Sensitivity to reduced backup capacity |
| Task | Algo | LR | HD | CL | EC | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| GenCos | IPPO | 5 | 256 | 48 | 4 | 3e-4 | [128,128] | 0.2 | 0.01 | 0.995 | 0.95 |
| TSO | PPO | 20 | 256 | 48 | 4 | 3e-4 | [256,256] | 0.2 | 0.01 | 0.995 | 0.95 |
| TSO | PPO-Lag | 20 | 256 | 48 | 4 | 5e-5 | [256,256] | 0.2 | 0.005 | 0.995 | 0.95 |
| DSO | PPO | 3 | 128 | 48 | 4 | 3e-4 | [128,128] | 0.2 | 0.01 | 0.995 | 0.95 |
| DSO | SAC | 3 | 64 | — | 1 | 3e-4 | [128,128] | — | — | 0.995 | — |
| DSO | Sauté PPO | 3 | 128 | 48 | 4 | 3e-4 | [128,128] | 0.2 | 0.01 | 0.995 | 0.95 |
| Method | Profit mean (GBP) | Profit 95% CI (GBP) | HHI |
|---|---|---|---|
| Truthful | 6,934 | 0.5664 | |
| Uniform-mid | 302,972 | 0.5663 | |
| IPPO | 395,030 | 0.5384 | |
| Max-markup | 597,864 | 0.5662 |
| Split | Truthful | Uniform-mid | IPPO | Max-markup |
|---|---|---|---|---|
| In-distribution | 6,934 | 302,972 | 395,030 | 597,864 |
| Demand shift | 9,634 | 354,978 | 475,620 | 702,906 |
| Renewable shock | 7,977 | 329,873 | 434,642 | 650,409 |
| Split | Method | Operating cost (£M/day) | Reserve-shortfall rate | Thermal-overload rate |
|---|---|---|---|---|
| in-distribution | PPO | 1.58 | 0.0728 | 0.3554 |
| in-distribution | Merit Order | 2.89 | 0.0096 | 0.1242 |
| in-distribution | PPO-Lag | 3.50 | 0.0000 | 0.0593 |
| in-distribution | All-On | 4.34 | 0.0000 | 0.0346 |
| line-tightening | PPO | 1.57 | 0.0547 | 0.4932 |
| line-tightening | Merit Order | 2.89 | 0.0096 | 0.3371 |
| Method | Total loss (MWh) | Voltage violations / step | Loss reduction |
|---|---|---|---|
| PPO | 1.9355 | 0.00033 | 34.45% |
| Sauté PPO | 1.9370 | 0.00292 | 34.35% |
| SAC | 2.3178 | 0.05475 | 20.54% |
| PPO-Lag | 2.4108 | 0.01733 | 15.84% |
| Droop | 2.8593 | 0.07042 | 0.74% |
| No control | 2.8909 | 0.20667 | 0.00% |
| Method | Active loss (MW) | Violation steps | Violation rate |
|---|---|---|---|
| IPPO-rs | 0.1974 | 0.00 | 0.0000 |
| IPPO | 0.1974 | 0.00 | 0.0000 |
| Voltage droop | 0.2031 | 0.00 | 0.0000 |
| IPPO-Lag | 0.2051 | 0.0067 | 0.00014 |
| No control | 0.2052 | 0.00 | 0.0000 |
| Method | Active loss (MW) | Violation steps | Violation rate |
|---|---|---|---|
| IPPO-rs | 0.1974 | 4.91 | 0.1022 |
| IPPO | 0.1974 | 4.91 | 0.1024 |
| Voltage droop | 0.2031 | 5.97 | 0.1243 |
| IPPO-Lag | 0.2051 | 7.94 | 0.1654 |
| No control | 0.2052 | 8.73 | 0.1819 |
| Method | Episode return | SLA rate | Spill cost |
|---|---|---|---|
| SAC | |||
| PPO | |||
| No control | |||
| Rule-based | |||
| Max renewable |
| Task | PowerZooJax | SBX/CUDA | SB3/CUDA | vs SBX |
|---|---|---|---|---|
| GenCos | 506,134 | 442 | 244 | 1,145 |
| DERs | 414,506 | 203 | 233 | 2,040 |
| DSO | 134,660 | 3,435 | 2,207 | 39 |
| TSO | 53,724 | 1,649 | 411 | 33 |
| DCMG | 25,132 | 1,167 | 1,298 | 22 |
| Task | Measured nenv range | Compile time | JAX throughput at |
|---|---|---|---|
| DCMG | 16–256 | 8.2 s | 307.7k steps/s |
| DSO | 16–256 | 2.9 s | 388.4k steps/s |
| DERs | 16–256 | 2.1 s | 283.8k steps/s |