Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: https://github.com/powerzoojax/PowerZooJax.
Figures & tables
Figure 1: Execution paradigm of PowerZooJax.
Figure 2: Overview of a modern power system model spanning generation, transmission, distribution, end-use DERs, and a data center microgrid.
Tasks
Power System Problems
RL Challenges
GenCos
Decentralized electricity market bidding with a 3-dim action space and a 12-dim observation space.
Multi-agent RL with partial observability (each agent’s local information).
TSO
Centralized unit commitment with a 108-dim action space and a 410-dim observation space.
Single-agent RL with a large, mixed action space of continuous and binary variables.
DSO
Centralized control of flexible loads with a 12-dim action space and a 195-dim observation space.
Single-agent RL with an easy-to-start setup to verify initial design ideas.
DERs
Decentralized control of three types of DERs (battery, PV, flexible loads) with a 2-dim action space and a 15-dim observation space.
Multi-agent RL with partial observability (each agent’s local and neighbor information) and heterogeneous agents.
DCMG
Centralized long-horizon scheduling with a 5-dim action space and a 24-dim observation space.
Single-agent RL with a long-horizon sequential decision making process and multi-objectives.
Table 1: Interrelationship between power system problems and RL challenges.
Figure 3: Cross-backend speed for five benchmark tasks. Top: wall-clock training time under matched budgets. Bottom: environment throughput with increasing parallelism, shown on log–log axes.
Wall-clock to matched budget (s)
Throughput at 212 envs (steps/s)
Task
Budget
Nenvs
PowerZooJax
SBX
SB3
speedup
PowerZooJax
SBX
SB3
speedup
GenCos
5M
256
67
9,602
10,868
161 ×
506,134
442
244
2,077 ×
TSO
20M
256
515
10,799
10,801
21 ×
53,724
1,649
411
131 ×
DSO
3M
128
32
1,667
3,801
117 ×
134,660
3,435
2,207
61 ×
DERs
10M
128
79
61,883
59,778
780 ×
414,506
203
233
2,040 ×
DCMG
1M
64
33
221
257
8 ×
25,132
1,167
1,298
22 ×
Table 2: Cross-backend speed per benchmark task for three training backends.
Figure 4: Generation company profit and market clearing price on case5 . (a) Total daily profit for each strategy and evaluation split. (b) Locational marginal price (LMP) vs system load on a representative evaluation. Error bars in (a) show 95% CIs across 5 seeds.
Figure 5: TSO economic–safety cost trade-off on case118 . (a) Operating cost vs thermal overload rate. (b) Operating cost vs reserve shortfall rate. Error bars show 95% CIs across 5 seeds.
Figure 6: DCMG daily scheduling on a representative day. (a) Power supply to the IT workload. (b) Daily battery state of charge (SoC) and grid electricity price.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
CMDP and policy
S,A
State and action spaces of one task.
P
Transition kernel; deterministic given exogenous time-series .
r,c∈Rk
Scalar reward and per-step cost vector with k task-specific entries.
γ,T
Discount factor and episode horizon ( T=48 or T=288 ).
b∈R≥0k
Cost thresholds (Form 1) or zero-violation set (Form 2).
Appendix
Table 3: Notation used across the main paper and the appendix.
IT power, cooling demand, grid import, SLA tracking, carbon emission
Appendix
Table 4: Physical kernels reused across tasks.
Task
Decision problem
Primary metric
Cost channels
GenCos
5-agent strategic bidding on 5-bus market, 48 half-hour steps
Total per-agent profit
Thermal overload (diagnostic)
TSO
1-agent SCUC on 118-bus system, 48 half-hour steps
Operating cost
Reserve, thermal (Form 2)
DSO
1-agent demand response on 33-bus feeder, 48 half-hour steps
Network loss (MWh)
Voltage band
DERs
12-agent cooperative DER control on 141-bus feeder, 48 half-hour steps
Active power loss (MW)
Voltage band, thermal, resource
DCMG
1-agent data center microgrid scheduling, 288 five-minute steps
Episode return
SLA, over-temperature, power balance
Appendix
Table 5: One-line task overview.
Figure 7: IEEE 5-bus market test system ( case5 ) used by the GenCos task: 5 buses B1–B5 (horizontal bars), 6 transmission lines L1–L6, 5 thermal generators G1–G5 (circled symbols, one per agent; G1 and G2 share bus B1, while G3, G4, G5 sit at buses B3, B4, B5 respectively), and 3 load sites D2, D3, D4 (outward arrows at buses B2, B3, B4). Each step, the SCED clearing of Section B.3 returns the per-bus LMPs and per-unit dispatches that drive each agent’s profit.
Figure 8: One-line diagram of the IEEE 118-bus transmission test system ( case118 ) used by the TSO task: 118 buses (horizontal bars), 186 branches (lines connecting buses), 54 thermal generators (circled G; arrows pointing into their bus), and 91 load sides (arrows pointing out of their bus). The TSO task in Section C.2 commits and dispatches the 54 generators on a 30-minute cadence under DC-OPF redispatch. Source: IIT Power Group, 2003.
Figure 9: IEEE 33-bus radial distribution network ( case33bw ) used by the DSO task. The substation (factory icon) at bus 1 supplies the network through a main feeder (buses 2–18) and three lateral branches (19–22, 23–25, 26–33). Gray house icons mark passive loads; orange house icons mark the six controllable flexible loads at buses 6, 14, 18, 22, 28, and 33. The DSO task in Section C.3 sets curtailment and load-shift fractions for these six devices on a 30-minute cadence under radial AC power flow.
Figure 10: 141-bus radial distribution feeder ( case141 ) used by the DERs task. The substation (factory icon) at bus 1 supplies the network through deep radial branches. Twelve DER agents are sited at fixed buses: 4 batteries (BESS, teal vertical bars) at buses 9, 17, 55, 122; 4 PV inverters (sun-panel icons) at buses 6, 72, 73, 82; and 4 flexible loads (orange houses) at buses 24, 41, 70, 135. Small black dots mark passive buses without an attached DER agent. The 12 agents share a team reward (negative network active power loss) and cooperate under partial observability with K=4 BFS-graph-neighbor windows (Section C.4 ).
Table 6: External time-series traces, splits, and seed and episode budgets per task. GB denotes Great Britain. Bold marks the primary split used for the headline metric reported in Section G ; for DCMG the appendix-only splits follow the semicolon.
Split type
Tasks / split names
Physical mechanism
Role
Routine in-distribution
TSO, DSO, DERs, GenCos, DCMG / in-distribution
Held-out episodes from the routine evaluation regime
Data center thermal load, workload mix, or service deadlines change
Long-horizon resource scheduling under thermal, workload, or service stress
Resource-availability stress
DCMG / dg derating
Backup-generation capacity is reduced
Sensitivity to reduced backup capacity
Appendix
Table 7: Split and stress-test taxonomy. Each row states the physical mechanism and the role of the split.
Task
Algo
Ttot
Nenvs
ns
E
LR
HD
CL
EC
γ
λ
GenCos
IPPO
5
256
48
4
3e-4
[128,128]
0.2
0.01
0.995
0.95
TSO
PPO
20
256
48
4
3e-4
[256,256]
0.2
0.01
0.995
0.95
TSO
PPO-Lag
20
256
48
4
5e-5
[256,256]
0.2
0.005
0.995
0.95
DSO
PPO
3
128
48
4
3e-4
[128,128]
0.2
0.01
0.995
0.95
DSO
SAC
3
64
—
1
3e-4
[128,128]
—
—
0.995
—
DSO
Sauté PPO
3
128
48
4
3e-4
[128,128]
0.2
0.01
0.995
0.95
Appendix
Table 8: Master hyperparameter table. Columns: Ttot total environment steps (millions); Nenvs parallel envs; ns steps-per-update; E PPO epochs (SAC update epochs); LR learning rate; HD hidden dims; CL clip-eps; EC entropy coef; γ discount; λ GAE lambda. Optimizer is Adam throughout.
Method
Profit mean (GBP)
Profit 95% CI (GBP)
HHI
Truthful
6,934
[6,412,7,432]
0.5664
Uniform-mid
302,972
[297,376,309,015]
0.5663
IPPO
395,030
[291,174,499,652]
0.5384
Max-markup
597,864
[587,181,609,873]
0.5662
Appendix
Table 9: GenCos in-distribution baseline comparison on case5 , 5 seeds, 30 evaluation episodes per seed. Higher total profit is better. Profit is the across-seed mean with a 95% bootstrap CI (Section F.4 ). HHI denotes the Herfindahl–Hirschman Index of cleared-energy shares.
Split
Truthful
Uniform-mid
IPPO
Max-markup
In-distribution
6,934
302,972
395,030
597,864
Demand shift
9,634
354,978
475,620
702,906
Renewable shock
7,977
329,873
434,642
650,409
Appendix
Table 10: GenCos total-profit mean (GBP) across splits. Each entry aggregates 5 seeds and 30 evaluation episodes per seed.
Figure 11: GenCos behavior on a representative in-distribution episode under the same physical load. Top-left: locational marginal price as a function of system load. Top-right: cumulative profit per step. Bottom-left: Herfindahl–Hirschman Index of cleared-energy shares. Bottom-right: per-step ramp-binding indicator.
Split
Method
Operating cost (£M/day)
Reserve-shortfall rate
Thermal-overload rate
in-distribution
PPO
1.58
0.0728
0.3554
in-distribution
Merit Order
2.89
0.0096
0.1242
in-distribution
PPO-Lag
3.50
0.0000
0.0593
in-distribution
All-On
4.34
0.0000
0.0346
line-tightening
PPO
1.57
0.0547
0.4932
line-tightening
Merit Order
2.89
0.0096
0.3371
Appendix
Table 11: TSO cost and safety results. Operating cost is reported per 24-hour episode in £M; reserve and thermal columns are per-step violation rates.
Figure 12: TSO cost–safety summary. Left: total operating cost vs the worst safety-violation rate (the maximum of the reserve-shortfall and thermal-overload rates). Right: the same methods separated by violation channel, showing that the reserve channel is satisfied for PPO-Lagrangian and All-On while the thermal channel is not satisfied for any method.
Figure 13: TSO per-episode operating cost vs thermal-overload cost on the in-distribution split (4 algorithms × 5 seeds × 50 episodes = 1,000 points). Marker area is proportional to the per-episode thermal-violation rate; the ellipse around each method covers its 95% covariance region.
Figure 14: TSO train-split checkpoint monitor over 20 M environment steps. Panels show evaluation operating cost, reserve-shortfall incidence, and thermal-overload incidence. Curves are smoothed means over five seeds; shaded bands show one standard deviation.
Method
Total loss (MWh)
Voltage violations / step
Loss reduction
PPO
1.9355
0.00033
34.45%
Sauté PPO
1.9370
0.00292
34.35%
SAC
2.3178
0.05475
20.54%
PPO-Lag
2.4108
0.01733
15.84%
Droop
2.8593
0.07042
0.74%
No control
2.8909
0.20667
0.00%
Appendix
Table 12: DSO in-distribution physical metrics. Lower total network loss is better; the voltage column is the per-step count of buses outside the [0.94,1.06] p.u. band.
Figure 15: DSO in-distribution physical metrics. Panels report total network loss, voltage-violation count per step, and loss reduction relative to no control.
Figure 16: DSO learning curves over 3 M environment steps. Left: episode total reward across the four learned algorithms. Middle: per-step voltage-violation count. Right: PPO-Lagrangian Lagrange multiplier λ .
Figure 17: DSO per-episode in-distribution behavior distributions over 6 algorithms × 5 seeds × 50 episodes. Each panel reports a single behavior metric; boxes are overlaid by jittered individual episodes so distributional tails (especially the voltage-violation panel) remain visible.
Method
Active loss (MW)
Violation steps
Violation rate
IPPO-rs
0.1974
0.00
0.0000
IPPO
0.1974
0.00
0.0000
Voltage droop
0.2031
0.00
0.0000
IPPO-Lag
0.2051
0.0067
0.00014
No control
0.2052
0.00
0.0000
Appendix
Table 13: DERs in-distribution results. Lower active-power loss is better. Loss values are mean across 5 seeds ×30 episodes.
Method
Active loss (MW)
Violation steps
Violation rate
IPPO-rs
0.1974
4.91
0.1022
IPPO
0.1974
4.91
0.1024
Voltage droop
0.2031
5.97
0.1243
IPPO-Lag
0.2051
7.94
0.1654
No control
0.2052
8.73
0.1819
Appendix
Table 14: DERs voltage-tightening stress split (band tightened from [0.94,1.06] to [0.96,1.04] p.u.). Violation steps are per episode; rates use the 48 -step horizon as denominator. Active loss matches the in-distribution column because tightening changes only the constraint band, not dispatch decisions.
Figure 18: DERs schedules on a single in-distribution episode. Panels compare battery state of charge, battery reactive power commands, PV curtailment, PV reactive-power commands, flexible load actions, and active power loss across IPPO, IPPO-rs, IPPO-Lagrangian, voltage droop, and no control.
Method
Episode return
SLA rate
Spill cost
SAC
−2594.59
4.17×10−7
9.05×10−5
PPO
−2845.29
0
1.03×10−3
No control
−3110.68
1.11×10−6
0
Rule-based
−5226.24
8.33×10−7
4.58×10−4
Max renewable
−6010.68
1.39×10−7
6.30×10−3
Appendix
Table 15: DCMG in-distribution results. Higher episode return is better. SLA rate is the per-step rate of expired tasks; spill cost is the per-step over-generation penalty. The power-deficit and over-temperature rates are 0 for every method and are not shown.
Figure 19: Representative DCMG dispatch from the SAC policy on one in-distribution episode, displayed as 30-minute means for readability. Left: real GB solar capacity factor and the data center IT load. Middle: net load after PV, grid import, diesel output, and battery charge/discharge. Right: battery state of charge with SoC bounds, alongside the GB MID grid price.
Task
PowerZooJax
SBX/CUDA
SB3/CUDA
× vs SBX
GenCos
506,134
442
244
1,145 ×
DERs
414,506
203
233
2,040 ×
DSO
134,660
3,435
2,207
39 ×
TSO
53,724
1,649
411
33 ×
DCMG
25,132
1,167
1,298
22 ×
Appendix
Table 16: Throughput at Nenv=212 parallel environments, with the environment-side speedup isolated. The headline speedup against the slower CPU baseline is in main-text Table 2 .
Figure 20: Per-task throughput scaling. PowerZooJax JAX/GPU is plotted over nenv∈{16,32,64,128,256} for every task; SB3/CUDA is overlaid where its SubprocVecEnv backend completes (through nenv=64 for DCMG, through nenv=128 for DSO; DERs has no paired SB3 sweep). Endpoint labels mark the JAX/GPU and SB3/CUDA throughput at the right end of each curve; arrows mark the JAX-vs-SB3 ratio at the top of each matched range.
Task
Measured nenv range
Compile time
JAX throughput at nenv=256
DCMG
16–256
8.2 s
307.7k steps/s
DSO
16–256
2.9 s
388.4k steps/s
DERs
16–256
2.1 s
283.8k steps/s
Appendix
Table 17: First JAX compilation time in task-specific scaling sweeps. Values are means over seeds.
Figure 21: Evaluation return against elapsed wall-clock time for DSO and TSO across PowerZooJax, SBX/CUDA, and SB3/CUDA.
Power markets are a natural testbed for multi-agent reinforcement learning (MARL), where multiple self-interested participants repeatedly submit bids. A market-clearing mechanism then determines dispatch and prices subject to power grid constraints and market settlement rules. However, existing MARL environments typically focus on a single market setting, implement simplified clearing mechanisms, or rely on CPU-based optimization solvers that slow large-scale training and limit the systematic study of bidding strategies and market behavior. We introduce PowerMarketJax, a benchmark suite for MARL across five power markets: day-ahead wholesale, real-time balancing, ancillary services, peer-to-peer double auctions, and local flexibility. Each environment implements its own clearing, pricing, and settlement rules while providing a common framework for learning and evaluation. We find that learned bidding behavior depends strongly on the market design: independent learners can miss better strategies when gains require many agents to change together, when more profitable strategies lie beyond a region of lower profit, or when profits disappear as more agents adopt the same strategy. PowerMarketJax implements both market simulation and policy training in JAX, allowing the entire pipeline to run on the GPU with 1,024 X 1,200 parallelisms across both environments and market participants, achieving up to 33X speedup over CPU-based baselines. Our open-source benchmark is available at: https://github.com/powermarketjax/PowerMarketJax.
Zhanhua Pan, Xin Qin, Xiao Liu +3
Nanyang Technological University, Singapore. · Cornell University, USA. · University of Bristol, UK.
Benchmarks are crucial in the development of machine learning algorithms, with available environments significantly influencing reinforcement learning (RL) research. Traditionally, RL environments run on the CPU, which limits their scalability with typical academic compute. However, recent advancements in JAX have enabled the wider use of hardware acceleration, enabling massively parallel RL training pipelines and environments. While this has been successfully applied to single-agent RL, it has not yet been widely adopted for multi-agent scenarios. In this paper, we present JaxMARL, the first open-source, Python-based library that combines GPU-enabled efficiency with support for a large number of commonly used MARL environments and popular baseline algorithms. Our experiments show that, in terms of wall clock time, our JAX-based training pipeline is around 14 times faster than existing approaches, and up to 12500x when multiple training runs are vectorized. This enables efficient and thorough evaluations, potentially alleviating the evaluation crisis in the field. We also introduce and benchmark SMAX, a JAX-based approximate reimplementation of the popular StarCraft Multi-Agent Challenge, which removes the need to run the StarCraft II game engine. This not only enables GPU acceleration, but also provides a more flexible MARL environment, unlocking the potential for self-play, meta-learning, and other future applications in MARL. The code is available at https://github.com/flairox/jaxmarl.
Alexander Rutherford, Benjamin Ellis, Matteo Gallici +18
University of Oxford · Universitat Politècnica de Catalunya · University College London +2
Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way to assess tool-using AI agents in software settings. Electric power engineering has not yet had an analogous benchmark: language-model use is still dominated by retrieval and text question answering, while agents acting on power-system artifacts remain mostly academic prototypes. We introduce the Power Systems Agent Benchmark, an executable benchmark for power-engineering agents. An agent receives a structured task and returns a structured solution; a deterministic evaluator recomputes the engineering quantities, checks operational constraints, and returns a feasibility flag, a normalized score, and explicit violations. The benchmark contains 41 task families across eight areas of power engineering, from power flow and protection to stability, microgrids, reliability, power quality, and forecasting. Each task is grounded in a citable source, standard, or documented engineering formulation. To resist contamination, held-out cases are synthesized on demand by per-family generators from private seeds: the construction is inspectable, but the instances remain private. In a reference evaluation with three command-line agents, the strongest score near the compact tier's ceiling, a smaller open model trails, and public and held-out performance are broadly consistent; a separate public-split grid with OpenCode and Aider probes harness effects. The reference evaluation doubles as quality control: unanimous failures flag candidate task or evaluator defects, and it exposed a latent evaluator bug missed by self-consistency checks. The evaluators are compact deterministic surrogates, but the task contract allows their internals to be upgraded to simulator-backed checks without changing how tasks are posed or solved.