Power markets are a natural testbed for multi-agent reinforcement learning (MARL), where multiple self-interested participants repeatedly submit bids. A market-clearing mechanism then determines dispatch and prices subject to power grid constraints and market settlement rules. However, existing MARL environments typically focus on a single market setting, implement simplified clearing mechanisms, or rely on CPU-based optimization solvers that slow large-scale training and limit the systematic study of bidding strategies and market behavior. We introduce PowerMarketJax, a benchmark suite for MARL across five power markets: day-ahead wholesale, real-time balancing, ancillary services, peer-to-peer double auctions, and local flexibility. Each environment implements its own clearing, pricing, and settlement rules while providing a common framework for learning and evaluation. We find that learned bidding behavior depends strongly on the market design: independent learners can miss better strategies when gains require many agents to change together, when more profitable strategies lie beyond a region of lower profit, or when profits disappear as more agents adopt the same strategy. PowerMarketJax implements both market simulation and policy training in JAX, allowing the entire pipeline to run on the GPU with 1,024 X 1,200 parallelisms across both environments and market participants, achieving up to 33X speedup over CPU-based baselines. Our open-source benchmark is available at: https://github.com/powermarketjax/PowerMarketJax.
Figures & tables
Figure 1: Five power markets in PowerMarketJax across different timescales and grid levels.
Figure 2: Execution paradigm of PowerMarketJax.
Figure 3: M1 day-ahead wholesale market (a) learning curves, (b) profit evaluation under joint and unilateral bidding strategies.
Figure 4: M2 real-time balancing market (a) bidding strategy, (b) learning curves, and M3 ancillary services market (c) profit under joint and unilateral bidding strategies.
Figure 5: P2P energy market (a) battery state of charge, (b) household gain, (c) local clearing price.
Figure 6: M5 local flexibility market (a) schematic return versus own offer price, (b) measured return in four system configurations, (c) procurement frequency versus the number of SAC-PS learners (seed mean with 95% bootstrap CI).
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Throughput of the peer-to-peer local energy market environment (1 200 households, 96 periods per episode) under three layouts: (a) against the number of parallel environments; (b) against the number of host CPU threads allotted, with 64 parallel environments. Each point is the mean of three runs with a bootstrap 95% confidence interval; each dotted line joins two points and the number beside it is their throughput ratio.
Figure 8: Throughput of the five market environments: (a) the three layouts with 64 parallel environments, the number above each host bar being the ratio of all on the GPU to that layout; (b) all on the GPU with 1 and with 128 parallel environments, the number above each pair being the ratio of the two bars. Each bar is the mean of three runs.
Figure 9: Throughput of the real-time balancing market environment in three implementations: PyPSA 1.3.0 and cvxpy 1.9.3 in their default usage, one environment at a time, and PowerMarketJax all on the GPU with 1 and with 64 parallel environments. Each bar is the mean of three runs; the label above each grey bar is how many times slower it is than PowerMarketJax with 64 parallel environments.
Market
Quantity
Data
Absolute difference
Relative difference
M1
LMP
36 days × 24 h, case29gb
2.2×10−8 $/MWh
1.4×10−10
Award
1.3×10−11 MW
1.7×10−15
Objective value
9.3×10−10 $
4.9×10−16
Settlement (profit)
2.8×10−9 $
6.3×10−16
M2
LMP
36 days × 48 half-hours, case29gb
9.1×10−8 $/MWh
9.1×10−12
Award
1.6×10−11 MW
2.1×10−15
Appendix
Table 1: Largest difference between the JAX implementation and the NumPy reference. The relative difference divides the absolute difference by the largest absolute reference value.
Market
Test system
Price
Periods
Price difference ($/MWh)
Relative objective difference
M1
case29gb
LMP
864
2.6×10−8
3.1×10−16
M1
case73rts
LMP
864
4.1×10−8
6.2×10−8
M1
case813nem
LMP
864
3.5×10−7
1.2×10−7
M2
case29gb
LMP
1 728
1.1×10−7
1.8×10−12
M3
case29gb
LMP
1 728
7.3×10−6
3.1×10−11
M3
case29gb
Reserve price
1 728
2.1×10−7
3.1×10−11
Appendix
Table 2: Largest difference between the JAX solver and HiGHS.
Figure 10: Production cost of the three-stage clearing above the exact optimum on the evaluation days of the three M1 test systems, under truthful offers. Each point is one evaluation day, in date order; the dashed line is the median.
Figure 11: One training iteration of IPPO, compiled as a single function. Rollout: N environments advance in parallel for T steps; in each step the n agents submit offers, the market clears, the awards are settled and the next observation is constructed, and the policy maps the observation to the next action inside the same function. Update: the policy is updated on the N×T experiences, with a second scan over epochs and minibatches (Table 3 ). The values of N and T for each market are in Table 4 .
Hyperparameter
IPPO
SAC
Optimiser
Adam
Adam
Learning rate
3×10−4 , constant
actor 3×10−4 , Q-networks and temperature 10−3
Discount factor γ
0.99
0.99
Hidden layers
two layers of 64 units, tanh
two layers of 256 units, ReLU
GAE parameter λ
0.95
—
Clipping range
0.2
—
Appendix
Table 3: Hyperparameters of the two IPPO and the two SAC learners, shared by the five markets.
Market
Seeds per learner
Iterations
Parallel environments and steps per iteration
Environment steps per iteration
M1 day-ahead wholesale
3
400
64 environments, 4 steps each
256
M2 real-time balancing
3
200
64 environments, 48 steps each
3 072
M3 ancillary services
3
200
64 environments, 48 steps each
3 072
M4 peer-to-peer local energy
3
400
64 environments, 96 steps each
6 144
M5 local flexibility
3 to 10
early stopping a , at most 300
64 environments, 24 steps each
1 536
Appendix
Table 4: Training scale.
Test system
Grid
Demand data
Time span
Training / evaluation days
Markup cap αˉ
case29gb
A reduced 29-bus model of the GB transmission network a (network diagram b ), 66 units
GB transmission system demand (NESO) c , day-ahead forecast from Elexon d
2023-07-10 to 2024-07-08
329 / 36
2
case73rts
The RTS-GMLC test system e , 73 units
Load, wind, solar and hydro of RTS-GMLC f
2020-01-01 to 2020-12-31
330 / 36
1.4
case813nem
An open grid model of the Australian National Electricity Market g , 151 units
Operational demand and day-ahead forecast of the four mainland regions from the Australian Energy Market Operator (AEMO) h
2025-02-01 to 2026-01-31
329 / 36
2
Appendix
Table 5: Data of the three test systems.
Figure 12: The test system case29gb. Left: the 66 units at their buses, coloured by fuel and sized by maximum output; corridor width scales with capacity. Middle: annual mean bus load and the corridors congested under truthful offers on the three example days. Right: the supply curve ordered by segment cost, against the range of hourly system demand over one year. Generation sits mostly in the north and load mostly in the south, so the congested corridors are central and southern. All 14 nuclear units, 4 359 MW in total, are among the 16 cheapest units, which sum to 19.2 GW; hourly demand ranges from 17.2 to 47.4 GW and exceeds 19.2 GW in almost every hour. The nuclear units therefore sit near the left of the supply curve with demand almost always to their right: a unilateral markup does not move the price, while a joint markup lifts the whole curve and the price with it.
Figure 13: The test system case73rts, drawn as Figure 12 : the three-area RTS-GMLC system with 73 units. Left: units by fuel and maximum output. Middle: annual mean bus load and the lines congested under truthful offers on the three example days. Right: the supply curve ordered by segment cost against the range of hourly net demand over one year. Gas units make up most of the supply curve; no line is congested on the low-load day and seven are on the high-load day.
Figure 14: The test system case813nem, drawn as Figure 12 : an open grid model of the mainland Australian National Electricity Market with 151 units. Only 7 of its 1 278 lines carry a published rating. Left: units by fuel and maximum output. Middle: annual mean bus load and the lines congested under truthful offers on the three example days. Right: the supply curve ordered by segment cost against the range of hourly demand over one year. Hydro and then coal fill the left of the supply curve; three lines are congested on all three days and a fourth on the mid- and high-load days.
Figure 15: Demand on the three test systems, and the renewable output netted out on case73rts. Left: daily means over the one-year window, with the 36 evaluation days marked. Right: hourly values in the week around the mid-load example day. The thick black line is the demand the market clears; the dashed line is the day-ahead forecast the agents observe. On case73rts the market clears net demand, total load minus the renewable output shown, raised to a floor of 2 500 MW; it sits at that floor in more than half of the hours of the year. Demand on case813nem has a floor of 11 500 MW, reached in a few hours.
Figure 16: Training return on the three test systems. Top: the four learners on case29gb. Bottom: IPPO-PS and IPPO-NoPS on case73rts and case813nem. Each line is the mean of three seeds, each seed first smoothed by an 11-iteration centred moving mean, and the shaded band is the bootstrap 95% confidence interval. Return is the mean daily profit per generator over all sampled environment steps of an iteration. On case29gb IPPO-PS is the only learner that rises and then falls; on the other two systems the return of both learners falls from its starting level.
Figure 17: Daily system quantities on the 36 evaluation days of each test system, truthful offers (black) against the learned offers of IPPO-PS (orange) and IPPO-NoPS (blue); each learned line is the mean of three seeds, with a bootstrap 95% confidence interval. Rows: daily demand, the load-weighted mean LMP, production cost, and the total profit of all units.
Figure 18: Supply curves at the hour of highest demand on each example day of case29gb, the evaluation days at the 10th, 50th and 90th percentiles of daily demand (2024-06-23, 2024-03-16 and 2024-01-23): committed units sorted by offer under truthful (grey) and learned (orange) offers, with the steps of the 14 nuclear units drawn thicker and the demand in that hour as a dashed line; the shaded bands are the range of LMP over the 29 buses under each set of offers. The learned offers are the mean over three IPPO-PS seeds at the end of training, with a bootstrap 95% confidence interval for each step; the learned LMP band is the range over buses of the seed-mean LMP. The nuclear steps lie far to the left of demand under both sets of offers, and the learned offers raise the whole curve.
Figure 19: Markup of each unit under the four learners on case29gb; rows are learners, each the mean over three seeds, columns the 66 units ordered by segment cost. Colour is the markup of that unit, averaged over the 36 evaluation days and the three seeds, on the full action range from 1.00 (truthful) to 2.00 (cap). Black triangles above the top row mark the 14 nuclear units.
Figure 20: Demand on case29gb. Left: the week around the mid-load day (2024-03-16): the half-hourly demand the real-time market clears (teal), the hourly day-ahead schedule summed over buses, which equals the hourly mean of that demand (purple steps), and the day-ahead forecast the units observe (dashed yellow). Right: the deviation of the demand in each half-hour from the mean of its hour, over all 365 days (grey) and the 36 evaluation days (teal). On the evaluation days the absolute deviation has a median of 214 MW, a 90th percentile of 638 MW and a maximum of 1 478 MW, against a mean demand of 27 730 MW. The day-ahead schedule follows the half-hourly demand to within a few hundred MW.
Figure 21: Training return on two test systems in the real-time market. Top: the four learners on case29gb. Bottom: IPPO-PS and IPPO-NoPS on case73rts. Each line is the mean over seeds, smoothed by an 11-iteration moving mean, and the shaded band is the bootstrap 95% confidence interval of the mean. Return is the profit of a unit per day, averaged over units and over the environment steps of an iteration. On case29gb the return of IPPO-PS falls the most, and SAC-NoPS is the only learner whose return rises. On case73rts the return of both learners rises slowly.
Figure 22: Learned offers against truthful offers on case29gb over the 36 evaluation days. Left: change in total profit of all 66 units compared to truthful offers (in million dollars per day); right: change in system cost compared to truthful offers (in percent). Each circle indicates one learner after training, placed at the mean markup of the 30 committed units. The bars are the 95% confidence interval of the mean over seeds; the dotted line is truthful offers. IPPO-PS returns to truthful offers; the other three learners end with higher total profit and lower system cost.
Figure 23: Each of the 30 committed units raises its offer alone while the others offer truthfully. Left: own profit change against markup (one line per unit, median in black). Right: own profit change against the profit change of the other units (with a least-squares fit, near-zero points are omitted). Means over the 36 evaluation days, in million dollars per day.
Figure 24: All 66 units at the same markup, from 1.00 to 2.00. Left: change in total profit (million dollars per day). Right: change in system cost (%). Both are relative to truthful offers (dotted line), averaged over the 36 evaluation days. A common markup leaves both nearly unchanged.
Test system
Markup cap αˉ
Value of lost reserve VOLR (£/MWh)
case29gb
2
250
case73rts
2
136
case813nem
2
147
Appendix
Table 6: Markup cap and value of lost reserve of the ancillary services market on the three test systems.
Figure 25: Training return on the three test systems. Top: the four learners on case29gb. Bottom: IPPO-PS and IPPO-NoPS on case73rts and case813nem. Return is the mean daily profit per generator in each iteration. Lines are the mean of three seeds, each smoothed over 11 iterations, and bands are bootstrap 95% confidence intervals. On case29gb, only the return of IPPO-PS falls, and on the other two systems, both learners rise.
Figure 26: Evaluation on the 36 evaluation days. (a) case29gb: change in total profit of all 66 units relative to truthful offers (million dollars per day), split into reserve payments and energy. (b) case29gb: change in production cost. Filled markers are trained policies; hollow markers are untrained networks. (c) case73rts, relative to the untrained network (truthful offers not evaluated), and (d) case813nem, relative to truthful offers: change in total profit against change in production cost for IPPO-PS (orange) and IPPO-NoPS (blue). The black square marks the reference. Markers are seed means with bootstrap 95% confidence intervals. On case29gb, all learners earn more than truthful offers at nearly unchanged cost, but only IPPO-PS earns less than its untrained network; on the other two systems, the differences are small.
Figure 27: Raising the reserve offer on case29gb from 0 to £150/MWh on both energy and reserve products. Teal: own profit change of each of the 29 committed units when it raises its offer alone, with the others truthful. Black: the median. Orange: total profit change of all 66 units when all raise their offers together. Values in million dollars per day. Raising alone gains almost nothing, while raising together gains £9.05 million per day at £150/MWh.
Figure 28: The P2P market of the test community. Each of the 1 200 households has PV and a battery. In every quarter-hour, sellers ask and buyers bid in a community double auction, which sets the clearing price λtloc between the export price πexp and the retail tariff πret ; trades in the community are settled at λtloc . Energy the auction leaves unmatched is settled with the grid.
Data source
Households
Period covered
Training and held-out
Licence
Quarter-hour injection and offtake from Fluvius smart meters a
1 200, all with PV
2024-04-01 to 2024-10-26, 209 days
3 consecutive days held out of every 15; 14 702 training and 2 702 held-out episode starts; evaluation on 1 024 held-out episodes, of which the figures below use 64
Fluvius open data licence
Appendix
Table 7: Data of the test community.
Figure 29: The test community of 1 200 households: community injection and offtake (left) and the share of households injecting more than they take off (right), by hour of day over the 64 held-out episodes.
Figure 30: Training return of the four learners up to iteration 400; the policy at iteration 400 is the one evaluated throughout this section. Shown is the mean over three seeds of the 11-iteration moving mean of the return on 64 sampled training episodes, with the shading the bootstrap 95% confidence interval. Both IPPO learners rise fastest in the first 50 iterations and both SAC learners reach their level within a few dozen; after iteration 50 all four change little.
Figure 31: Results on the 64 held-out episodes ordered by start time. Top: mean clearing price per episode. Bottom: profit per household per episode. Black is truthful offers; each learner is the mean of three seeds with a bootstrap 95% confidence interval. IPPO-NoPS lowers the price and raises the profit on every episode, while IPPO-PS and SAC-NoPS gain on 54 and 57 episodes. The seeds of SAC-PS disagree, and two of them raise the price.
Figure 32: Learned actions of the four learners on the 64 held-out episodes (mean of three seeds, bootstrap 95% confidence interval). Top: state of charge averaged over households by hour of day (last 12 hours of each episode only). The band marks the PV surplus hours, those in which the community injects more than it draws on most days. Middle and bottom: asks and bids, scaled from the export price (0) to the retail tariff (1).
Figure 33: On the 64 held-out episodes, k households follow a fixed arbitrage schedule (charge 11:00–15:00, discharge 18:00–22:00) and the rest keep their batteries idle. Left: gain per household, for those arbitraging and those idle, relative to all batteries idle. Right: mean clearing price in the charging and discharging windows. The horizontal axis of both panels is k . Arbitrage pays while few households take part and turns into a loss at 210 to 240 households, before the two window prices cross.
Figure 34: Feeder 459_0 in Switzerland. Bottom right: its location in the Lake Geneva region. Left: the feeder, with a small red box marking the congested line. Top right: an aerial image of the congested line. Lines behind the congested line are blue, and only batteries on these lines can relieve it.
Figure 35: Demand and supply of flexibility. Left: the requirement in each hour of the 36 test days, with all batteries idle. Right: the power of each of the 24 batteries in the 2040 placement, with the 12 downstream of the congested line first.
Data
Source
Licence
Training days
Validation days
Test days
Swiss distribution feeder 459_0
SwissDN a , Zapparoli et al. (2025) b
CC BY 4.0
293 days
36 days (the 5th, 15th and 25th of each month)
36 days (the 1st, 11th and 21st of each month)
Energy price (category C2, 2026)
ElCom c
opendata.swiss Open use
—
—
—
Appendix
Table 8: Overview of Swiss test system data.
Figure 36: Training return of the four learners against training iteration; lines are the geometric mean over the seeds of each configuration and shaded bands the bootstrap 95% confidence interval, each seed held at its last value after it stops. All four learners reach a plateau within about 100 iterations.
Figure 37: Source of the flexibility requirement over the 36 test days of the four configurations. Bars show truthful offers, the untrained shared and per-agent networks pooled in one bar, and the four learners at validation-selected checkpoints (seed mean and bootstrap 95% confidence interval). Two SAC-PS return bars exceed the axis and are cut at its top, with their means shown above. Rows show the share of periods with procurement; daily procurement split into what truthful offers also buy and the excess; and daily return per aggregator. Learners raise procurement above the truthful-offer level, and their returns rise with it.
Figure 38: Unilateral offer increase. Left: offer above the floor against price action (log scale). Near the floor, the gap shrinks by about a factor of e for each unit decrease in the action. Cleared prices for all IPPO seeds are within 1% of the floor. Right: aggregator 7’s hourly award on the example day at offers of the floor, 150.25 and 9 900 CHF/MWh, with all other aggregators at the floor.
Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: https://github.com/powerzoojax/PowerZooJax.
Zhanhua Pan, Xiao Liu, Zhilong Cao +2
Nanyang Technological University, Singapore. · Cornell University, USA. · University of Bristol, UK.
Benchmarks are crucial in the development of machine learning algorithms, with available environments significantly influencing reinforcement learning (RL) research. Traditionally, RL environments run on the CPU, which limits their scalability with typical academic compute. However, recent advancements in JAX have enabled the wider use of hardware acceleration, enabling massively parallel RL training pipelines and environments. While this has been successfully applied to single-agent RL, it has not yet been widely adopted for multi-agent scenarios. In this paper, we present JaxMARL, the first open-source, Python-based library that combines GPU-enabled efficiency with support for a large number of commonly used MARL environments and popular baseline algorithms. Our experiments show that, in terms of wall clock time, our JAX-based training pipeline is around 14 times faster than existing approaches, and up to 12500x when multiple training runs are vectorized. This enables efficient and thorough evaluations, potentially alleviating the evaluation crisis in the field. We also introduce and benchmark SMAX, a JAX-based approximate reimplementation of the popular StarCraft Multi-Agent Challenge, which removes the need to run the StarCraft II game engine. This not only enables GPU acceleration, but also provides a more flexible MARL environment, unlocking the potential for self-play, meta-learning, and other future applications in MARL. The code is available at https://github.com/flairox/jaxmarl.
Alexander Rutherford, Benjamin Ellis, Matteo Gallici +18
University of Oxford · Universitat Politècnica de Catalunya · University College London +2
The increasing penetration of renewable energy has introduced substantial volatility into wholesale electricity markets, complicating the optimal bidding strategies for power producers. Traditional Reinforcement Learning (RL) approaches often struggle to balance profit maximization with risk management, frequently overfitting to specific market conditions or failing to account for the stochastic spread between Day-Ahead (DA) and Real-Time (RT) settlements. To address these challenges, this paper makes two primary contributions. First, we introduce and open-source a high-fidelity gymnasium environment for two-settlement electricity market bidding. Grounded in extensive empirical data from the PJM Interconnection, the environment explicitly models the interplay between DA commitments and RT deviations, providing a standardized testbed for general and risk-sensitive agents. Second, we propose MARS-DA (Multi-Agent Regime-Switching for Day-Ahead markets), a novel hierarchical framework that orchestrates distinct sub-policies for risk management and profit seeking. MARS-DA utilizes a top-level Meta-Controller to dynamically blend the actions of two specialized base agents: a "Safe Agent" that optimizes for reliable DA allocation and a "Speculator Agent" that targets volatile RT arbitrage opportunities. Extensive experiments demonstrate that MARS-DA achieves superior risk-adjusted returns compared to state-of-the-art baselines while maintaining robust regime alignment during periods of extreme market volatility.
Jiayi Chen, Xuan Zhang, Guiling Wang
Department of Computer Science, New Jersey Institute of Technology