Climate Surrogates for Scalable Multi-Agent Reinforcement Learning: A Case Study with CICERO-SCM
Organizations: Technical University of Denmark Kongens Lyngby, Denmark
Abstract
Climate policy analysis requires models that capture multi-gas climate effects, but such models are too slow to embed in reinforcement learning loops at scale. In collaboration with the European Environment Agency, we develop a multi-agent reinforcement learning (MARL) framework that integrates a higher-fidelity climate surrogate as the environment transition, enabling regional agents to learn policies under multi-gas dynamics. We train a recurrent surrogate on multi-gas emission pathways to emulate CICERO-SCM. The surrogate achieves near-simulator accuracy (global-mean temperature RMSE ) with faster one-step inference and yields end-to-end MARL training speed-up. We show policy agreement with the simulator in tractable settings and propose a replay- and rank-consistency test (Kendall's ) for assessing policy fidelity when simulator-in-the-loop training is infeasible. This enables large-scale multi-agent policy experiments while retaining high-fidelity multi-gas climate response.
Figures & tables
| Test data performance & inference speed | Policy-induced performance & MARL speed | ||||||
|---|---|---|---|---|---|---|---|
| Climate engine | Test data (RMSE, ) | Mean inference [s] (CPU/GPU) | Speed-up (CPU/GPU) | Scenario (i) (RMSE, rank- ) | Scenario (ii) (RMSE, rank- ) | Mean env-step [s] | Speed-up |
| CICERO–SCM | – | / – | – | – | – | – | |
| LSTM | ( , 0.99) | / | / | ( , 0.996) | ( , 0.990) | ||
| GRU | ( , 0.99) | / | / | ( , 0.996) | ( , 0.997) | ||
| TCN | ( , 0.99) | / | / | ( , 0.994) | ( , 0.982) | ||
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Species | Class | Forcing Sign | Model Treatment |
|---|---|---|---|
| CO 2 (FF) | Long-lived GHG | Warming | Carbon-cycle |
| CO 2 (AFOLU) | Long-lived GHG | Warming | Carbon-cycle |
| CH 4 | Short-lived GHG | Warming | Simplified decay (multi- ) |
| N 2 O | Long-lived GHG | Warming | Fixed lifetime decay |
| SO 2 | Aerosol precursor | Cooling | Linear forcing proxy |
| CFC-11 | Long-lived GHG | Warming | Fixed lifetime decay |
| Group | Parameter | Value | Description |
| Climate response parameters | rlamdo | 15.0836 | Air–sea heat exchange parameter [W m -2 K -1 ] |
| akapa | 0.6568 | Vertical heat diffusivity [cm 2 s -1 ] | |
| cpi | 0.2077 | Polar amplification factor | |
| W | 2.2059 | Upwelling velocity [m yr -1 ] | |
| beto | 6.8982 | Ocean heat exchange coefficient [W m -2 K -1 ] | |
| lambda | 0.6063 | Climate sensitivity parameter [K W -1 m 2 ] |
| Item | Specification |
|---|---|
| Agents ( ) | agents with shares |
| Climate engines | CICERO-SCM and surrogates |
| Controlled gases | CO2_FF, CO2_AFOLU, CH4, N2O, SO2 |
| Levers and levels | |
| Energy | Levels: |
| Methane | Levels: |
| Item | Specification |
|---|---|
| Agents ( ) | agents with shares |
| Climate engines | Surrogates only |
| Controlled gases | CO2_FF, CO2_AFOLU, CH4, N2O, SO2 |
| Levers and levels | |
| Energy | Levels: |
| Methane | Levels: |
| Topic | Scenario | Description of figure |
|---|---|---|
| Training time comparison | Homogeneous | Fig. A.3 — Wall-clock training time for(surrogate vs. simulator) |
| Reward convergence | Homogeneous | Fig. A.4 — Reward convergence per agent (surrogate vs. simulator) |
| Reward convergence | Heterogeneous | Fig. A.5 — Reward convergence per agent (surrogate only) |
| Mean lever policies | Homogeneous | Fig. A.6 — Mean lever convergence (surrogate vs. simulator) |
| Mean lever policies | Heterogeneous | Fig. A.7 — Mean lever efforts convergence (surrogates only) |
| Per-agent levers (GRU) | Homogeneous | Fig. A.8 — Per-agent lever convergence (SCM vs. GRU) |