cs.CEOct 7, 2026

Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach

Authors: Yasaman Cheraghi, Reidar B. Bratvold, Aojie Hong, Ressi B. Muhammad, Sergey Alyaev

Organizations: Department of Energy and Petroleum Engineering, University of Stavanger, Norway · Independent Researcher, Stavanger, Norway · NORCE Norwegian Research Centre, Bergen, Norway

Abstract

The global challenge of climate change has driven significant steps to reduce CO2 emissions, guided by international agreements like the Paris Agreement of 2015. Acting too slowly could result in future losses and reputational damage, while moving too quickly could jeopardize shareholder value due to the marginal profitability or potential losses due to technology immaturity of many renewable projects. To navigate this complex transition, energy companies must adopt Sequential Decision Making (SDM) strategies to maximize value creation from decision flexibility under uncertainties. To support this, we developed a custom simulation environment to model the dynamic energy landscape up to 2050. Building on this, we designed a multi-criteria SDM framework that explores various decision strategies related to different portfolios for allocating funds across three sectors: oil & gas, renewables, and CO2 reduction. It aims to maximize value during the transition while accounting for uncertainties in productions, energy prices, and costs. This framework has three objectives: maximizing profit, minimizing CO2 social costs, and enhancing competitive advantage in the renewable energy sector. This research evaluates the use of Reinforcement Learning (RL) to identify optimal investment policies within the defined SDM framework. The agent's sequential decisions shape a virtual dynamic environment by influencing key variables such as oil and gas production, renewable energy output, CO2 emissions, and revenues. Through repeated interaction, the RL algorithm explores the state space and learns an optimal policy under uncertainty. We benchmark the RL strategy against a set of manually defined baseline policies and find it consistently outperforms them in adaptability and long-term value creation.

Explore similar work

CardsList
  1. Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources

    Jun 23, 2026Haoyuan Deng, Yihong Zhou, Thomas Morstyn +1Reinforcement LearningRL Fine-Tuning

  2. Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration

    May 23, 2026Ninglin Ou, Mohammad A. Razzaque, Iftekher Islam Shovon +5RL ControlMulti-Objective Reinforcement Learning

  3. Accounting for Optimal Control in the Sizing of Isolated Hybrid Renewable Energy Systems Using Imitation Learning

    Jan 7, 2026Simon Halvdansson, Lucas Ferreira Bernardino, Brage Rugstad KnudsenPower Systems