May 14, 2026 · eess.SYJ/K move · Enter open · S save
Maximilian Bloor, Martha White, Ehecatl Antonio del Rio Chanona, Calvin Tsay
Sargent Centre for Process Systems Engineering, Imperial College London, London, SW7 2AZ, UK · Department of Computer Science, University of Alberta, Edmonton, AB, Canada · Department of Computing, Imperial College London, London, SW7 2AZ, UK
Electrified chemical processes are incentivized by exposure to time-varying electricity markets to operate flexibly, but participating in demand response schemes can require satisfying terminal constraints over long horizons. Specifically, terminal constraints may be required when computing optimal schedules in order to preserve dynamic stability. Model-based optimization methods are computationally costly, and data-driven scheduling via reinforcement learning (RL) faces severe credit-assignment challenges. We integrate Goal-Space Planning (GSP) with Deep Deterministic Policy Gradient (DDPG), using learned temporally abstract models over discrete subgoals to propagate value across extended horizons. Using a simulated air separation benchmark, we demonstrate the proposed approach improves sample efficiency over standard DDPG while satisfying terminal storage constraints, mitigating myopic control behavior.