eess.SYNov 25, 2025

Reinforcement Learning to Initialize Newton-Raphson for AC Power Flow with Quantum Annealing-Based Environment Updates

Authors: Zeynab Kaseb, Matthias Moller, Lindsay Spoor, Jerry J. Guo, Yu Xiang, Peter Palensky, Pedro P. Vergara

Organizations: Electrical Sustainable Energy, Delft University of Technology, P.O. Box 5031, 2600 GA Delft, The Netherlands. · Applied Mathematics, Delft University of Technology, P.O. Box 5031, 2600 GA Delft, The Netherlands. · Leiden Institute of Advanced Computer Science, Leiden University, P.O. Box 9512, 2300 RA Leiden, The Netherlands. · Alliander N.V., P.O. Box 50, 6920 AB, Arnhem, The Netherlands. · Intelligent Systems, Delft University of Technology, P.O. Box 5031, 2600 GA Delft, The Netherlands. · Electrical Engineering, Eindhoven University of Technology, P.O. Box 513, 5600 MB, Eindhoven, The Netherlands.

Abstract

The Newton-Raphson (NR) method is widely used for solving power flow (PF) equations due to its quadratic convergence. However, its performance deteriorates under poor initialization or extreme operating scenarios, e.g., high levels of renewable energy penetration. We propose the use of reinforcement learning (RL) to optimize the initialization of NR, and introduce a quantum-enhanced RL environment update mechanism that addresses the combinatorially large action space at each RL timestep by formulating the voltage adjustment task as a Quadratic Unconstrained Binary Optimization (QUBO) problem, solved with an Ising machine. RL initialization is benchmarked against flat start and start from the DC (linearized) PF solution on a standard 4-bus system, Iwamoto's ill-conditioned 11-bus system, and the IEEE 118-bus system under normal and stressed loading and reactive power limits, with verified operational solutions. On all systems, a supervised initializer refined by RL requires fewer NR iterations than flat and DC starts and than the same initializer without RL, for all seeds. For example, on the 118-bus system under normal and stressed loading, it reached 2.04 and 2.86 NR iterations, compared with 3.02 and 5.13 from DC start and 2.61 and 3.09 without RL. In wall-clock time, this pays off only for an initializer integrated into the solver and reused for many solves on a fixed topology. On the 4-bus system, a quantum-enhanced RL agent with a quantum-inspired annealer moved challenging initial states that required 29 and 44 NR iterations to initializations that required three NR iterations within one RL timestep.

Figures & tables

Explore similar work

May 11, 2026cs.LG

Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models

Neural warm starts can sharply reduce the number of Newton-Raphson iterations required to solve the AC power flow problem, but existing supervised approaches generalize poorly on heavily loaded instances near voltage collapse. We prove a lower bound on the Newton-Raphson iteration count that depends on the direction of the warm start error rather than on its magnitude, and show as a corollary that the bound becomes vacuous as the smallest singular value of the power-flow Jacobian shrinks, identifying the failure mode of supervised regression near the saddle-node bifurcation. Motivated by this analysis, we introduce Newton's Lantern, a finetuning pipeline that combines group relative policy optimization with a learned reward model trained on perturbations of the base model's predictions, using the iteration count itself as the supervisory signal. Across IEEE 118-bus, GOC 500-bus, and GOC 2000-bus benchmarks, Newton's Lantern is the only method that converges on every test snapshot while attaining the smallest mean iteration count.
May 14, 2026cs.LG

QuantFPFlow: Quantum Amplitude Estimation for Fokker--Planck Policy Optimisation in Continuous Reinforcement Learning

We introduce \textbf{QuantFPFlow}, a reinforcement learning framework that integrates quantum amplitude estimation into the Fokker--Planck~(FP) formulation of stochastic policy optimisation. Classical continuous-space RL agents must estimate the FP partition function Z=∫e−V(x)/D dxZ = \int e^{-V(\mathbf{x})/D}\,d\mathbf{x} at cost \calO(1/ε2)\calO(1/\varepsilon^{2}); QuantFPFlow replaces this with a Grover-amplified amplitude estimator achieving \calO(1/ε)\calO(1/\varepsilon) -- a provable quadratic speedup. While the full quantum acceleration requires fault-tolerant hardware, the quantum-inspired classical simulation demonstrated here already exhibits the \calO(1/ε)\calO(1/\varepsilon) algorithmic structure. The estimated stationary distribution \rhostar\rhostar drives a theoretically grounded exploration bonus \Raug=\Renv+αlog⁡(1/\rhostar(s))\Raug = \Renv + α\log(1/\rhostar(s)). This bonus steers the agent toward globally optimal regions of multimodal reward landscapes while simultaneously constraining policy variance through FP diffusion matching. On a continuous-control task specifically designed to expose local-optima failure, QuantFPFlow achieves mean reward 1,295.7±423.21{,}295.7 \pm 423.2 versus 1,284.0±474.01{,}284.0 \pm 474.0 for Soft Actor-Critic~(SAC), while discovering the global optimum \textbf{10.4,% more frequently} (33.9,% vs.\ 30.7,%). Policy entropy remains near H(π)≈6.5H(π)\approx 6.5,nats throughout training, whereas SAC collapses to 1.51.5,nats, confirming that FP diffusion matching actively prevents premature convergence. Dimensionality experiments further show computational scaling of \calO(d0.35)\calO(d^{0.35}) for QuantFPFlow versus \calO(d0.76)\calO(d^{0.76}) for classical FP estimation.
Sep 26, 2025cs.LG

Physics-informed GNN for medium-high voltage AC power flow with edge-aware attention and line search correction operator

Physics-informed graph neural networks (PIGNNs) have emerged as fast AC power-flow solvers that can replace the classic NewtonRaphson (NR) solvers, especially when thousands of scenarios must be evaluated. However, current PIGNNs still need accuracy improvements at parity speed; in particular, the soft constraint on the physics loss is inoperative at inference, which can deter operational adoption. We address this with PIGNN-Attn-LS, combining an edge-aware attention mechanism that explicitly encodes line physics via per-edge biases to form a fully differentiable knownoperator layer inside the computation graph, with a backtracking line-search-based globalized correction operator that restores an operative decrease criterion at inference. Training and testing use a realistic High-/Medium-Voltage scenario generator, with NR used only to construct reference states. On held-out HV cases consisting of 4-32-bus grids, PIGNN-Attn-LS achieves a test RMSE of 0.00033 p.u. in voltage and 0.08 deg in angle, outperforming the PIGNN-MLP baseline by 99.5% and 87.1%, respectively. With streaming micro-batches, it delivers 2-5x faster batched inference than NR on 4-1024-bus grids.