Recent neural power-flow solvers, including emerging foundation models, achieve accurate voltage predictions, yet such accuracy does not necessarily imply physically consistent solutions. Even small complex voltage errors can yield large AC power-balance residuals. We study this accuracy-consistency gap across PIGNN-GC, GridSFM, gridfm-graphkit, and LUMINA on realistic 2224-bus Great Britain network (GBnetwork) scenarios, with cross-grid evaluation of GridSFM over 31 systems. Using a singular value decomposition (SVD) basis fitted to training AC power-flow solutions, we find that neural prediction errors contain substantial components outside the dominant solution subspace. To address this mismatch, calibrated solution-subspace projection (CSP) suppresses off-subspace prediction components after train-only bias calibration, reducing Mean PB by 67.0%, 37.8%, 40.5%, and 68.9% for PIGNN-GC, GridSFM, gridfm-graphkit, and LUMINA, respectively, relative to calibrated predictions, while improving voltage-magnitude accuracy in all four models. These results identify output-error geometry as an important factor in physics-consistent neural AC power flow. Code: https://github.com/Kimchangheon/neural-acpf-error-geometry
Figures & tables
Figure 1: A low-dimensional subspace is derived from the training solutions. CSP first removes train-estimated bias and projects the corrected prediction onto this subspace, suppressing off-subspace errors and improving voltage accuracy and AC power-balance consistency.
Model
Output
V RMSE
θ RMSE
Rw2↑
aw
Mean PB
Max PB
Inference
↓
( ∘ ) ↓
↓
↓
(ms/scen.)
Per-bus mean
Baseline
.02341
6.847
.000
.000
.12084
5.88
–
PIGNN-GC
Raw
.02850±.00232
1.108±.175
.671±.050
.546±.072
.31592±.03240
195.89±52.54
10.034±.023
C
.01393±.00105
.777±.161
.671±.050
.546±.072
.16661±.02352
190.00±37.53
10.028±.023
P16
.01624±.00254
1.091±.166
.829±.045
.541±.071
.07594±.00422
5.07±.85
10.036±.023
CSP 16
.01222±.00134
.835±.141
.829±.045
.541±.071
.05497±.00257
3.93±.07
10.034±.032
Table 1: GBnetwork results at k=16 . All models report mean ± standard deviation over three model seeds. C denotes calibration, P16 projection, and CSP 16 their composition. Inference time is GPU-synchronized forward execution plus the corresponding output transform, reported per scenario.
Figure 2: Fixed-norm directional intervention on GBnetwork, with Relative Mean PB normalized at α=1 .
Basis
Ref. var.
Ref.
Pred.
retained (%)
PB
PB
SVD ( 16/16 )
98.789/99.683
.0415
.0569
Laplacian matched ( 1750/2122 )
98.819/99.688
26.15
26.22
Spatial matched ( β=.357/.208 )
98.801/99.684
6.20
6.22
Table 2: Matched basis comparison on GBnetwork for PIGNN-GC. Parentheses denote V/θ ranks or smoothing strengths.
Figure 3: Grid-by-grid physical consistency of GridSFM before and after CSP 16 across 31 systems ordered by bus count. Mean PB is shown on a logarithmic scale.
Model
Start
V RMSE
θ RMSE
Mean PB
(10−3)
( ∘ )
(10−3)
PIGNN-GC
C
1.80±.12
.119±.013
2.84±.46
CSP 16
1.17±.10
.073±.009
1.48±.13
GridSFM
C
11.86±6.77
.681±.380
2.50±1.54
CSP 16
.97±.20
.059±.013
1.22±.29
gridfm-
C
.943±.232
.0843±.0236
.941±.325
Table 3: One Newton update from calibrated or CSP 16 outputs on GBnetwork. V RMSE and Mean PB are reported in 10−3 . Values are means ± standard deviations over three model seeds.
Condition
V RMSE
Rw2↑
aw
Mean
Med.
p95
p99
Max
Control
.01281
.8476
.4866
.0581
.0131
.2665
.5462
3.58
N –1 (T)
.01314
.8391
.4750
.0603
.0131
.2713
.5632
33.16
N –2 (T)
.01288
.8458
.4870
.0617
.0132
.2713
.5721
27.45
GBcorr (T)
.01966
.4819
.4025
.1003
.0231
.4512
.9746
7.00
GBcorr (R)
.01620
.5150
.4247
.0680
.0174
.3031
.6415
4.99
Table 4: PIGNN-GC with fixed-rank CSP 16 under topology and operating-distribution shifts. T transfers the control calibration and basis, whereas R refits both on condition-matched training data.
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany · Institute of Electrical Energy Systems, Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany