We present the first complete machine learning method for accelerating plane-wave density functional theory (DFT) in materials under the projector augmented wave (PAW) formalism. We formalize seven criteria that a Complete Neural Electronic Initializer must satisfy for practical end-to-end PAW DFT acceleration. Applying these to prior work reveals two structure-dependent components, augmentation occupancies and spin initialization, whose absence prevents existing acceleration methods from providing complete reference-free initialization. We show that omitting these components can eliminate or reverse the acceleration obtained via models that only predict the smooth valence density. We satisfy the missing requirements by introducing AugNet, a general equivariant model for PAW augmentation occupancies, and the first general spin density model for materials, which predicts the smooth spin-difference density and spin-difference PAW augmentation occupancies using predicted magnetic moments to constrain the global magnetic state. Combined with existing valence density models, our full method satisfies all seven criteria and forms a fully reference-free electronic initializer for materials DFT, requiring no electronic quantities from a converged target calculation. We show that perfect initialization could cut PAW DFT wall time by 40-52%, and our method recovers up to 62% of this saving, reducing end-to-end DFT wall time by up to ~25% on unseen structures while preserving converged energies.
Figures & tables
Figure 1 : Workflow for end-to-end PAW DFT acceleration, illustrated using the models employed in this work. For a given atomic structure and PAW setup, the fixed PAW datasets define the frozen core and basis information, while a Complete Neural Electronic Initializer provides the three structure-dependent components defined in Table 1 : the smooth valence density ρ~+ , PAW augmentation occupancies da+ , and spin/magnetic initialization. ELECTRAFI predicts ρ~+ , AugNet predicts da+ , and the corresponding spin-dependent components ρ~− and da− are predicted by spin-ELECTRAFI and spin-AugNet, with magnetic moment predictions constraining the global magnetic state. Together, these components enable complete reference-free initialization of PAW DFT.
Reference-free initialization components
Demonstrated evaluation
Method
Valence density C1
PAW augmentation C2
Spin / magnetic initialization C3
Density accuracy C4
SCF acceleration C5
Component ablations C6
Reference-free acceleration C7
GPWNO ( Kim and Ahn, 2024 )
✓
×
×
✓
×
×
×
InfGCN ( Cheng and Peng, 2024 )
✓
×
×
✓
×
×
×
SCDP ( Fu et al., 2024 )
✓
×
×
✓
×
×
×
BOA ( Klockow et al., 2026 )
✓
×
×
✓
×
×
×
EdenGNN ( Li et al., 2025 )
✓
✓
×
✓
×
×
×
Table 1 : Criteria for complete neural electronic initialization for accelerating materials DFT. C1-C3 are structure-dependent quantities. C4-C7 refer to the required evaluation for demonstrating complete initialization. ✓ denotes a capability in the cited work, × denotes that it was not demonstrated. † denotes system-specific models: CJM ( Focassio et al., 2024 ) is specific to MoS 2 , while de Blasio et al. (2023) and EAC-Net ( Qin et al., 2026 ) predict spin densities for Na 3 V 2 (PO 4 ) 3 and Fe, respectively, without demonstrating SCF acceleration.
Spin channel
Aug. only
Valence only
Default
Valence + aug.
mm only
Smooth grids
Val. + aug. + mm
Oracle
Valence density ρ~+
×
×
✓
×
✓
×
✓
✓
✓
Spin density ρ~−
✓
×
×
×
×
mm ✓
✓
mm ✓
✓
Valence aug. d+
×
✓
×
×
✓
×
×
✓
✓
Spin aug. d−
✓
✓
×
×
×
mm ✓
×
mm ✓
✓
Non-mag. [%]
−9.7
−14.3
+0.5
0
+13.5
+12.0
+10.2
+47.6
+49.0
Mag. [%]
−70.7
−65.0
−29.4
0
−11.0
−3.1
−4.2
+24.6
+55.4
Table 2: Component ablation for PAW DFT initialization. The matrix specifies the electronic components in each experiment, with SCF step savings relative to the default SAD initialization reported separately for non-magnetic and magnetic Materials Project structures. ✓ denotes converged (Oracle) initialization, × denotes SAD/default initialization. mm denotes spin initialization using atomic magnetic moments.
Magnetic initialization
Subset
MAE
Oracle valence + augmentation
ELECTRAFI + AugNet
ChargE3Net + AugNet
Oracle (ρ~−,d−)
Non-mag.
–
49.0%
23.3%
29.1%
Magnetic
–
55.4%
29.3%
31.9%
Oracle moments
Non-mag.
0.028
47.6%
22.6%
29.0%
Magnetic
4.658
24.7%
6.2%
9.5%
Models
CHGNet moments
Non-mag.
0.347
38.7%
19.4%
25.9%
Table 3 : Spin-difference density accuracy and SCF step reduction for magnetic initializations using Oracle and ML valence and augmentation on the Materials Project test set.
Model:
CJM
AugNet
AugNet
spin-AugNet
spin-AugNet
Data:
MoS2
MP
GNoME
MP
GNoME
MAE ↓
0.0130
0.0041
0.0062
0.0025
0.0020
RMSE ↓
0.0459
0.0119
0.0262
0.0098
0.0129
Table 4 : PAW augmentation occupancy prediction errors. CJM is evaluated on its system-specific MoS2 structure. AugNet is evaluated on the MP and GNoME test data.
Dataset
Metric
Default
Oracle
CNEI-EFI
CNEI-C3Net
MP
NMAE [%] ↓
–
–
0.58
0.54
SCF steps ↓
22.05
10.41
19.05
18.36
SCF steps saved [%] ↑
–
52.78
13.62
16.76
Total time [s] ↓
623.84
302.04
530.04
585.26
Total time saved [%] ↑
–
51.58
15.04
6.18
OAC [%] ↑
0.0
100.0
29.15
11.99
Table 5 : End-to-end performance of our Complete Neural Electronic Initializer (CNEI), using either ELECTRAFI (EFI) or ChargE3Net (C3Net) as the valence backbone. Total time includes ML initialization and DFT execution. SCF step savings are reported relative to Default. Oracle acceleration captured (OAC) is calculated as defined in Equation 10 .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Initialization components
vs. Default [%]
vs. Oracle [%]
Smooth density
PAW augmentation
Initialization
ρ~+
ρ~−
d+
d−
Non-mag.
Mag.
Non-mag.
Mag.
Baselines
Default
×
×
×
×
–
–
–
–
Oracle
✓
✓
✓
✓
+49.0
+55.4
–
–
Electronic component ablations
Appendix
Table 6: Paired SCF-step savings (%) on the Materials Project test set, separated into non-magnetic and magnetic structures. The initialization components are the smooth spin-summed valence density ρ~+ , smooth spin-difference density ρ~− , spin-summed PAW augmentation occupancies d+ , and spin-difference PAW augmentation occupancies d− , where d± denotes the collection of atom-wise occupancies {da±}a . Default denotes standard SAD initialization, while Oracle uses all converged electronic components from the corresponding completed calculation. ✓ denotes a converged component and × denotes SAD/default initialization. OM, MP and CG denote spin initialization from converged atomic magnetic moments, MPRelaxSet, and CHGNet-predicted magnetic moments, respectively. Positive values indicate faster calculations and negative values indicate slower calculations relative to the corresponding baseline.
Variant ( ISPIN=1] )
Default (SAD)
Oracle
ChargE3Net 1,∗
DFT Steps ↓
15.33 ± 8.02
9.21 ± 7.82
12.34 ± 8.69
DFT Time ↓
163.30 ± 326.32 s
114.89 s ± 276.57
141.49 ± 298.70 s
DFT Steps Saved ↑
–
39.90 %
19.51 %
DFT Time Saved ↑
–
29.64 %
13.35 %
Appendix
Table 7: 1 Koker et al. (2024) . ∗ ChargE3Net is initialized with converged augmentation occupancies. DFT convergence analysis with different initializations. The tests were performed on the non-magnetic part of the MP test set.
Dataset
Training set
RMSE ↓
MAE ↓
MaxAE ↓
MaxAE top11↓
MP
1k
0.0340 ± 0.0206
0.0120 ± 0.0065
4.504
1.487
10k
0.0183 ± 0.0113
0.0064 ± 0.0034
2.829
0.796
50k
0.0130 ± 0.0096
0.0046 ± 0.0028
1.931
0.664
Full
0.0118 ± 0.0091
0.0041 ± 0.0026
1.114
0.646
GNoME
1k
0.0464 ± 0.3916
0.0120 ± 0.0287
366.85
1.724
10k
0.0302 ± 0.3915
0.0078 ± 0.0284
366.58
0.868
Appendix
Table 8 : Physical augmentation occupancy prediction errors versus training-set size on the full MP and GNoME test sets. RMSE and MAE are means over per-structure values, while MaxAE is the largest single-coefficient error. MaxAE top 11 reports the largest error after excluding the ten most extreme coefficients.
Model
Steps
Trainable params.
MAE ↓
RMSE ↓
MaxAE ↓
CJM ( Focassio et al., 2024 )
1k
1,758
0.0130
0.0459
0.9137
AugNet, zero-shot
–
–
0.3854
1.3439
14.9708
AugNet, from scratch
1k
3M
0.0010
0.0258
0.4969
AugNet, head fine-tune
1k
73k
0.1801
0.7039
8.2880
10k
73k
0.0122
0.0400
0.9490
AugNet, full fine-tune
1k
3M
0.0157
0.0352
0.8535
Appendix
Table 9: Transfer of MP-pretrained AugNet to the MoS2 PAW setup of Focassio et al. (2024) . Their calculations use a different Mo PAW dataset from Materials Project, so the target augmentation representation changes and direct zero-shot transfer is not expected. “Trainable params.” denotes the number of parameters optimized during adaptation.
Dataset
Metric
Default
Oracle
CNEI-EFI
CNEI-C3Net
MP
ρ~+ NMAE ↓
–
–
0.58%
0.54%
ML init time ↓
–
–
(0.24+0.37) s
(78.73+0.37) s
SCF steps ↓
22.05
10.41
19.05
18.36
DFT time ↓
623.84 s
302.04 s
529.43 s
506.16 s
Total time ↓
623.84 s
302.04 s
530.04 s
585.26 s
SCF steps saved ↑
–
52.78%
13.62%
16.76%
Appendix
Table 10 : Detailed comparison of reference-free ML PAW initializations and resulting DFT performance on the test sets of MP and GNoME. Total time includes both ML initialization and DFT execution. For both CNEI-EFI and CNEI-C3Net, we add the time it takes to evaluate spin-ELECTRAFI (0.24s/0.15s) and spin-AugNet (0.05s both) as well as CHGNet (0.03s both) to ELECTRAFI and ChargE3Net.
Dataset
Metric
Default
Oracle
CNEI-EFI
CNEI-C3Net
MP
ρ~+ NMAE ↓
–
–
0.67 %
0.78 %
ML Init Time ↓
–
–
(0.24 + 0.37) s
(87.25 + 0.37) s
SCF steps ↓
27.82
12.43
24.97
24.48
DFT time ↓
882.45 s
377.16 s
777.63 s
744.13 s
Total time ↓
882.45 s
377.16 s
778.24 s
831.75 s
SCF steps saved ↑
–
55.30%
10.22%
11.98%
Appendix
Table 11 : The magnetic subset counterpart of table 10 . The spin-ELECTRAFI model measures (0.24s/0.15s).
Dataset
Metric
Default
Oracle
CNEI-EFI
CNEI-C3Net
MP
ρ~+ NMAE ↓
–
–
0.55 %
0.50 %
ML Init Time ↓
–
–
(0.17+0.30) s
(72.11+0.30) s
SCF steps ↓
16.83
8.58
13.68
12.80
DFT time ↓
389.41 s
233.95 s
304.45 s
290.45 s
Total time ↓
389.41 s
233.95 s
304.92 s
362.86 s
SCF steps saved ↑
–
49.01%
18.71%
23.91%
Appendix
Table 12 : The non-magnetic subset counterpart of table 10 . The spin-ELECTRAFI model measures (0.17s/0.11s) on these subsets.
Site
Partial waves
ε
1
(2,2,0,0,1,1)
4.6×10−7
2–5
(0,0,1,1)
1.1 – 10.1×10−7
Appendix
Table 13 : Parameter-free verification of the equivariant transformation law over 241 held-out rotations.
Group
Hyperparameter
Value
Backbone
Hidden width
64
Max. spherical order ℓmax
3
Cutoff radius rmax (Å)
6.0
Interaction layers
2
Correlation order
3
Avg. number of neighbors
64.3
Appendix
Table 14: Hyperparameters used for training AugNet and spin-AugNet. The spin-summed and spin-difference models share all architectural and optimization settings and differ only in the predicted channel and reference baseline.
Group
Parameter
Value
Backbone (EScAIP)
Layers
2
Hidden size
256
Attention heads
32
Atom embedding size
128
Edge distance embedding
512 (expansion 600 )
Node direction embedding
256 (expansion 13 )
Appendix
Table 15: Hyperparameters for the ELECTRAFI and spin-ELECTRAFI models, mirroring the choices in Elsborg et al. (2026b) .
We present \textsc{dm-PhiSNet}, a physically constrained \textsc{PhiSNet}-based equivariant model that predicts one-electron reduced density matrices (1-RDMs) directly from molecular geometries in an atomic-orbital (AO) basis for accelerated self-consistent field (SCF) workflows. Training follows a two-stage schedule with progressively introduced physically motivated objectives, and the resulting predictions are refined by a lightweight analytic block. This block enforces electron-number conservation, drives the 1-RDM toward generalized idempotency in the AO metric, and regularizes the occupation spectrum of the Löwdin-orthogonalized density. Across six closed-shell systems -- H2O, CH4, NH3, HF, ethanol, and NO3− -- the refined 1-RDMs provide SCF initial guesses that substantially reduce iteration steps by 49--81% relative to standard initializations. Beyond SCF acceleration, the learned 1-RDMs yield accurate one-shot total energies and Hellmann--Feynman atomic forces without force supervision, indicating that the model captures chemically meaningful electronic structure. These results demonstrate that combining equivariant learning with analytic constraint enforcement provides a simple, general route to solver-ready density-matrix initializations and accelerated SCF workflows.
Zuriel Y. Yescas-Ramos, Andrés Álvarez-García, Huziel E. Sauceda
Instituto de Física, Universidad Nacional Autónoma de México, Cd. de México C.P. 04510, Mexico
The cost of Kohn-Sham density functional theory (KS-DFT) calculations scales with the number of solver iterations, which depends on the quality of the initial guess. Machine learning methods that predict initial guesses from molecular geometry can reduce this cost, but matrix-prediction models fail when extrapolating to larger molecules, degrading rather than accelerating convergence [Liu et al., 2025]. We show that this failure is a supervision problem, not an extrapolation problem: models trained on ground-state targets fit those targets well out of distribution, yet produce initial guesses that slow convergence. Solver-Aligned Initialization Learning (SAIL) resolves this for both Hamiltonian and density matrix models by differentiating through the self-consistent field (SCF) solver end-to-end. We introduce the Effective Relative Iteration Count (ERIC), a correction to the commonly used RIC that accounts for hidden Fock-build overhead. On QM40, which contains molecules up to 4× larger than the training distribution, SAIL reduces ERIC by 37% (PBE), 33% (SCAN), and 28% (B3LYP), more than doubling the previous state-of-the-art reduction on B3LYP. On QMugs molecules 10× larger than the training set, SAIL delivers a 1.35× wall-time speedup at the hybrid level of theory, extending ML SCF acceleration to large drug-like molecules.
Eike S. Eberhard, Viktor Kotsev, Timm Güthle +1
1Technical University of Munich (TUM) · 2Munich Data Science Institute (MDSI) · 3Munich Center for Machine Learning (MCML)
Machine-learned (ML) operator models can be trained to predict density functional theory (DFT) Hamiltonian/density matrices at significantly reduced computational cost, thus extending electronic-structure calculations to previously unfeasible scales. Here, we introduce MALOQ (Massively Accelerated Learning of Operators for Quantum Transport), an application built to train on and predict electronic-structure matrices for systems made of few to 100k atoms, described by large basis sets, and covering a wide range of atomic elements. Based on a state-of-the-art, SO(2)-equivariant backbone architecture, MALOQ provides (i) custom data-processing kernels to handle high-rank Hamiltonian matrix data and (ii) a scalable edge-wise distribution of atomic graph(s). Trained on the largest molecular Hamiltonian datasets available today, it reduces time-per-epoch by over 30% compared to a molecule-wise-distributed framework, and enables inference on material graphs of arbitrary size. We demonstrate scalable training and inference for 3,000-12,000 atoms on the Alps supercomputer, up to 192 GPUs and 256 GPUs, respectively.
Manasa Kaniselvan, Alexander Maeder, Denghui Lu +2