Physics-Aligned Electronic Ground-State Learning Improves Generalization
Organizations: Technical University of Munich (TUM) · Munich Data Science Institute (MDSI) · Munich Center for Machine Learning (MCML) · NVIDIA
Abstract
Machine-learned interatomic potentials (MLIPs) excel at in-distribution tasks, accelerating drug and material development, yet they struggle to generalize out-of-distribution. We propose to push the cost-accuracy Pareto frontier by designing observable-agnostic electronic ground-state descriptor models (GSMs) with computational costs situated between MLIPs and Kohn-Sham density functional theory (KS-DFT). We align the learning objectives and architectures of GSMs with the governing equations of KS-DFT by enforcing physical constraints and removing optimization pressure on unphysical or irrelevant degrees of freedom. In our size-extrapolation experiments from QM9 to QM40, our combined contributions OrthoNormal-Loss (ON-Loss) and Grassmann Restricted Occupied-Orbital Training (GROOT) reach a 79.1% energy and 83.4% force mean absolute error (MAE) reduction over previous state-of-the-art density GSMs. For Hamiltonian GSMs, ON-Loss and Residual Optimal-gauge Conditioning-aware KS-Eq. Training (ROCKET) together reduce the energy and force MAEs of the strongest baseline by 99.8% and 95.9%, respectively. Using a self-consistency rejection criterion, we filter out extrapolation errors on QMugs, rejecting fewer than 0.4% of predictions while reaching an energy MAE of 0.07 mHa. Finally, we demonstrate the efficiency of label-free self-consistency fine-tuning, and transfer GSMs to reactive chemistry in Transition1x, reaching energy errors below chemical accuracy.
Figures & tables
| Observable | Leading order in | |
| Unconstrained | ||
| Manifold Constrained | ||
Appendix figures & tables46 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value | Notes |
| Basis set | def2-SVP | Ahlrichs split-valence double-zeta basis with polarization functions ( Weigend and Ahlrichs, 2005 ) . |
| XC functional | B3LYP | Three-parameter hybrid combining exact exchange ( Becke, 1993 ) with LYP ( Lee et al., 1988 ) and VWN local correlation ( Vosko et al., 1980 ) in the parameterization of Stephens et al. (1994) ; we use the VWN-RPA form, as Gaussian does. |
| Density fitting | RI-J, RI-K | Density fitting of both the Coulomb ( Vahtras et al., 1993 ) and the exact-exchange term ( Weigend, 2002 ) . |
| Auxiliary basis | def2-universal-JKFIT | One auxiliary basis serves both the RI-J and the RI-K contractions ( Weigend, 2008 ) . |
| Spin treatment | Restricted | Closed-shell restricted Kohn–Sham. |
| Quadrature grid | Level 3 | PySCF grid level for the XC energy and potential ( Becke, 1988 ; Sun et al., 2020 ) . |
| Stage | Cost | Computed quantities |
| Cheap | One Fock build | The predicted density matrix is the model output for density readouts and the density matrix after diagonalization of the predicted ( eq. 1 ) for Hamiltonian readouts. From we build once. Every quantity that needs no further Fock build and no nuclear gradient counts as cheap. |
| SCF | Up to reference | A full SCF started from with the settings of table 2 . A run is counted as not converged after twice the cycle count of the MINAO-started reference SCF. |
| Forces | One nuclear gradient | Nuclear forces evaluated at the density matrix , compared with the reference forces . |
| Parameter | Value | Notes |
| Release | molecules | Small organic molecules with up to 9 heavy atoms (C, N, O, F), enumerated from GDB-17, with equilibrium geometries relaxed at B3LYP/6-31G(2df,p). |
| Our copy | molecules | We drop the molecules that the authors flag as uncharacterized. |
| Geometries | 0 K minima | Only the released thermochemistry refers to 298.15 K. |
| Elements | H, C, N, O, F | molecules contain F. |
| Size range | 3–29 atoms, 1–9 heavy atoms, 10–74 electrons, 24–226 basis functions. | |
| Training labels | The QM9 checkpoints are trained on labels converged to and otherwise identical to table 2 . The total energies agree with the evaluation reference to a median of , with a maximum of . |
| Parameter | Value | Notes |
| Release | molecules | Drug-like molecules from ZINC14 with 10–40 heavy atoms (C, N, O, S, F, Cl), covering the size range of 88% of FDA-approved drugs. Geometries are pre-optimized with GFN2-xTB and relaxed at B3LYP/6-31G(2df,p), the level of theory of QM9; molecules with imaginary frequencies are removed. |
| Geometries | 0 K minima | |
| Element filter | molecules | H, C, N, O, F only, 78–272 electrons. No molecule with 20 or 21 heavy atoms passes the filter, since all of them contain S or Cl. |
| Electron bins | Ten bins of 20 electrons. The two molecules with 78 electrons fall below the first bin. The sparsest bin, 80–100 electrons, holds 311 molecules. | |
| Validation (hyperparameter selection) | ||
| Rule | 50 regular, 25 hard | Used to select learning rate and weight decay (App. L ). We remove them from the pool before drawing the test sets. |
| Parameter | Value | Notes |
| Release | conformers of molecules | Neutral, closed-shell drug-like molecules from ChEMBL 27 with 3–100 heavy atoms (H, C, N, O, F, P, S, Cl, Br, I). Each molecule has three conformers, selected by clustering GFN2-xTB metadynamics snapshots and then optimized with GFN2-xTB. |
| Geometries | 0 K minima | GFN2-xTB local minima; 300 K is only the temperature of the conformer search. |
| Our copy | molecules | All conformers via OpenQDC, 22–850 electrons. Molecules are identified by SMILES, which merges some ChEMBL entries. |
| Training split | ||
| Rule | Every conformer with at most 150 electrons ( conformers of molecules; 3–22 heavy atoms, 4–54 atoms) that is not in the validation split; all conformers kept. | |
| Elements | H, C, N, O, F, P, S, Cl | Molecules with Br or I are excluded. |
| Parameter | Value | Notes |
| Release | frames | A single flexible molecule, 3-(benzyloxy)pyridin-2-amine, with three consecutive rotatable bonds. The MD frames come from 25 ps Langevin trajectories at 300, 600 and 1200 K (1 fs time step) started in five conformational pockets; the dihedral scans are constrained DFT geometry optimizations. The released labels are B97X/6-31G(d), computed with ORCA. |
| Molecule | C 12 H 12 N 2 O | 27 atoms, 106 electrons, 270 basis functions. |
| Used subsets | Dihedral scans | frames. We do not use the two training sets (500 frames each) or the MD test sets ( / / frames at 300 / 600 / 1200 K). |
| Dihedrals | Three consecutive rotatable bonds of the benzyl-ether linker: about the pyridine C–O bond, about the O–CH 2 bond, and about the CH 2 –phenyl bond. A fourth constraint fixes the amino group. | |
| Grid | in steps of ; in steps of . | |
| Geometries | 0 K | Constrained minima at each with fixed. |
| Parameter | Value | Notes |
| Release | frames | Configurations on and around the reaction paths of reactions from Grambow et al. (2020) , each with up to seven heavy atoms (C, N, O). The frames are images of nudged-elastic-band and climbing-image NEB runs, labeled at B97x/6-31G(d). |
| Elements | H, C, N, O | 4–23 atoms. |
| Geometries | Reaction paths | Non-equilibrium configurations: many lie on or near reaction barriers. |
| Release split | / 225 / 287 reactions | Training / validation / test reactions of the release. |
| Reaction coordinate | on the reactant side and on the product side, where is the Kabsch RMSD; R, TS and P sit at , and . | |
| Training split (label-free) | ||
| model type | supervision target | ||
| (AO) | |||
| (AO) | |||
| fitted gauge | |||
| fitted gauge (AO) |
| Hyperparameter | Value | Notes |
| Number of layers | 4 | |
| Sphere channels | 192 | Shared by backbone and decoder. |
| Maximum degree | 4 | def2-SVP off-diagonal requires . |
| Edge channels | 128 | |
| Node hidden channels | 128 | |
| Radial basis | 128 | fixed Gaussians on Å |
| PaiNN | NequIP | MACE | |
| Interaction layers | 3 | 4 | 3 |
| Features / channels | 128 | 128 per | 256 |
| Max. angular order | (vectors) | ||
| Correlation order | – | – | 3 (body order 4) |
| Parity | – | off | – |
| Radial basis | Gaussian, 20 | Bessel, 8 | Bessel, 8 |
| Hyperparameter | Value | Notes |
| Batch size | ||
| Optimizer | Muon | Muon on the 2D hidden weight matrices, AdamW on all remaining parameters. Both optimizer branches use Nesterov momentum, following the Optax (contrib) defaults ( Babuschkin et al., 2020 ) . We also tried the same momentum as used for standard Adam but observed no improvements. |
| Muon momentum | Default ( Jordan et al., 2024 ) . | |
| Muon Newton–Schulz steps | Default ( Jordan et al., 2024 ) . | |
| Adam | Default ( Kingma and Ba, 2015 ) . | |
| Muon consistent RMS | Each orthogonalized Muon update is scaled by , giving approx. shape-independent update RMS ( Liu et al., 2025 ) . |
| Hyperparameter | Value | Notes |
| Optimal shift std | Standard deviation of the -optimal gauge shifts over the training split. | |
| Restraint weight | Calibrated as over the full training split ( molecules). | |
| EMA decay | Decay of the EMA of the optimal overlap shift center. Initialized to training split optimal . |
| QM40 loss | ||||||||
| ep | lr | wd | seed | val loss | reg. 50 | hard 25 | Muon | Adam |
| abs -learning (no-rescale) | ||||||||
| 16 | 3e-4 | 1e-2 | 995756 | 5.73 | 16.3 | 29.6 | 0.16 | 0.55 |
| 16 | 1e-3 | 1e-2 | 995756 | 4.82 | 14.8 | 27.8 | 0.17 | 0.44 |
| 16 | 3e-3 | 1e-2 | 995756 | 4.48 | 13.8 | 26.0 | 0.26 | 0.40 |
| " | " | " | 684851 | 4.40 | 14.1 | 36.9 | 0.25 | 0.41 |
| QM40 loss | ||||||||
| ep | lr | wd | seed | val loss | reg. 50 | hard 25 | Muon | Adam |
| ON-Loss | ||||||||
| 12 | 1e-3 | 1e-2 | 0 | 1.46 | 7.54 | 9.06 | 0.17 | 0.48 |
| " | " | " | 749811 | 1.47 | 7.10 | 9.12 | 0.17 | 0.47 |
| 12 | 1e-3 | 3e-2 | 0 | 1.40 | 7.99 | 9.51 | 0.10 | 0.35 |
| 12 | 3e-3 | 1e-3 | 0 | 1.43 | 6.03 | 8.60 | 0.42 | 0.59 |
| QM40 loss | ||||||||
| ep | lr | wd | seed | val loss | reg. 50 | hard 25 | Muon | Adam |
| + GROOT proj | ||||||||
| 16 | 1e-3 | 3e-2 | 193360 | 0.38 | 1.24 | 2.48 | 0.09 | 0.32 |
| 16 | 3e-3 | 3e-3 | 193360 | 0.39 | 2.33 | 2.52 | 0.36 | 0.47 |
| 16 | 3e-3 | 1e-2 | 193360 | 0.39 | 1.17 | 2.97 | 0.22 | 0.36 |
| " | " | " | 141727 | 0.39 | 1.06 | 2.44 | 0.21 | 0.35 |
| QM40 loss | ||||||||
| ep | lr | wd | seed | val loss | reg. 50 | hard 25 | Muon | Adam |
| Density GROOT proj , pre/post supervision | ||||||||
| 16 | 3e-3 | 1e-2 | 193360 | (2.17) | 3.75 | 5.77 | 0.23 | 0.38 |
| Density GROOT proj , AO-frame prediction | ||||||||
| 16 | 3e-3 | 1e-2 | 193360 | 0.58 | 1.39 | 3.03 | 0.21 | 0.34 |
| Hamiltonian ROCKET, AO-frame prediction | ||||||||
| QM40 loss | ||||||||
| ep | lr | wd | seed | QM9 val | reg. 50 | hard 25 | Muon | Adam |
| abs -learning | ||||||||
| 16 | 1e-4 | 1e-3 | 684464 | 0.78 | 3.95 | 12.2 | 0.18 | 0.62 |
| -learning | ||||||||
| 16 | 3e-6 | 1e-3 | 684464 | 2.12 | 8.7 | 18.4 | 0.18 | 0.63 |
| 16 | 1e-5 | 1e-3 | 684464 | 1.12 | 5.3 | 15.1 | 0.18 | 0.63 |
| QM40 loss | ||||||||
| ep | lr | wd | seed | QM9 val | reg. 50 | hard 25 | Muon | Adam |
| -learning | ||||||||
| 16 | 1e-4 | 1e-3 | 684464 | 11.8 | 12.1 | 56 | 0.18 | 0.63 |
| -learning | ||||||||
| 16 | 3e-6 | 1e-3 | 684464 | 36.0 | 31.6 | 87 | 0.18 | 0.63 |
| 16 | 1e-5 | 1e-3 | 684464 | 9.17 | 12.1 | 110 | 0.18 | 0.63 |
| QM40 loss | ||||||||
| ep | lr | wd | seed | QM9 val | reg. 50 | hard 25 | Muon | Adam |
| -learning | ||||||||
| 16 | 3e-5 | 3e-3 | 670079 | 1.60 | 3.6 | 9.7 | 0.18 | 0.63 |
| " | " | " | 435126 | 1.61 | 3.7 | 9.9 | 0.18 | 0.62 |
| " | " | " | 301732 | 1.63 | 3.7 | 10.0 | 0.18 | 0.63 |
| -learning | ||||||||
| QM40 loss | ||||||||
| ep | lr | wd | seed | QM9 val | reg. 50 | hard 25 | Muon | Adam |
| 1st stage (unrestrained gauge) | ||||||||
| 16 | 1e-4 | 1e-2 | 443905 | 0.96 | 41.0 | 32.5 | 0.17 | 0.60 |
| 16 | 3e-4 | 1e-2 | 443905 | 0.75 | 14.2 | 24.5 | 0.16 | 0.56 |
| 16 | 1e-3 | 1e-2 | 443905 | 0.67 | 24.6 | 26.8 | 0.17 | 0.44 |
| 2nd stage (restrained gauge) | ||||||||
| LR | Validation self-consistency residual after epoch | |||||||||||
| 1 † | 2 † | 3 | 4 † | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 † | |
| GROOT rot – density residual, (dimensionless) | ||||||||||||
| ∗ | 00 12.4 | 000 9.10 | 8.47 | 7.59 | 7.00 | 6.98 | ||||||
| 000 9.08 | 000 7.41 | 7.00 | 6.65 | 6.20 | 5.33 | 5.26 | 4.98 | 4.88 | 4.84 | 4.81 | 4.81 | |
| 000 8.13 | 000 7.42 | 6.71 | 5.97 | 5.86 | 5.25 | 5.01 | 4.66 | 4.34 | 4.28 | 4.23 | 4.19 | |
| 00 10.6 | 000 8.75 | 8.45 | 7.38 | 6.36 | 5.50 | 5.19 | 4.71 | 4.41 | 4.15 | 4.07 | 4.04 | |
| [mEh] | dipole [D] | [ ] | |||||||||
| lr | wd | loss | reg. | hard | reg. | hard | reg. | hard | cata. | ||
| NON losses | |||||||||||
| Density | |||||||||||
| abs. | 3 | 3e-3 | 1e-2 | 259 | 1 655 | {\color[rgb]{1,0,0}31.9} | {\color[rgb]{1,0,0}0.50} | {\color[rgb]{1,0,0}4.13} | 11.7 | ||
| 3 | 3e-3 | 1e-2 | 76 | 263 | {\color[rgb]{1,0,0}4.50} | {\color[rgb]{1,0,0}0.38} | {\color[rgb]{1,0,0}1.4} | {\color[rgb]{1,0,0}1.3} | |||
| Hamiltonian | |||||||||||
| (base 2) | |||||
| Model | Seeds | NON | ON | NON | ON |
| [m ] | |||||
| (NON) | 3 | ||||
| (ON) | 3 | ||||
| (NON) | 1 | ||||
| (ON) | 3 | ||||
| Model | LR | WD | Excluded | Self-consistent | Dipole | Forces | HOMO–LUMO | ERIC | |
| (%) | energy (mHa) | (D) | (e) | (mHa/Å) | (mHa) | (ratio) | |||
| Density-space models | |||||||||
| (NON) | 0 62.9 [62, 66] | 0 9.65 000 0 (12.1) [9.1, 10] (12, 13) | 0 0.0419 0 (0.0553) [0.040, 0.043] (0.055, 0.056) | 0 0.0333 0 (0.0428) [0.032, 0.034] (0.043, 0.043) | 0 0.974 0 0 (1.26) [0.94, 1.0] (1.3, 1.3) | 0 0.578 0 0 (0.774) [0.56, 0.62] (0.77, 0.78) | 0 0.745 (0.754) [0.74, 0.75] (0.75, 0.75) | ||
| (NON), no rescale | 0 61.8 [60, 64] | 10.4 0000 0 (12.2) [10, 11] (12, 13) | 0 0.0421 0 (0.0545) [0.042, 0.042] (0.054, 0.055) | 0 0.0337 0 (0.0427) [0.034, 0.034] (0.042, 0.043) | 0 1.00 00 0 (1.27) [1.00, 1.0] (1.3, 1.3) | 0 0.570 0 0 (0.759) [0.56, 0.59] (0.75, 0.78) | 0 0.745 (0.754) [0.74, 0.75] (0.75, 0.75) | ||
| (NON) | 0 21.2 [19, 25] | 0 8.47 000 00 (9.32) [8.4, 8.5] (9.2, 9.6) | 0 0.0512 0 (0.0563) [0.051, 0.052] (0.055, 0.057) | 0 0.0399 0 (0.0439) [0.039, 0.040] (0.043, 0.045) | 0 1.19 00 0 (1.31) [1.2, 1.2] (1.3, 1.3) | 0 0.741 0 0 (0.813) [0.73, 0.77] (0.79, 0.86) | 0 0.751 (0.755) [0.75, 0.75] (0.75, 0.76) | ||
| Model | LR | WD | Excluded | Self-consistent | First-order corrected | Dipole | Forces | HOMO–LUMO | ERIC | |
| (%) | energy (mHa) | energy (mHa) | (D) | (e) | (mHa/Å) | (mHa) | (ratio) | |||
| Density-space models | ||||||||||
| (NON) | 100 [100, 100] | – . 000 0 (338) (245, 431) | – . 0000 00 (11.5) (8.0, 16) | – . 0000 00 (0.952) (0.76, 1.1) | – . 0000 0 (0.616) (0.55, 0.67) | – . 000 00 (10.1) (8.5, 12) | – . 000 00 (6.72) (5.8, 7.7) | – . 000 (0.830) (0.82, 0.84) | ||
| (NON), no rescale † | 100 | – . 000 0 (175) | – . 0000 000 (4.07) | – . 0000 00 (0.594) | – . 0000 0 (0.477) | – . 000 000 (8.45) | – . 000 00 (5.36) | – . 000 (0.821) | ||
| (NON) | 100 [100, 100] | – . 000 00 (75.7) (61, 98) | – . 0000 000 (2.17) (1.3, 3.8) | – . 0000 00 (0.581) (0.50, 0.72) | – . 0000 0 (0.446) (0.41, 0.52) | – . 000 000 (7.51) (6.7, 9.0) | – . 000 00 (5.06) (4.9, 5.3) | – . 000 (0.816) (0.81, 0.82) | ||
| Model | LR | WD | Excluded | Self-consistent | Dipole | Forces | HOMO–LUMO | ERIC | |
| (%) | energy (mHa) | (D) | (e) | (mHa/Å) | (mHa) | (ratio) | |||
| Density-space models | |||||||||
| (NON) | 100 [100, 100] | – . 00 (46920) (1214, 137230) | – . 000 (145) (8.3, 369) | – . 000 (12.9) (5.5, 25) | – . 00 0 (314) (69, 756) | – . 00 0 (76.3) (59, 104) | – . 000 (1.23) (1.1, 1.4) | ||
| (ON) | 100 [100, 100] | – . 00 (13720) (887, 31790) | – . 000 0 (14.2) ( 1.0 , 31) | – . 000 0 (2.03) (1.3, 3.2) | – . 00 00 (60.1) (8.7, 154) | – . 00 0 (10.8) (8.5, 14) | – . 000 (0.817) (0.81, 0.83) | ||
| GROOT rot | 0 61.5 [53, 66] | 10.4 0 000 ( 36.9 ) [10, 11] ( 28 , 44) | 0 0.279 00 (3.29) [0.27, 0.30] (3.3, 3.3) | 0 0.329 0 ( 0.647 ) [0.32, 0.34] ( 0.62 , 0.69) | 0 2.41 000 ( 3.27 ) [2.4, 2.4] ( 3.2 , 3.4) | 0 2.17 00 ( 3.48 ) [2.1, 2.3] ( 3.5 , 3.5) | 0 0.743 ( 0.776 ) [0.74, 0.74] ( 0.77 , 0.79) | ||
| Model | LR | WD | Excluded | Self-consistent | First-order corrected | Dipole | Forces | HOMO–LUMO | ERIC | |
| (%) | energy (mHa) | energy (mHa) | (D) | (e) | (mHa/Å) | (mHa) | (ratio) | |||
| Supervised, 24 epochs | ||||||||||
| (NON) | 0 4.01 (245) | 0.0142 0 (59.0) | 0.127 0 0 (3.92) | 0.0343 (1.43) | – . 000 0 (16.2) | 0.123 0 (9.80) | – . 000 (0.818) | |||
| (ON) | 0 53.2 | 10.1 0 0 (31.4) | 0.140 0 00 (0.365) | 0.0638 0 ( 0.143 ) | 0.0495 (0.155) | 0 0.696 00 (2.01) | 0.406 0 (1.10) | 0 0.628 (0.668) | ||
| GROOT rot | 00 1.12 | 0 3.22 00 (9.29) | 0.0712 0 (36.8) | 0.114 0 0 (4.69) | 0.101 0 (0.313) | 0 0.678 00 (3.11) | 0.481 0 (1.18) | 0 0.643 (0.654) | ||
| Model | LR | WD | Excluded | First-order corrected | Self-consistent | Dipole | Forces | HOMO–LUMO | |
| (%) | energy (mHa) | energy (mHa) | (D) | (e) | (mHa/Å) | (mHa) | |||
| (NON) | 100 | – . 00000 (0.204) | – . 0000 (13.5) | – . 0000 0 (0.187) | – . 0000 0 (0.148) | – . 00 0 (4.12) | – . 000 0 (3.32) | ||
| (ON) | 00 0.0426 | 0 0.00634 ( 0.00643 ) | 0 4.34 00 0 (4.35) | 0 0.108 0 0 ( 0.108 ) | 0 0.0572 0 ( 0.0572 ) | 0 1.27 0 ( 1.27 ) | 0 0.995 0 (0.995) | ||
| (NON) | 00 0.213 | 0 0.104 00 (0.604) | 0 0.157 0 0 (1.26) | 0 0.235 0 0 (0.285) | 0 0.183 0 0 (0.188) | 0 4.35 0 (4.54) | 0 1.19 0 0 (1.29) | ||
| Model | LR | Excluded | First-order corrected | Self-consistent | Dipole | Forces | HOMO–LUMO | ERIC | |
| (%) | energy (mHa) | energy (mHa) | (D) | (e) | (mHa/Å) | (mHa) | (ratio) | ||
| Before self-consistency training (the QMugs-trained parents) | |||||||||
| (NON) | 0 95.9 | 0 0.0363 (33.4) | 13.3 00 (142) | 0 0.0680 (2.09) | 0 0.0396 (0.667) | 0 2.80 (47.0) | 0 0.784 (17.8) | 0 0.769 (0.980) | |
| (ON) | 0 84.7 | 0 0.0386 0 ( 3.58 ) | 44.1 00 (116) | 0 0.0416 ( 0.805 ) | 0 0.0323 ( 0.269 ) | 0 2.64 ( 21.8 ) | 0 0.455 0 ( 7.99 ) | 0 0.741 ( 0.896 ) | |
| GROOT rot | 0 72.4 | 0 0.202 0 (63.7) | 0 0.574 0 (33.7) | 0 0.146 0 (4.14) | 0 0.0713 (1.01) | 0 4.37 (66.4) | 0 1.49 0 (23.9) | 0 0.794 (0.999) | |