NEMORA: Neural Equivariant Multipole Operators for Long-Range Atomistic Learning
Authors: Jay L. Kaplan, Samuel Varner, Rebecca Willett, Juan J. de Pablo
Organizations: Courant Institute School of Mathematics, Computing, and Data Science, New York University, New York, NY 10012, USA · Department of Chemical and Biomolecular Engineering, Tandon School of Engineering, New York University, Brooklyn, NY 11201, USA · Department of Statistics, The University of Chicago, Chicago, IL 60637, USA · Department of Computer Science, The University of Chicago, Chicago, IL 60637, USA · NSF-Simons National Institute for Theory and Mathematics in Biology, Chicago, IL 60611, USA · Department of Physics, New York University, New York, NY 10012, USA
Equivariant graph neural networks have emerged as foundational architectures for machine-learned interatomic potentials, approaching quantum-chemical accuracy at a fraction of the computational cost. These models describe local atomic environments accurately, but finite spatial cutoffs truncate long-range information flow, and stacking message-passing layers can lead to over-smoothing and over-squashing. Existing long-range extensions either prescribe a fixed analytical propagation kernel, restrict long-range communication to scalars or degree-preserving channels, are only approximately equivariant, or incur super-linear computational cost. Combining learnable long-range equivariant transport with multiscale many-body expressivity and efficient scaling for larger systems remains a central challenge. We introduce Neural Equivariant Multipole Operators (NEMORA), a neural equivariant extension of the Fast Multipole Method (FMM) for learning long-range tensorial representations. NEMORA generalizes the FMM's analytical multipole expansion and translation operators to learned equivariant counterparts on an adaptive spatial hierarchy. Its operators couple angular degrees and form many-body interactions across length scales, retaining the FMM's hierarchical organization and analytical radial factors as physical inductive biases while learning data-dependent long-range couplings. NEMORA evaluates in linear time and memory complexity, allowing it to treat larger systems than other long-range methods reaching hundreds of thousands of atoms, and it augments both symmetry-constrained and unconstrained short-range backbones. On non-local benchmarks, it reduces force and energy errors relative to the short-range backbones by over an order of magnitude and up to three orders of magnitude, respectively, which is better than or competitive with existing long-range extensions in accuracy.
Tensor expansion obtained from harmonic and biharmonic potentials and their derivatives; not a scalar radial pair [ Gumerov & Duraiswami, 2006 , Greengard et al., 2021 ] .
Appendix
Table 4: Analytical free-space kernel families in three dimensions. The first three rows use the separated form above; the remaining families require the indicated additional structure. Overall physical prefactors are omitted, and tensor kernels are stated away from their singularities.
Method
Equiv.
Learn.
MB
ℓ -mix
Hier.
Scaling
Scalar long-range corrections
SpookyNet [ Unke et al., 2021a ] , TensorMol [ Yao et al., 2018 ] , AIMNet [ Zubatyuk et al., 2019 , Anstine & Isayev, 2023 ]
○
○
○
○
○
O(N2)
SO3LR [ Kabylda et al., 2025 ] , FENNOL [ Plé et al., 2023 , Plé et al., 2024 ]
○
○
○
○
○
O(N⋅⟨NLR⟩)
4G–HDNNP [ Ko et al., 2021 ] , CENT/CENT2 [ Ghasemi et al., 2015 , Khajehpasha et al., 2022 ] , CELLI [ Fuchs et al., 2025 ] , FeNNix-Bio1 [ Plé et al., 2025 ] , Maruf et al. [2025]
○
○
○
○
○
O(NlogN) a
LES [ Cheng, 2025 , Kim et al., 2025 ] , Neural P 3 M [ Wang et al., 2024b ] , Ewald MP [ Kosmala et al., 2023 ] , Reciprocal Space NN [ Yu et al., 2022 ]
○
○
○
○
○
O(NlogN)
LODE [ Grisafi & Ceriotti, 2019 , Huguenin-Dumittan et al., 2023 , Faller et al., 2024 ]
○
○
◐
○
○
O(NlogN)
Appendix
Table 5: Long-range neural methods along six architectural axes. Equiv. : ℓ>0 tensorial long-range messages. Learn. : learnable propagation parameters, excluding source construction and readout. Many-body : pairwise (○), single-mode many-body via aggregation/path (◐), or both atom-level and coarse-level long-range coupling (●). ℓ -mix : angular channel mixing during long-range transport. Hier. : hierarchical multi-scale structure. Scaling : cost of the long-range component at fixed feature widths and block count; linear costs require bounded interaction lists, leaf occupancy or graph degree. Symbols describe the long-range mechanism, not the total body order of the complete potential. ○ = no, ◐ = partial/approximate, ● = yes.
Model
Cumulene
Carbon chain
NaCl
Au 2 /MgO
MACE
3
3
3
3
MACE–LES
3
3
3
3
MACE–RANGE
3
3
3
2
LOREM
3
3
3
3
EFA
3
3
3
2
MACE–MFN
3
3
3
NR
Appendix
Table 7: Numbers of runs contributing to the benchmark means and SEMs. Counts apply to energy and force unless noted. The separate 3–64-carbon sweep uses three runs for LOREM, EFA and NEMORA and one for the other models. NR denotes that no run count or no result is reported. Additional cumulene controls have three-run means but no available SEM in the source summary; SpookyNet is a published reference with no run count given there.
Free-space Coulomb, N=96
p
θ
M2L rows
ϵE
ϵF,∞
2
0.5
4
3.01×10−4
1.89×10−4
4
0.5
4
1.28×10−5
1.01×10−5
6
0.5
4
5.66×10−7
1.09×10−6
4
0.6
10
8.95×10−6
2.38×10−5
4
0.4
0
2.56×10−16
3.87×10−16
Appendix
Table 14: Fixed-charge numerical errors. In the free-space Coulomb test, ϵE=∣E−Eref∣/∣Eref∣ and ϵF,∞=∥F−Fref∥∞/∥Fref∥∞ use direct summation as reference. The periodic test uses one 256-ion configuration with charges ±0.6e and a converged Ewald reference (reciprocal-grid parameter dl=0.6 Å); its last column is the absolute force component RMS difference in meV/Å. Here w is the explicit image-shell radius. All entries use float64 and contain no prediction error against training labels.
Model
Non-periodic
Periodic
MACE
508,552
496,256
LES
37,096
23,128
RANGE
233,400
233,552
EFA
43,960
618,112
LOREM
≈ 115,000
11,520
MFN
512
512
Appendix
Table 19: Largest number of atoms N that fits on one H200 with fp32 precision for each model, in non-periodic clusters and periodic boxes. Note that LOREM and EFA use a different short-range architecture and therefore are not directly comparable.
Molecules
Atoms
Training
Validation
Test
128
384
788
328
474
256
768
578
219
346
Total
—
1366
547
820
Appendix
Table 21: Periodic water split sizes, in configurations.
Hyperparameter
Value
Interactions
2
ℓmax
3
Hidden irreps
128x0e + 128x1o + 128x2e
ν
3
Polynomial cutoff exponent
5
Radial basis functions
8
Appendix
Table 22: Shared parameters for short-range MACE backbone upon which all MACE-based long-range models are built.
Hyperparameter
Value
Optimizer
Adam
Learning rate
0.01
Energy:Force weight
10:100
SWA Energy:Force weight
1000:1000
EMA decay
0.99
Appendix
Table 23: Default optimizer settings for MACE, MACE–LES and NEMORA; per-system overrides are listed below.
Hyperparameter
Cumulene
Cumulene (3-64)
NaCl
Carbon
Cutoff radius (Å)
3.0
3.0
3.0
3.0
ℓmax,lr
2
2
2
2
ℓmax
6
6
6
6
Feature channels
128
128
128
128
Message-passing layers
1
1
1
1
Radial basis functions
8
8
8
8
Appendix
Table 24: LOREM hyperparameters for the cumulene and nonlocal charge-transfer benchmark systems.
Hyperparameter
Cumulene
Cumulene (3-64)
NaCl
Carbon
Cutoff radius (Å)
2.0
2.0
3.0
2.0
Feature channels
144
144
256
144
Message-passing layers
3
3
2
3
MP maximum degree
3
3
3
3
MP radial basis functions
32
32
32
32
ERA maximum degree
2
2
2
2
Appendix
Table 25: EFA hyperparameters for the cumulene and nonlocal charge-transfer benchmark systems. The EPE maximum frequency is π for non-periodic systems and π/4 for periodic systems.
Arm
Change relative to FULL
NOHIER
Remove the hierarchy, retaining the standalone Kelvin branch. Periodic root injection is also removed.
NOKELVIN
Remove the standalone Kelvin branch.
NOKESL
Remove the per-box Kelvin contribution to the downward pass.
FIXEDTREE
Replace adaptive coarsening by a fixed top-down octree.
MAXLEV1/3
Set maximum hierarchy depth to one or three.
LMAX0/1/3
Set source-feature and operator degree together. On cumulene, degrees below two also disable the quadrupole-dependent Kelvin head.
Appendix
Table 27: Definitions of the component-ablation arms. All arms are retrained. Joint changes are identified explicitly.
Accurate interatomic potentials enable molecular dynamics of materials, molecules, and interfaces beyond density-functional-theory length and time scales. Equivariant neural network potentials have improved the representation of local geometry. However, their deployable energy surfaces ultimately manifest through invariant scalar channels, whose aggregation and spectral resolution remain comparatively underexamined. Here we use Physics-Aware Neighborhood (PAN) pooling and Physics-Guided Spectral (PGS) mixers as controlled scalar-pathway probes: lightweight, symmetry-preserving modifications that act only on ℓ=0 channels while leaving the equivariant tensor backbone unchanged. Using MACE as a high-body-order mechanistic scaffold, PAN adds coordination-sensitive amplitude modulation, whereas PGS augments edge and readout scalar features with radial and tapered spectral bases. Across metallic Ag, covalent Si, a short-range ionic LiF/Li--F subset, and MD17/rMD17 molecules, this scalar-pathway correction reduces MACE force errors by 22--27% and energy errors by 19--22%; on systems with stress labels, stress errors decrease by 27--28%, at approximately 5% additional inference-FLOPs cost. Directionally consistent gains in Allegro and NequIP further indicate that the correction is portable across distinct short-range equivariant backbones, although effect sizes remain architecture-dependent. These results identify scalar-pathway fidelity as a practical design dimension for short-range equivariant interatomic potentials.
Jia Bi, Alin Marin Elena, Samuel Pinilla
Science and Technology Facilities Council, Harwell Campus, Didcot, OX11 0QX, Oxfordshire, United Kingdom. · Science and Technology Facilities Council, Keckwick Lane, Daresbury, WA4 4AD, Cheshire, United Kingdom. · Diamond Light Source, Harwell Science and Innovation Campus,2026 Didcot, OX11 0QX, Oxfordshire, United Kingdom.
Equivariant machine learning interatomic potentials (MLIPs) have revolutionized atomistic modeling, but accurate treatment of complex materials and molecular systems demands expensive models. This limits simulation length- and time-scales, with tensor products a key computational bottleneck. The recent emergence of foundation-scale MLIPs further exacerbates this challenge. We present Branch Interatomic Potential (BranchIP), a single-model framework for learned adaptive tensor product computation, trained with a novel distillation loss. In our experiments on two systems of physical interest, a heterogeneous catalysis system and a proton-conducting solid acid electrolyte, BranchIP accelerates MLIPs across model sizes by up to 2.4× while reducing memory usage by up to 2.6×. This is achieved while maintaining physical fidelity. Furthermore, the learned adaptive computation provides model interpretability by revealing which interactions demand deeper computation and showing how computational depth relates to chemical complexity and dynamics.
Laura Zichi, Gil Harari, Chuin Wei Tan +6
Harvard University · Robert Bosch LLC Research and Technology Center
Machine-learned (ML) operator models can be trained to predict density functional theory (DFT) Hamiltonian/density matrices at significantly reduced computational cost, thus extending electronic-structure calculations to previously unfeasible scales. Here, we introduce MALOQ (Massively Accelerated Learning of Operators for Quantum Transport), an application built to train on and predict electronic-structure matrices for systems made of few to 100k atoms, described by large basis sets, and covering a wide range of atomic elements. Based on a state-of-the-art, SO(2)-equivariant backbone architecture, MALOQ provides (i) custom data-processing kernels to handle high-rank Hamiltonian matrix data and (ii) a scalable edge-wise distribution of atomic graph(s). Trained on the largest molecular Hamiltonian datasets available today, it reduces time-per-epoch by over 30% compared to a molecule-wise-distributed framework, and enables inference on material graphs of arbitrary size. We demonstrate scalable training and inference for 3,000-12,000 atoms on the Alps supercomputer, up to 192 GPUs and 256 GPUs, respectively.
Manasa Kaniselvan, Alexander Maeder, Denghui Lu +2