Authors: Sandy Adhitia Ekahana, Aalok Tiwari, Pratik Saud, Aaron Bostwick, Chris Jozwiak, Eli Rotenberg, Jyoti Katoch
Organizations: Department of Physics, Carnegie Mellon University, Pittsburgh, 15213, PA, USA. · MAESTRO, Advanced Light Source, Lawrence Berkeley National Laboratory, Berkeley, CA, USA.
Artificial intelligence (AI) is becoming an increasingly useful tool across the experimental sciences, including angle-resolved photoemission spectroscopy (ARPES), which routinely produces large, multidimensional datasets of electronic structure. Recent advances in AI and machine learning (ML) have opened new opportunities across the entire ARPES workflow, from automated sample preparation and real-time data acquisition to post-experiment data analysis and comparison with theoretical calculations. Despite this progress, a comprehensive review of ML applications, their capabilities, and reliability across the different stages of ARPES workflow is still lacking. In this review, we first introduce ML methods that are most relevant to experimentalists working in condensed matter physics and materials science. We then follow the ARPES workflow, reviewing existing ML applications at each step and discussing their advantages, limitations and potential for future development. We also examine the current ARPES data landscape, where several open databases are available but remain relatively small and fragmented compared with large, shared datasets such as ImageNet. Given these limitations, we suggest that the community focus on sharing pretrained models that can be further trained, adapted to specific tasks, and redistributed, while working toward a larger and standardized open ARPES dataset repository. Finally, we discuss our perspectives on the future of AI within the ARPES workflow using a six-level framework of laboratory automation, highlighting the opportunities and challenges in moving toward a fully autonomous, self-driving ARPES laboratory.
Figures & tables
Figure 1: Schematic of ARPES Workflow The typical ARPES workflow progresses from left to right, from sample growth to analysis via measurement in an ultrahigh vacuum (UHV) chamber. Solid straight arrows indicate the nominal forward path, while solid curved arrows mark active learning–based iterations designed to acquire a rapid and complete dataset during a small experimental window. It represents an iterative closed loop between measuring, analyzing, and deciding the next step. Combined with other complementary studies, this reveals the underlying electronic properties of quantum materials
Figure 2: A practical AI terminology for experimental physicists. Machine learning splits by architecture into traditional compact models and deep learning networks. The four learning protocols represent a separate, independent axis of training models. The asterisk marks reinforcement learning, which in current practice is typically implemented with deep networks rather than as an architecture-independent protocol. The figure is a visual guide rather than a comprehensive glossary. Detailed examples appear in the text.
Figure 3: Automated sample preparation and preliminary characterizations. ( a ) The five pre experiment stages for ARPES experiment, shaded by the extent of automation. ( b ) An autonomous inorganic synthesis laboratory (A-lab) for air-stable targets, in which an ML algorithm proposes recipes, while robotic synthesis and learned phase identification form a self-revising loop (adapted from Szymanski et al. (2023) ). ( c ) Robotic four dimensional pixel assembly, from wafer scale growth and patterning to automated stacking, with control over layer number, composition, lateral position and twist angle. Scale bar: 50 μ m (adapted from Mannix et al. (2022) ).
Figure 4: Route to experimental automation ( a ) KMGP-driven data collection identifying distinct major phases on the sample, colored by cluster label Melton et al. (2020) . ( b ) The GPR-guided acquisition loop showing the autonomous data acquisition process. The colored boxes represent asynchronously running processes, while the red arrows indicate real-time, ML-based communication between the instrument and Python scripts Ágústsson et al. (2024b) . ( c ) GPR score, mean uncertainty ( ΔPmean ), and gain versus number of scans for Bi 2 Te 3 spin-ARPES, illustrating a quantitative stopping criterion for the acquisition loop Iwasawa et al. (2024) . ( d ) A synthetic trained CNN scores the measured spectral quality across a Bi 2 Se 3 grid scan. The predicted map is compared with a conventional benchmark score and decides where to measure Na et al. (2025) . Combined, these steps form a closed-loop experimental workflow: locating sample variations, selecting target regions, measuring until an objective stopping criterion is reached, evaluating spectral quality, and iterating to optimize spatial and electronic coverage.
Figure 5: Denoising network training sequence on Poisson subsampled ARPES data ( a ) Generation of noisy training data from high count Au(111) surface states spectra. From original high count data x , low count noisy copies x′ are generated via Poisson subsampling, simulating rapid acquisition. The denoising process is denoted f(x′) , while the generation of the noisy training set is denoted f−1(x) . ( b ) Training loop for the denoising neural network, where the loss function drives parameter optimization. The loss function L(x,f(x′)) combines mean absolute error (MAE) and multi-scale structural similarity index measure (MS-SSIM). The weighted formulation ensures that reconstructed spectra retain sharp spectral features essential for lineshape and second-derivative analysis, enabling quantitative comparison between short and long acquisitions on the same ARPES cuts (adapted from Kim et al. (2021) ).
Figure 6: Unsupervised clustering for spatial domain identification Panel (a) shows raw spatially resolved spectral intensity. Panel (b) applies k -means clustering with three clusters (1,2,3), assigning each point exclusively to one cluster. Panel (c) applies fuzzy c-means soft clustering, where each point receives probabilistic membership weights across clusters; the background shading indicates low membership confidence or non-sample regions, revealing sharp domain boundaries (adapted from Iwasawa et al. (2022) )
Figure 7: Band extraction of NbIrTe 4 using DFT, conventional analysis, and machine learning. ( a ) Density functional theory calculation of the band structure of NbIrTe 4 . The bands crossings are marked with red arrows. ( b ) Corresponding experimental ARPES data showing the same band structure. ( c–e ) Results from three band extraction methods applied to the experimental data: maximum curvature tracing ( c ), minimum gradient method ( d ), and CNN ( e ). The CNN approach recovers band positions comparable to conventional methods while reducing manual intervention and improving robustness on weak or overlapping features (adapted from Ref. Peng et al. (2020) ).
Figure 8: Flow chart of an ML regression procedure for extracting hidden self energies in Bi2212. A Boltzmann machine represents the normal ( Σnor(k,ω) ) and anomalous self energies ( Σano(k,ω) ), and is optimized against the measured spectral weight A(k,ω) through nested training and test loops until convergence. Adapted from Ref. Yamaji et al. (2021) ; Yamaji et al. (2023)
Figure 9: DeepH: An example of a theoretical framework where a DFT Hamiltonian surrogate predicts the Hamiltonian from the atomic structure. Adapted from Ref. Li et al. (2022) . ( a ) The DFT Hamiltonian H^DFT({R}) , which is a function of atomic coordinates, can be obtained by self-consistent field (SCF) calculations or learned by a neural network for efficient ab initio electronic structure calculations, yielding its physical properties ( b ) The network exploits the nearsightedness principle of electronic matter making the surrogate size transferable. The Hamiltonian matrix elements in the localized basis are nonzero only between neighboring atoms (within a cutoff RC ) and are influenced only by the local neighborhood (within RN ).
Figure 10: The six automation levels of ARPES , adapted from Ref. Le Houx (2026) . Level 0 (Manual) represents fully manual ARPES measurements with only the minimal necessary motorization. Level 1 (Scripted) represents the stage at which full motorization is realized and predefined scripts control all motion. Level 2 (Reflexive) represents the stage at which a subset of decisions is transferred to pre defined scalar feedback logic. Level 3 (Heuristic) represents the stage at which decisions are made heuristically through physics-informed inference. Level 4 (Supervisory) represents the stage at which an expert supervises an agentic bot controlling the experiment, while retaining key decision points and the right to veto the agent’s actions. Level 5 (Autonomous) represents the stage at which the human engages with the agent at the level of scientific discussion, with minimal direct decision-making. Representative example is an agent that can propose and independently execute a new experiment based on prior discussion.
Data-driven machine learning (ML) techniques have become an essential tool in many domains of science. Their application to atomistic simulations of matter is particularly widespread and impactful. This success is due largely to the existence of a well-developed and established physics-based modeling framework, ranging from first-principles electronic-structure calculations to molecular dynamics and statistical sampling, into which ML was integrated naturally to reshape long-standing trade-offs between accuracy, efficiency, and scale. Nevertheless, this integration raises both conceptual and practical challenges, from choosing between data-centric and physics-based modeling approaches to adapting established software stacks to modern hardware accelerators and ML libraries. As the field evolves rapidly, fueled in part by widespread enthusiasm but also by tangible impact, it seems appropriate to take a moment to consider the current state of the art and open challenges, and reflect on what can be done to better coordinate efforts across the community. With this goal in mind, several members of this community met in Lausanne in January 2026 at CECAM to discuss algorithms, models, software and hardware infrastructure, and the most promising scientific applications that have become possible thanks to the use of artificial intelligence in atomic-scale simulations. This strategic roadmap paper summarizes the outcomes of these discussions, suggesting some long-term goals, and some concrete actions, to establish a healthy, sustainable and impactful atomistic ML ecosystem.
While large language models (LLMs) have transformed AI agents into proficient executors of computational materials science, performing a hundred simulations does not make a researcher. What distinguishes research from routine execution is the progressive accumulation of knowledge - learning which approaches fail, recognizing patterns across systems, and applying understanding to new problems. However, the prevailing paradigm in AI-driven computational science treats each execution in isolation, largely discarding hard-won insights between runs. Here we present QMatSuite, an open-source platform closing this gap. Agents record findings with full provenance, retrieve knowledge before new calculations, and in dedicated reflection sessions correct erroneous findings and synthesize observations into cross-compound patterns. In benchmarks on a six-step quantum-mechanical simulation workflow, accumulated knowledge reduces reasoning overhead by 67% and improves accuracy from 47% to 3% deviation from literature - and when transferred to an unfamiliar material, achieves 1% deviation with zero pipeline failures.
Haonan Huang
Department of Physics, Princeton University, Princeton, NJ 08540, USA.
Protein dynamics underlie many biological functions, yet remain difficult to characterize due to the high computational cost of molecular dynamics simulations and the scarcity of dynamic structural data. This survey reviews recent advances in artificial intelligence for protein dynamics from three perspectives: learning from structural ensembles and trajectories, learning from physical energy signals, and learning to accelerate molecular simulations. We summarize representative methods for conformation ensemble generation, trajectory generation, Boltzmann generators, physics-aware adaptation, machine learning potentials, coarse-grained modeling, and collective variable discovery. We further discuss available datasets and key open challenges, such as scalability, thermodynamic consistency, kinetic fidelity, and integration with experimental constraints.
Haocheng Tang, Liang Shi, Ya-Shi Zhang +3
Mila – Quebec AI Institute · University of Pittsburgh · Work conducted during an internship at Mila +4