Authors: Sandy Adhitia Ekahana, Aalok Tiwari, Pratik Saud, Aaron Bostwick, Chris Jozwiak, Eli Rotenberg, Jyoti Katoch
Organizations: Department of Physics, Carnegie Mellon University, Pittsburgh, 15213, PA, USA. · MAESTRO, Advanced Light Source, Lawrence Berkeley National Laboratory, Berkeley, CA, USA.
Artificial intelligence (AI) is becoming an increasingly useful tool across the experimental sciences, including angle-resolved photoemission spectroscopy (ARPES), which routinely produces large, multidimensional datasets of electronic structure. Recent advances in AI and machine learning (ML) have opened new opportunities across the entire ARPES workflow, from automated sample preparation and real-time data acquisition to post-experiment data analysis and comparison with theoretical calculations. Despite this progress, a comprehensive review of ML applications, their capabilities, and reliability across the different stages of ARPES workflow is still lacking. In this review, we first introduce ML methods that are most relevant to experimentalists working in condensed matter physics and materials science. We then follow the ARPES workflow, reviewing existing ML applications at each step and discussing their advantages, limitations and potential for future development. We also examine the current ARPES data landscape, where several open databases are available but remain relatively small and fragmented compared with large, shared datasets such as ImageNet. Given these limitations, we suggest that the community focus on sharing pretrained models that can be further trained, adapted to specific tasks, and redistributed, while working toward a larger and standardized open ARPES dataset repository. Finally, we discuss our perspectives on the future of AI within the ARPES workflow using a six-level framework of laboratory automation, highlighting the opportunities and challenges in moving toward a fully autonomous, self-driving ARPES laboratory.
Figures & tables
Figure 1: Schematic of ARPES Workflow The typical ARPES workflow progresses from left to right, from sample growth to analysis via measurement in an ultrahigh vacuum (UHV) chamber. Solid straight arrows indicate the nominal forward path, while solid curved arrows mark active learning–based iterations designed to acquire a rapid and complete dataset during a small experimental window. It represents an iterative closed loop between measuring, analyzing, and deciding the next step. Combined with other complementary studies, this reveals the underlying electronic properties of quantum materials
Figure 2: A practical AI terminology for experimental physicists. Machine learning splits by architecture into traditional compact models and deep learning networks. The four learning protocols represent a separate, independent axis of training models. The asterisk marks reinforcement learning, which in current practice is typically implemented with deep networks rather than as an architecture-independent protocol. The figure is a visual guide rather than a comprehensive glossary. Detailed examples appear in the text.
Figure 3: Automated sample preparation and preliminary characterizations. ( a ) The five pre experiment stages for ARPES experiment, shaded by the extent of automation. ( b ) An autonomous inorganic synthesis laboratory (A-lab) for air-stable targets, in which an ML algorithm proposes recipes, while robotic synthesis and learned phase identification form a self-revising loop (adapted from Szymanski et al. (2023) ). ( c ) Robotic four dimensional pixel assembly, from wafer scale growth and patterning to automated stacking, with control over layer number, composition, lateral position and twist angle. Scale bar: 50 μ m (adapted from Mannix et al. (2022) ).
Figure 4: Route to experimental automation ( a ) KMGP-driven data collection identifying distinct major phases on the sample, colored by cluster label Melton et al. (2020) . ( b ) The GPR-guided acquisition loop showing the autonomous data acquisition process. The colored boxes represent asynchronously running processes, while the red arrows indicate real-time, ML-based communication between the instrument and Python scripts Ágústsson et al. (2024b) . ( c ) GPR score, mean uncertainty ( ΔPmean ), and gain versus number of scans for Bi 2 Te 3 spin-ARPES, illustrating a quantitative stopping criterion for the acquisition loop Iwasawa et al. (2024) . ( d ) A synthetic trained CNN scores the measured spectral quality across a Bi 2 Se 3 grid scan. The predicted map is compared with a conventional benchmark score and decides where to measure Na et al. (2025) . Combined, these steps form a closed-loop experimental workflow: locating sample variations, selecting target regions, measuring until an objective stopping criterion is reached, evaluating spectral quality, and iterating to optimize spatial and electronic coverage.
Figure 5: Denoising network training sequence on Poisson subsampled ARPES data ( a ) Generation of noisy training data from high count Au(111) surface states spectra. From original high count data x , low count noisy copies x′ are generated via Poisson subsampling, simulating rapid acquisition. The denoising process is denoted f(x′) , while the generation of the noisy training set is denoted f−1(x) . ( b ) Training loop for the denoising neural network, where the loss function drives parameter optimization. The loss function L(x,f(x′)) combines mean absolute error (MAE) and multi-scale structural similarity index measure (MS-SSIM). The weighted formulation ensures that reconstructed spectra retain sharp spectral features essential for lineshape and second-derivative analysis, enabling quantitative comparison between short and long acquisitions on the same ARPES cuts (adapted from Kim et al. (2021) ).
Figure 6: Unsupervised clustering for spatial domain identification Panel (a) shows raw spatially resolved spectral intensity. Panel (b) applies k -means clustering with three clusters (1,2,3), assigning each point exclusively to one cluster. Panel (c) applies fuzzy c-means soft clustering, where each point receives probabilistic membership weights across clusters; the background shading indicates low membership confidence or non-sample regions, revealing sharp domain boundaries (adapted from Iwasawa et al. (2022) )
Figure 7: Band extraction of NbIrTe 4 using DFT, conventional analysis, and machine learning. ( a ) Density functional theory calculation of the band structure of NbIrTe 4 . The bands crossings are marked with red arrows. ( b ) Corresponding experimental ARPES data showing the same band structure. ( c–e ) Results from three band extraction methods applied to the experimental data: maximum curvature tracing ( c ), minimum gradient method ( d ), and CNN ( e ). The CNN approach recovers band positions comparable to conventional methods while reducing manual intervention and improving robustness on weak or overlapping features (adapted from Ref. Peng et al. (2020) ).
Figure 8: Flow chart of an ML regression procedure for extracting hidden self energies in Bi2212. A Boltzmann machine represents the normal ( Σnor(k,ω) ) and anomalous self energies ( Σano(k,ω) ), and is optimized against the measured spectral weight A(k,ω) through nested training and test loops until convergence. Adapted from Ref. Yamaji et al. (2021) ; Yamaji et al. (2023)
Figure 9: DeepH: An example of a theoretical framework where a DFT Hamiltonian surrogate predicts the Hamiltonian from the atomic structure. Adapted from Ref. Li et al. (2022) . ( a ) The DFT Hamiltonian H^DFT({R}) , which is a function of atomic coordinates, can be obtained by self-consistent field (SCF) calculations or learned by a neural network for efficient ab initio electronic structure calculations, yielding its physical properties ( b ) The network exploits the nearsightedness principle of electronic matter making the surrogate size transferable. The Hamiltonian matrix elements in the localized basis are nonzero only between neighboring atoms (within a cutoff RC ) and are influenced only by the local neighborhood (within RN ).
Figure 10: The six automation levels of ARPES , adapted from Ref. Le Houx (2026) . Level 0 (Manual) represents fully manual ARPES measurements with only the minimal necessary motorization. Level 1 (Scripted) represents the stage at which full motorization is realized and predefined scripts control all motion. Level 2 (Reflexive) represents the stage at which a subset of decisions is transferred to pre defined scalar feedback logic. Level 3 (Heuristic) represents the stage at which decisions are made heuristically through physics-informed inference. Level 4 (Supervisory) represents the stage at which an expert supervises an agentic bot controlling the experiment, while retaining key decision points and the right to veto the agent’s actions. Level 5 (Autonomous) represents the stage at which the human engages with the agent at the level of scientific discussion, with minimal direct decision-making. Representative example is an agent that can propose and independently execute a new experiment based on prior discussion.