Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression
Organizations: Sandia National Laboratories (SNL), Albuquerque, NM, 87185, USA. · Sandia National Laboratories, Livermore, CA 94451, USA. · Oakridge National Laboratories (ORNL), Oak Ridge, TN, 37830, USA. · Pacific Northwest National Laboratory (PNNL), Richland, WA 99352, USA. · Los Alamos National Laboratory, Los Alamos, NM, 87545, USA. · Lawrence Livermore National Labs (LLNL), Livermore, CA 94550, USA. · Schmidt Sciences, New York, NY, 10011, USA. · RAND, Santa Monica, CA 90401, USA. · University of Tennessee, Knoxville, TN 37996, USA. · Software Engineering Institute (SEI), Carnegie Mellon University, Pittsburgh, PA, 15213, USA. · Georgetown University, Washington D.C., 20057, USA. · University of Montreal and LawZero, Montreal, H2S 3G9, Canada.
Abstract
This article presents a structured framework of behavioral indicators that may signal progression toward potentially catastrophic threats from artificial intelligence systems. We adopt a pragmatic approach, inspired by established methodologies in cybersecurity and national security. By establishing clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior, this framework enables researchers and policymakers to implement evidence-based monitoring protocols.