cs.CYSep 2, 2026

Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression

Authors: T. BauerW. P. KegelmeyerE. BegoliA. SadovnikT. EmersonC. CorleyN. GenerousJ. Moore+10 more

Organizations: Sandia National Laboratories (SNL), Albuquerque, NM, 87185, USA. · Sandia National Laboratories, Livermore, CA 94451, USA. · Oakridge National Laboratories (ORNL), Oak Ridge, TN, 37830, USA. · Pacific Northwest National Laboratory (PNNL), Richland, WA 99352, USA. · Los Alamos National Laboratory, Los Alamos, NM, 87545, USA. · Lawrence Livermore National Labs (LLNL), Livermore, CA 94550, USA. · Schmidt Sciences, New York, NY, 10011, USA. · RAND, Santa Monica, CA 90401, USA. · University of Tennessee, Knoxville, TN 37996, USA. · Software Engineering Institute (SEI), Carnegie Mellon University, Pittsburgh, PA, 15213, USA. · Georgetown University, Washington D.C., 20057, USA. · University of Montreal and LawZero, Montreal, H2S 3G9, Canada.

Abstract

This article presents a structured framework of behavioral indicators that may signal progression toward potentially catastrophic threats from artificial intelligence systems. We adopt a pragmatic approach, inspired by established methodologies in cybersecurity and national security. By establishing clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior, this framework enables researchers and policymakers to implement evidence-based monitoring protocols.

Explore similar work

Jul 17, 2026cs.AI

Harmonizing AI Safety Thresholds

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.
Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza +1
May 27, 2026cs.AI

Measuring Progress Toward AGI: A Cognitive Framework

Despite widespread discussion of AGI, there is no clear framework for measuring progress toward it. This ambiguity fuels subjective claims, makes it difficult to track progress, and risks hindering responsible governance. As a starting point to address this gap, we present a framework for understanding system capabilities in relation to human cognitive abilities. Drawing from decades of research in psychology, neuroscience, and cognitive science, we introduce a Cognitive Taxonomy that deconstructs general intelligence into 10 key cognitive faculties. We then propose a rigorous evaluation protocol in which a system's performance is measured across a suite of targeted, held-out cognitive tasks, generating a 'cognitive profile' that can be used to understand a system's strengths and weaknesses. We hope this framework will provide a practical roadmap and an initial step toward more rigorous, empirical evaluation of AGI.
Ryan Burnell, Yumeya Yamamori, Orhan Firat +10
Jul 2, 2026cs.CY

Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond

The society and emerging risk-based regulatory frameworks for AI underscore the need for rigorous risk assessment to ensure safe and reliable AI systems. In response to this imperative, this paper presents an overview of AI risk assessment (identification and analysis) and management methodologies. It begins by reviewing the worldwide regulatory landscape that drives the need for systematic AI risk assessment. Then we characterize the spectrum of AI-related risks identified in the literature, from technical failures to ethical and social impacts. Subsequently, it reviews key risk assessment methodologies proposed for AI systems, focusing on general frameworks. The paper highlights best practices and illuminates methodological gaps, highlighting areas for further research on AI risk assessment.
Javier Irigoyen, Roberto Daza, Aythami Morales +5