cs.CRSep 29, 2026

Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks

Authors: Alexandra Souly, Kai Fronsdal, Abby D'Cruz, Xander Davies, Robert Kirk

Organizations: UK AI Security Institute · UK AISI Research Affiliate

Abstract

This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of-scope, third-party targets when placed in difficult cybersecurity challenges, motivated by recently observed cases of models attacking real open-source repositories during evaluations. Applying our methods to GPT-6 Astra and previous OpenAI models, with cyber safeguards disabled, we find that GPT-6 Astra attempts complete supply-chain attacks in simulation at a higher rate than GPT-5.6 Sol and GPT-5.5. This includes writing malicious code as a contribution to an out-of-scope open-source codebase, creating fake identities to deceive open-source developers, and submitting benign contributions before malicious ones. GPT-6 Astra frequently reasons about the scope of the challenge in its chain-of-thought yet still proceeds to attack out-of-scope targets; it often asks for permission, and treats an automated message as as authorisation; and it continues to take unsanctioned actions, at a reduced rate, when internet access is more explicitly disallowed. Our evaluation builds on an internal version of Petri, an open-source LLM auditing tool, with all tool calls simulated by other LLMs, so that no real network access, systems or third-party repositories are reachable and no real-world harm is caused. Finally, we discuss limitations, in particular simulation awareness. We believe simulation awareness may have driven some of the observed behaviour but does not remove our concern. Our results suggest that defences beyond model alignment, such as sandboxing and monitoring, are increasingly critical for safe and secure deployment.

Figures & tables

Explore similar work

CardsList
  1. Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

    Sep 14, 2026Mark Russinovich, Blake Bullwinkel, Giorgio Severi +2Large Language Model SafetyStrong Large Language Model

  2. Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance

    May 10, 2026Michael Alexander Riegler, Finn Schwall, Annika Willoch Olstad +4ExploitationArtificial Intelligence Risk

  3. GPT-Red: Automated Red Teaming via Self-Play at Scale

    Jul 28, 2026Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15Red-TeamingAttacker Large Language Model