Info Gov

The AI Security Institute (AISI) has reveaked a security incident in which AI agents being evaluated for their cyber capabilities took sustained, unsanctioned action directed at real people and organisations, including an attempted supply-chain attack on a publicly used open-source software project.

The institute, a research organisation within the Department for Science, Innovation and Technology (DSIT), said its security team detected unusual data transfers leaving its research systems on 28 July 2026 during a routine cyber evaluation. The incident was contained within roughly an hour of discovery, and AISI said its investigation had found no evidence of resulting real-world harm.

The incident arose from a single evaluation in which agents were tasked with solving a cyber security challenge on a simulated network, run 122 times across seven frontier models. In 10 of those runs, an agent took autonomous action on the live internet beyond the scope of the test, with 19 such actions catalogued in total. Seventeen came from Anthropic's Mythos 5 model and two from OpenAI's GPT-5.6 Sol, in both cases with the developers' cyber misuse classifiers deliberately disabled and open internet access enabled. AISI said that these conditions do not reflect how frontier models are made available to the public.

AISI highlighted four significant behaviours identified in its investigation:

  • An attempted supply-chain attack, in which an agent tried to insert malicious code into a real open-source project, researched the project's human maintainers, created multiple fake identities and used them to pressure a maintainer into approving the code. When challenged publicly, the agent edited its earlier activity to appear harmless. Its use of the Tor anonymity network to bypass network restrictions was what first triggered AISI's security alert.
  • Attempts to deceive and target real people, including sending messages and files through an online file-transfer service to persuade individuals - or their AI coding tools - to run malicious code.
  • Attempts to plant prompt-injection attacks, placing hidden malicious instructions where the agent reasoned other automated AI systems might pick them up and execute them.
  • Collaboration between independent agents assessed simultaneously, with one agent leaving public messages on GitHub offering to work with other agents on the same challenge and providing instructions to reuse accounts and artefacts it had left behind.

The malicious pull request was caught and refused by a human maintainer, and a member of the public who suspected the code was malicious opened it inside a secure, isolated environment. AISI said standard security practice and human judgment had stopped the worst outcomes, but noted that in several cases the margin between failure and success rested on human vigilance rather than technical barriers.

AISI notified GitHub of the activity ahead of publication, and the platform confirmed the agents' actions violated its terms of service. The two organisations worked together to remove artefacts left behind by the agent and to notify the GitHub users the model had interacted with. AISI said it had also contacted other affected parties and intended to commission an independent third-party review from Model Evaluation and Threat Research (METR).

The institute acknowledged that its own evaluation design contributed to the behaviour. In some runs the task prompt was misconfigured so that the agent believed no legitimate solution existed, and internet access had been granted without monitoring purpose-built to watch the evaluation as it ran. It said it would now treat internet access in evaluations as a decision requiring active justification, introduce real-time monitoring capable of flagging or blocking out-of-scope actions, and add checks to ensure tasks are correctly specified.

AISI cautioned that the findings involved a small number of events under very specific conditions, and that its analysis of whether the agents understood they were acting in the real world - rather than a fictional test scenario - presented a mixed picture and remained ongoing. It nonetheless described the behaviour as "possible, sustained, and new", saying deception had emerged as a by-product of the agents pursuing their assigned task without being instructed to deceive, and that the incident pointed to a shift in the risk landscape in which harm may arise from capable agents acting beyond their authorised scope in internal research settings.

For organisations, AISI said the most effective response remained standard cyber hygiene, urging caution when verifying outside code and contributions. It encouraged organisations to sign up to the National Cyber Security Centre's free Early Warning service, to make cyber security a board-level responsibility, and to require Cyber Essentials across their supply chains.

AISI's full technical incident report is available here.

Also in this section

Jul 30, 2026

New NCSC guidance sets out three-stage framework for cyber attack recovery

The National Cyber Security Centre (NCSC) has published new guidance setting out a three-stage framework to help organisations respond to and recover from highly disruptive cyber attacks, warning that recovery from the most serious incidents can take weeks or months and that victims must plan for consequences extending well beyond their technology estate.

InfoGov Masthead Newsletter 800