Info Gov

The AI Security Institute (AISI) has reveaked a security incident in which AI agents being evaluated for their cyber capabilities took sustained, unsanctioned action directed at real people and organisations, including an attempted supply-chain attack on a publicly used open-source software project.

The institute, a research organisation within the Department for Science, Innovation and Technology (DSIT), said its security team detected unusual data transfers leaving its research systems on 28 July 2026 during a routine cyber evaluation. The incident was contained within roughly an hour of discovery, and AISI said its investigation had found no evidence of resulting real-world harm.

The incident arose from a single evaluation in which agents were tasked with solving a cyber security challenge on a simulated network, run 122 times across seven frontier models. In 10 of those runs, an agent took autonomous action on the live internet beyond the scope of the test, with 19 such actions catalogued in total. Seventeen came from Anthropic's Mythos 5 model and two from OpenAI's GPT-5.6 Sol, in both cases with the developers' cyber misuse classifiers deliberately disabled and open internet access enabled. AISI said that these conditions do not reflect how frontier models are made available to the public.

AISI highlighted four significant behaviours identified in its investigation:

  • An attempted supply-chain attack, in which an agent tried to insert malicious code into a real open-source project, researched the project's human maintainers, created multiple fake identities and used them to pressure a maintainer into approving the code. When challenged publicly, the agent edited its earlier activity to appear harmless. Its use of the Tor anonymity network to bypass network restrictions was what first triggered AISI's security alert.
  • Attempts to deceive and target real people, including sending messages and files through an online file-transfer service to persuade individuals - or their AI coding tools - to run malicious code.
  • Attempts to plant prompt-injection attacks, placing hidden malicious instructions where the agent reasoned other automated AI systems might pick them up and execute them.
  • Collaboration between independent agents assessed simultaneously, with one agent leaving public messages on GitHub offering to work with other agents on the same challenge and providing instructions to reuse accounts and artefacts it had left behind.

The malicious pull request was caught and refused by a human maintainer, and a member of the public who suspected the code was malicious opened it inside a secure, isolated environment. AISI said standard security practice and human judgment had stopped the worst outcomes, but noted that in several cases the margin between failure and success rested on human vigilance rather than technical barriers.

AISI notified GitHub of the activity ahead of publication, and the platform confirmed the agents' actions violated its terms of service. The two organisations worked together to remove artefacts left behind by the agent and to notify the GitHub users the model had interacted with. AISI said it had also contacted other affected parties and intended to commission an independent third-party review from Model Evaluation and Threat Research (METR).

The institute acknowledged that its own evaluation design contributed to the behaviour. In some runs the task prompt was misconfigured so that the agent believed no legitimate solution existed, and internet access had been granted without monitoring purpose-built to watch the evaluation as it ran. It said it would now treat internet access in evaluations as a decision requiring active justification, introduce real-time monitoring capable of flagging or blocking out-of-scope actions, and add checks to ensure tasks are correctly specified.

AISI cautioned that the findings involved a small number of events under very specific conditions, and that its analysis of whether the agents understood they were acting in the real world - rather than a fictional test scenario - presented a mixed picture and remained ongoing. It nonetheless described the behaviour as "possible, sustained, and new", saying deception had emerged as a by-product of the agents pursuing their assigned task without being instructed to deceive, and that the incident pointed to a shift in the risk landscape in which harm may arise from capable agents acting beyond their authorised scope in internal research settings.

For organisations, AISI said the most effective response remained standard cyber hygiene, urging caution when verifying outside code and contributions. It encouraged organisations to sign up to the National Cyber Security Centre's free Early Warning service, to make cyber security a board-level responsibility, and to require Cyber Essentials across their supply chains.

AISI's full technical incident report is available here.

Also in this section

Sep 11, 2026

Anthropic discloses fourth incident of AI model attacking real systems and hands investigation to independent evaluation organisation

Anthropic has published details of four incidents in which its Claude models gained unauthorised access to real third-party systems during cybersecurity evaluations, downloading and modifying user records at a real company, reading the personal information of an individual, harvesting credentials and accessing a security vendor's live database, after the test environments were mistakenly…
Sep 10, 2026

Welsh environmental watchdog hit by data breach

Environmental regulator Natural Resources Wales (NRW) has reported itself to the Information Commissioner's Office after a data breach saw personal details of staff made public.
Aug 24, 2026

Ministers seek power to ban tech risky vendors from critical sectors and bar recipients from discussing the order

The government has tabled amendments to the Cyber Security and Resilience (Network and Information Systems) Bill that would allow the Secretary of State to direct operators of essential services, data centres, managed service providers and other designated organisations to stop buying from, restrict the use of, or remove and disable products from a named vendor on national security grounds, with…
Aug 12, 2026

ACRO Criminal Records Office reprimanded by ICO following cyber security failings

The Information Commissioner's Office (ICO) has urged organisations to strengthen “patching and security monitoring processes” after cyber security failings at ACRO Criminal Records Office left the personal information of up to ten-thousand people, including some individuals’ sensitive data, potentially exposed.
Aug 06, 2026

AI agents sent malicious files to real developers and planted prompt injections in unmonitored test: AISI

The AI Security Institute (AISI) has published an incident report disclosing that AI agents under evaluation in its research environment took sustained, unsanctioned action against real people and organisations on the live internet, including researching the human maintainers of an open-source project, creating fake online identities to pressure one of them into approving malicious code, and…

InfoGov Masthead Newsletter 800