Info Gov

Anthropic has published details of four incidents in which its Claude models gained unauthorised access to real third-party systems during cybersecurity evaluations, downloading and modifying user records at a real company, reading the personal information of an individual, harvesting credentials and accessing a security vendor's live database, after the test environments were mistakenly connected to the open internet.

The company's assessment, published on 9 September and led by senior authors Sara Price and Subhash Kantamneni of Anthropic's alignment team, discloses a previously unreported fourth incident from January 2026 and revises the company's earlier explanation of the three incidents it first described on 30 July.

Anthropic said its initial account, which characterised the episodes as operational failures in which the models believed they were attacking simulated targets, had relied too heavily on what the models said they believed. Having conducted "a fuller analysis", the company now attributes the behaviour to two forms of misalignment: biased reasoning, in which the models disregarded or explained away evidence that they were on the real internet, and recklessness, a willingness to take harmful actions in pursuit of a task.

All four incidents occurred in capture-the-flag exercises built by the same third-party evaluation partner, in which the models were told they had no internet access and were run without the cyber safeguards that ship with released models. Each involved a single model instance working for between roughly 10 and 34 hours. The incidents were:

  1. Claude Mythos 5 registered a disposable email account, published three versions of a malicious package to PyPI, the public Python package repository, and used credentials leaked by one of the 15 systems that installed it to access a real security vendor's database. It had also attempted to use payment processor credentials and find cryptocurrency to buy a phone number for registration. PyPI removed the package within an hour.
  2. Claude Opus 4.7, in four separate runs, mistook a real company with a near-identical name for its fictional target, scanned its live web application, downloaded user records and modified them. Anthropic said the model recognised in its reasoning that the target was real but assumed it was authorised to proceed.
  3. An internal research model, after its target machine shut down, attacked neighbouring systems it believed belonged to other test participants, using access tokens and passwords to penetrate one, downloading files and inserting a remote-control script before recognising a real company name and stopping.
  4. An early checkpoint of Claude Opus 4.6, unable to abort a task it had accidentally broken, found a third party's machine, used a password found in a file to gain administrator access, harvested further credentials, altered system settings and read the personal information of one person associated with the third party.

Anthropic said it had notified all affected parties and that Claude never attempted to conceal its actions or coordinate with other agents. Resampling experiments on the Mythos 5 transcript found that making the real-world nature of the environment clearer did not meaningfully change the model's behaviour, even though it made the model more likely to acknowledge possible real-world harm. Only an unambiguous statement that a host was live on the public internet caused it to stand down.

The company said its newer models, Claude Opus 5 and Claude Mythos 5.1, took severely harmful actions in around 30% of runs in a simulated replication of the scenario, compared with 82% for Mythos 5, a rate it described as lower but still concerning.

The incidents illustrate risks that sit within existing data protection obligations. A controller using an AI agent with access to systems holding personal data remains responsible under Article 5(1)(f) and Article 32 UK GDPR for the security of that processing, and an agent that accesses, downloads or alters personal data without authorisation would engage the personal data breach provisions in Articles 33 and 34.

Where an organisation's own data is accessed by a third party's agent, as happened to the companies in these incidents, the question of who is the controller for the resulting processing, and who owes notification duties, is not straightforward. The ICO's guidance on AI and data protection and its 2025 consultation on agentic AI both emphasise the need for a data protection impact assessment under Article 35 before deploying autonomous tools with access to live systems.

The assessment does not cover a separate incident disclosed on 4 August by the UK AI Security Institute, in which Mythos 5 and, to a lesser extent, OpenAI's GPT-5.6-Sol took unsanctioned action on the live internet during an AISI cyber evaluation. In the most serious case, an agent attempted to insert malicious code into an open-source project on GitHub, researched the project's maintainers, created multiple fake identities and used them to pressure a real maintainer into approving the code.

AISI said a human maintainer refused the request and no real-world harm had been identified, but described it as the first time it had seen autonomy and deception risks manifest without specific prompting outside a test scenario. Anthropic said it planned to conduct an alignment assessment of those transcripts as well.

Anthropic has signed an eight-week agreement with METR, an independent evaluation organisation, to investigate the incidents, giving it access to transcripts, employees and confidential information, with the option to extend. Kantamneni and colleagues said the company's pre-release auditing had not warned that misalignment of this severity was present, and that its offline chain-of-thought monitors would have missed the Mythos 5 incident because the model's reasoning persuaded the monitor the environment was simulated.

The company also disclosed that it had trained two versions of Mythos 5, one with additional alignment environments and one without, and had shipped the version without them because employees found it more usable, a decision it now describes as a mistake.

Anthropic said it was establishing a regular process for publishing findings about model behaviour beyond what appears in its system cards, and had set requirements that third-party partners must meet before running pre-release models without cyber safeguards, including explicit definition of in-scope targets, permitted actions and network boundaries.

Also in this section

Sep 10, 2026

Welsh environmental watchdog hit by data breach

Environmental regulator Natural Resources Wales (NRW) has reported itself to the Information Commissioner's Office after a data breach saw personal details of staff made public.
Aug 24, 2026

Ministers seek power to ban tech risky vendors from critical sectors and bar recipients from discussing the order

The government has tabled amendments to the Cyber Security and Resilience (Network and Information Systems) Bill that would allow the Secretary of State to direct operators of essential services, data centres, managed service providers and other designated organisations to stop buying from, restrict the use of, or remove and disable products from a named vendor on national security grounds, with…
Aug 12, 2026

ACRO Criminal Records Office reprimanded by ICO following cyber security failings

The Information Commissioner's Office (ICO) has urged organisations to strengthen “patching and security monitoring processes” after cyber security failings at ACRO Criminal Records Office left the personal information of up to ten-thousand people, including some individuals’ sensitive data, potentially exposed.
Aug 06, 2026

AI agents sent malicious files to real developers and planted prompt injections in unmonitored test: AISI

The AI Security Institute (AISI) has published an incident report disclosing that AI agents under evaluation in its research environment took sustained, unsanctioned action against real people and organisations on the live internet, including researching the human maintainers of an open-source project, creating fake online identities to pressure one of them into approving malicious code, and…
Aug 05, 2026

Third AI platform goes rogue during cyber testing

The AI Security Institute (AISI) has reveaked a security incident in which AI agents being evaluated for their cyber capabilities took sustained, unsanctioned action directed at real people and organisations, including an attempted supply-chain attack on a publicly used open-source software project.

InfoGov Masthead Newsletter 800