Info Gov

OpenAI has admitted that a combination of its AI models autonomously broke out of a testing environment and hacked into Hugging Face, the open-source model and dataset platform, in what the company has called an "unprecedented cyber incident, involving state-of-the-art cyber capabilities".

In a statement published on 21 July, OpenAI said the incident was driven by a combination of its models - including the publicly available GPT-5.6 Sol and "an even more capable pre-release model", all with reduced cyber refusals for evaluation purposes - while being internally tested on a benchmark of cyber capabilities. The company ran the evaluation without the production classifiers normally used to prevent models from pursuing high-risk cyber activity, in order to estimate their maximal capability.

According to OpenAI, the models were "hyperfocused" on solving the evaluation, known as ExploitGym, and went to "extreme lengths" to achieve the narrow testing goal. Operating inside a sandboxed environment, they spent a substantial amount of inference compute finding a way onto the open internet, identifying and exploiting a zero-day vulnerability in a package registry cache proxy to do so. The models then performed privilege escalation and lateral movement until they reached a node with internet access, inferred that Hugging Face might host solutions for the benchmark, and chained together stolen credentials and further vulnerabilities to find a remote code execution path on Hugging Face's servers.

Hugging Face, which disclosed the intrusion separately on 16 July before the source was known, said it had identified unauthorised access to a limited set of internal datasets and to several credentials used by its services, and was still completing its assessment of whether any partner or customer data was affected. The company said it had found no evidence of tampering with public, user-facing models, datasets or Spaces, and that its software supply chain was verified clean. It has reported the incident to law enforcement and engaged outside forensic specialists, and advised users to rotate access tokens and review recent account activity as a precaution.

Hugging Face said a malicious dataset abused two code-execution paths in its dataset processing pipeline - a remote-code dataset loader and a template injection in a dataset configuration - to run code on a processing worker, from where the attacker escalated to node-level access, harvested cloud and cluster credentials, and moved laterally over a weekend. The campaign executed many thousands of individual actions across a swarm of short-lived sandboxes, and Hugging Face later reconstructed more than 17,000 recorded events to establish what had happened. 

When its responders first tried to use frontier models behind commercial APIs to examine the attacker's commands, exploit payloads and command-and-control artefacts, the requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. The team ran the forensic analysis instead on GLM 5.2, an open-weight model, on its own infrastructure which had the added benefit that no attacker data, and none of the credentials it referenced, left the company's environment.

Hugging Face said that the keeping of attacker data within a controlled environment during forensic analysis underlines the importance data minimisation and security principles that govern how organisations handle personal data during incident response.

The attacker's agents were bound by no usage policy, while the defenders' own forensic work was blocked by the guardrails of the hosted models they first tried.

OpenAI's security team discovered the anomalous activity internally, while Hugging Face's team detected and contained it on their infrastructure before the two connected. OpenAI said it has disclosed the zero-day vulnerability to the vendor, is implementing stricter infrastructure controls at the cost of research velocity, and has brought Hugging Face into its trusted access programme for cyber defence.

Hugging Face CEO Clem Delangue said the incident, "possibly the first of its kind", proved that AI safety would not be solved by any single company working in secret but "in the open, collaboratively, with broad access to AI for every defender, everywhere".

Also in this section

Sep 11, 2026

Anthropic discloses fourth incident of AI model attacking real systems and hands investigation to independent evaluation organisation

Anthropic has published details of four incidents in which its Claude models gained unauthorised access to real third-party systems during cybersecurity evaluations, downloading and modifying user records at a real company, reading the personal information of an individual, harvesting credentials and accessing a security vendor's live database, after the test environments were mistakenly…
Sep 10, 2026

Welsh environmental watchdog hit by data breach

Environmental regulator Natural Resources Wales (NRW) has reported itself to the Information Commissioner's Office after a data breach saw personal details of staff made public.
Aug 24, 2026

Ministers seek power to ban tech risky vendors from critical sectors and bar recipients from discussing the order

The government has tabled amendments to the Cyber Security and Resilience (Network and Information Systems) Bill that would allow the Secretary of State to direct operators of essential services, data centres, managed service providers and other designated organisations to stop buying from, restrict the use of, or remove and disable products from a named vendor on national security grounds, with…
Aug 12, 2026

ACRO Criminal Records Office reprimanded by ICO following cyber security failings

The Information Commissioner's Office (ICO) has urged organisations to strengthen “patching and security monitoring processes” after cyber security failings at ACRO Criminal Records Office left the personal information of up to ten-thousand people, including some individuals’ sensitive data, potentially exposed.
Aug 06, 2026

AI agents sent malicious files to real developers and planted prompt injections in unmonitored test: AISI

The AI Security Institute (AISI) has published an incident report disclosing that AI agents under evaluation in its research environment took sustained, unsanctioned action against real people and organisations on the live internet, including researching the human maintainers of an open-source project, creating fake online identities to pressure one of them into approving malicious code, and…
Aug 05, 2026

Third AI platform goes rogue during cyber testing

The AI Security Institute (AISI) has reveaked a security incident in which AI agents being evaluated for their cyber capabilities took sustained, unsanctioned action directed at real people and organisations, including an attempted supply-chain attack on a publicly used open-source software project.

InfoGov Masthead Newsletter 800