Info Gov

OpenAI has admitted that a combination of its AI models autonomously broke out of a testing environment and hacked into Hugging Face, the open-source model and dataset platform, in what the company has called an "unprecedented cyber incident, involving state-of-the-art cyber capabilities".

In a statement published on 21 July, OpenAI said the incident was driven by a combination of its models - including the publicly available GPT-5.6 Sol and "an even more capable pre-release model", all with reduced cyber refusals for evaluation purposes - while being internally tested on a benchmark of cyber capabilities. The company ran the evaluation without the production classifiers normally used to prevent models from pursuing high-risk cyber activity, in order to estimate their maximal capability.

According to OpenAI, the models were "hyperfocused" on solving the evaluation, known as ExploitGym, and went to "extreme lengths" to achieve the narrow testing goal. Operating inside a sandboxed environment, they spent a substantial amount of inference compute finding a way onto the open internet, identifying and exploiting a zero-day vulnerability in a package registry cache proxy to do so. The models then performed privilege escalation and lateral movement until they reached a node with internet access, inferred that Hugging Face might host solutions for the benchmark, and chained together stolen credentials and further vulnerabilities to find a remote code execution path on Hugging Face's servers.

Hugging Face, which disclosed the intrusion separately on 16 July before the source was known, said it had identified unauthorised access to a limited set of internal datasets and to several credentials used by its services, and was still completing its assessment of whether any partner or customer data was affected. The company said it had found no evidence of tampering with public, user-facing models, datasets or Spaces, and that its software supply chain was verified clean. It has reported the incident to law enforcement and engaged outside forensic specialists, and advised users to rotate access tokens and review recent account activity as a precaution.

Hugging Face said a malicious dataset abused two code-execution paths in its dataset processing pipeline - a remote-code dataset loader and a template injection in a dataset configuration - to run code on a processing worker, from where the attacker escalated to node-level access, harvested cloud and cluster credentials, and moved laterally over a weekend. The campaign executed many thousands of individual actions across a swarm of short-lived sandboxes, and Hugging Face later reconstructed more than 17,000 recorded events to establish what had happened. 

When its responders first tried to use frontier models behind commercial APIs to examine the attacker's commands, exploit payloads and command-and-control artefacts, the requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. The team ran the forensic analysis instead on GLM 5.2, an open-weight model, on its own infrastructure which had the added benefit that no attacker data, and none of the credentials it referenced, left the company's environment.

Hugging Face said that the keeping of attacker data within a controlled environment during forensic analysis underlines the importance data minimisation and security principles that govern how organisations handle personal data during incident response.

The attacker's agents were bound by no usage policy, while the defenders' own forensic work was blocked by the guardrails of the hosted models they first tried.

OpenAI's security team discovered the anomalous activity internally, while Hugging Face's team detected and contained it on their infrastructure before the two connected. OpenAI said it has disclosed the zero-day vulnerability to the vendor, is implementing stricter infrastructure controls at the cost of research velocity, and has brought Hugging Face into its trusted access programme for cyber defence.

Hugging Face CEO Clem Delangue said the incident, "possibly the first of its kind", proved that AI safety would not be solved by any single company working in secret but "in the open, collaboratively, with broad access to AI for every defender, everywhere".

Also in this section

Jul 16, 2026

Scattered Spider pair jailed over TfL hack

Two members of the Scattered Spider hacking collective have each been jailed for five and a half years at Woolwich Crown Court on 16 July over a 2024 cyber attack on Transport for London that compromised the personal data of millions of customers and cost the transport authority £39 million.
Jul 15, 2026

Government introduces mandatory data breach reporting protocol

The Cabinet Office has published a Model Action Plan setting out a single cross-government framework that all departments and arms-length bodies must follow when responding to significant personal data breaches, introducing mandatory central reporting of such incidents for the first time.
Jun 29, 2026

"Five eyes" warn on accelerating cyber security threat from AI

The National Cyber Security has called on organisational leaders to treat cyber resilience as a core business and governance responsibility rather than a technical matter in the light of a report by the Five Eyes intelligence alliance on the "fundamental" impact of AI on the speed and scale of cyber threats.
Jun 15, 2026

Government AI hackathons uncover 407 vulnerabilities across nine departments, including critical remote code execution flaw

A pilot programme using frontier AI models to scan public-sector code repositories has identified 407 security findings across nine government organisations, including a critical vulnerability that could have allowed an external attacker to execute arbitrary code on a key digital service, the Department for Science, Innovation and Technology (DSIT) and the National Cyber Security Centre (NCSC)…

InfoGov Masthead Newsletter 800