Info Gov

A pilot programme using frontier AI models to scan public-sector code repositories has identified 407 security findings across nine government organisations, including a critical vulnerability that could have allowed an external attacker to execute arbitrary code on a key digital service, the Department for Science, Innovation and Technology (DSIT) and the National Cyber Security Centre (NCSC) have revealed.


The Government Cyber Coordination Centre (GC3) - a joint NCSC/DSIT body - ran a series of weekly hackathons over a month in which teams used frontier AI systems to scan open-source government code for previously unidentified weaknesses. All critical vulnerabilities identified during the exercise have now been remediated, and the departments involved said no evidence of exploitation had been found for any of the findings.

Teams were given access to frontier models, including Claude Mythos and GPT-5.5, and allowed to design their own tooling rather than follow a mandated methodology. Across the nine participating organisations, the pilot generated 407 findings in total, spanning authentication bypass, data exposure and remote code execution risks. Some had already been identified and mitigated through existing compensating controls; others were previously unknown. The total cost of the exercise, measured in AI model token usage, was reported as £13,000.

The department said AI models were able to trace vulnerabilities across service boundaries, connecting business logic with technical detail in ways traditional static-analysis scanners cannot, though all findings were subject to human validation before entering departmental remediation pipelines.

The most significant finding affected legacy GitHub Actions workflows in a repository supporting a major government digital service. The vulnerability allowed an external user to trigger a chain of automated workflows simply by posting a specially crafted comment on an open pull request. This bypassed the safeguards normally applied to contributions from unverified users, because the trigger was the comment itself rather than the pull request. The department said this level of access could have supported wider repository compromise, including manipulating pull requests, approving workflow activity and altering trusted contributor permissions.

The department said the exercise demonstrated that the architecture surrounding an AI model mattered more than the choice of model itself, with AI Security Institute research cited as showing that near-frontier and frontier models perform comparably when given the right task structure. Effective triage was identified as essential, given that AI agents generate candidate findings far faster than human reviewers can validate them.

GC3 said a second phase of the pilot would extend the approach to additional departments and models, and broaden the scope from public code repositories to closed-source government IT estates, as part of implementation of the Government Cyber Action Plan.

Further details of the exercise can be found at When AI Leaves the Lab: Testing Frontier Models in Government Cyber Defence.

Also in this section

Sep 11, 2026

Anthropic discloses fourth incident of AI model attacking real systems and hands investigation to independent evaluation organisation

Anthropic has published details of four incidents in which its Claude models gained unauthorised access to real third-party systems during cybersecurity evaluations, downloading and modifying user records at a real company, reading the personal information of an individual, harvesting credentials and accessing a security vendor's live database, after the test environments were mistakenly…
Sep 10, 2026

Welsh environmental watchdog hit by data breach

Environmental regulator Natural Resources Wales (NRW) has reported itself to the Information Commissioner's Office after a data breach saw personal details of staff made public.
Aug 24, 2026

Ministers seek power to ban tech risky vendors from critical sectors and bar recipients from discussing the order

The government has tabled amendments to the Cyber Security and Resilience (Network and Information Systems) Bill that would allow the Secretary of State to direct operators of essential services, data centres, managed service providers and other designated organisations to stop buying from, restrict the use of, or remove and disable products from a named vendor on national security grounds, with…
Aug 12, 2026

ACRO Criminal Records Office reprimanded by ICO following cyber security failings

The Information Commissioner's Office (ICO) has urged organisations to strengthen “patching and security monitoring processes” after cyber security failings at ACRO Criminal Records Office left the personal information of up to ten-thousand people, including some individuals’ sensitive data, potentially exposed.
Aug 06, 2026

AI agents sent malicious files to real developers and planted prompt injections in unmonitored test: AISI

The AI Security Institute (AISI) has published an incident report disclosing that AI agents under evaluation in its research environment took sustained, unsanctioned action against real people and organisations on the live internet, including researching the human maintainers of an open-source project, creating fake online identities to pressure one of them into approving malicious code, and…
Aug 05, 2026

Third AI platform goes rogue during cyber testing

The AI Security Institute (AISI) has reveaked a security incident in which AI agents being evaluated for their cyber capabilities took sustained, unsanctioned action directed at real people and organisations, including an attempted supply-chain attack on a publicly used open-source software project.

InfoGov Masthead Newsletter 800