Info Gov

A pilot programme using frontier AI models to scan public-sector code repositories has identified 407 security findings across nine government organisations, including a critical vulnerability that could have allowed an external attacker to execute arbitrary code on a key digital service, the Department for Science, Innovation and Technology (DSIT) and the National Cyber Security Centre (NCSC) have revealed.


The Government Cyber Coordination Centre (GC3) - a joint NCSC/DSIT body - ran a series of weekly hackathons over a month in which teams used frontier AI systems to scan open-source government code for previously unidentified weaknesses. All critical vulnerabilities identified during the exercise have now been remediated, and the departments involved said no evidence of exploitation had been found for any of the findings.

Teams were given access to frontier models, including Claude Mythos and GPT-5.5, and allowed to design their own tooling rather than follow a mandated methodology. Across the nine participating organisations, the pilot generated 407 findings in total, spanning authentication bypass, data exposure and remote code execution risks. Some had already been identified and mitigated through existing compensating controls; others were previously unknown. The total cost of the exercise, measured in AI model token usage, was reported as £13,000.

The department said AI models were able to trace vulnerabilities across service boundaries, connecting business logic with technical detail in ways traditional static-analysis scanners cannot, though all findings were subject to human validation before entering departmental remediation pipelines.

The most significant finding affected legacy GitHub Actions workflows in a repository supporting a major government digital service. The vulnerability allowed an external user to trigger a chain of automated workflows simply by posting a specially crafted comment on an open pull request. This bypassed the safeguards normally applied to contributions from unverified users, because the trigger was the comment itself rather than the pull request. The department said this level of access could have supported wider repository compromise, including manipulating pull requests, approving workflow activity and altering trusted contributor permissions.

The department said the exercise demonstrated that the architecture surrounding an AI model mattered more than the choice of model itself, with AI Security Institute research cited as showing that near-frontier and frontier models perform comparably when given the right task structure. Effective triage was identified as essential, given that AI agents generate candidate findings far faster than human reviewers can validate them.

GC3 said a second phase of the pilot would extend the approach to additional departments and models, and broaden the scope from public code repositories to closed-source government IT estates, as part of implementation of the Government Cyber Action Plan.

Further details of the exercise can be found at When AI Leaves the Lab: Testing Frontier Models in Government Cyber Defence.

Also in this section

Aug 12, 2026

ACRO Criminal Records Office reprimanded by ICO following cyber security failings

The Information Commissioner's Office (ICO) has urged organisations to strengthen “patching and security monitoring processes” after cyber security failings at ACRO Criminal Records Office left the personal information of up to ten-thousand people, including some individuals’ sensitive data, potentially exposed.
Aug 05, 2026

Third AI platform goes rogue during cyber testing

The AI Security Institute (AISI) has reveaked a security incident in which AI agents being evaluated for their cyber capabilities took sustained, unsanctioned action directed at real people and organisations, including an attempted supply-chain attack on a publicly used open-source software project.
Jul 30, 2026

New NCSC guidance sets out three-stage framework for cyber attack recovery

The National Cyber Security Centre (NCSC) has published new guidance setting out a three-stage framework to help organisations respond to and recover from highly disruptive cyber attacks, warning that recovery from the most serious incidents can take weeks or months and that victims must plan for consequences extending well beyond their technology estate.

InfoGov Masthead Newsletter 800