Info Gov

The National Cyber Security Centre has published interim practical advice on managing the cyber security risks of agentic AI, setting out seven safeguards for organisations deploying AI systems that can plan, use tools and take actions with significant autonomy.

The guidance, published as a blog in the NCSC website on 20 August, is aimed at system designers and operators building environments in which AI agents operate with limited human intervention, and at those concerned about agents carrying out unintended actions because of the instructions they receive, the tools available to them or the systems they can access.

The NCSC says it has been researching and experimenting in the area for some time and is working with partners on formal guidance that will build on and ultimately supersede the blog.

The blog follows what the NCSC describes as several recent incidents involving AI models and agentic systems carrying out unsanctioned or unintended activity, which it says highlight the need for organisations to consider carefully how the technology is deployed, constrained, observed and responded to.

The seven key considerations are:

1. Identify what could go wrong, documenting what is in and out of scope, setting red lines and carrying out threat modelling before deployment, remembering that an agent may interpret goals literally or in unexpected ways.
2. Prompt carefully, making explicit what the agent should and should not do and when it should stop for human approval, while not relying on prompting alone.
3. Set the right level of oversight, choosing between human-in-the-loop, human-on-the-loop and human-out-of-the-loop models, with named individuals responsible for agentic activity and real-time monitoring where consequences would be significant.
4. Control the agent's environment with a robust sandbox, restricting network access on a deny-by-default basis, isolating compute, and limiting the credentials available to the agent, which should have its own unique identity distinct from human users.
5. Log, audit and monitor agentic activity as part of security operations, retaining chain-of-thought traces and transcripts alongside environment logs, protecting them from modification and treating agent activity as user activity within 24/7 monitoring.
6. Make AI activity easy to attribute, so that third parties can identify traffic as originating from the organisation.
7. Maintain the ability to pull the plug, with controls that can halt agent activity, restrict network access and interrupt communications with the model inference infrastructure immediately.

The NCSC says the recommendations should be applied proportionately to the level of autonomy each system is intended to have. Organisations should first assess how much autonomy a use case requires and the risk they are prepared to tolerate, then understand the safeguards built into the model, inference service and harness they are using.

Those built-in controls, it warns, may be bypassed, may not provide adequate protection in higher-risk environments and should not be treated as sufficient on their own. Where the consequences of failure exceed tolerance, additional safeguards such as classifiers and deterministic provers should be layered on top.

For data controllers, agentic deployments that process personal data fall within the requirement in Article 32 of the UK GDPR, and section 66 of the Data Protection Act 2018 for law enforcement processing, to implement technical and organisational measures appropriate to the risk, taking into account the state of the art.

The guidance includes two four-level maturity models. For network access, level one is unrestricted access, level two an allowlist of approved domains, level three access restricted to the model's API only, and level four no external network access with the model hosted locally within the sandbox. For compute isolation, level one is no isolation, with the agent running alongside other workloads; level two uses properly configured process separation and containers, with residual risk of kernel breakout; level three uses virtualisation; and level four runs the agent on dedicated hardware.

The NCSC notes that agents may be able to discover and exploit configuration weaknesses or vulnerabilities in technical controls, potentially escaping a sandbox, and recommends multiple layers of isolation, regular validation of configurations and explicit instructions not to attempt escape or to connect outside the allowlist. It also flags that organisations should consider whether their log collection infrastructure itself creates an attack surface an agent could abuse. It names OpenClaw and Hermes Agent as examples of autonomous tools designed to interact with large numbers of systems, observing that wider access increases the potential blast radius if an agent malfunctions or is compromised.

The guidance suggests initial experiments be run during office hours when human oversight is available, expanding to overnight or weekend execution only once controls are shown to be effective. Abuse reports about an agent's activity should be handled like vulnerability reports under an incident management playbook.

The blog builds on the NCSC's May 2026 advice Thinking carefully before adopting agentic AI, which summarised joint Five Eyes guidance on careful adoption of agentic AI services, and points readers to ETSI EN 304 223, the baseline cyber security standard for AI models and systems.

Also in this section

Aug 03, 2026

NHS England to revise Palantir platform metrics after FOI request reveals internal doubts

NHS England is to change the methodology behind two of its most widely cited claims about the Palantir-built Federated Data Platform, after internal emails disclosed under a freedom of information request showed the health service's own analysts had questioned the validity of the baseline used to measure the system's success, the Financial Times has reported.
Jul 22, 2026

AI in education: the importance of transparency

The debate around AI in education often focuses on capability. What can it do? How much time can it save? Yet some of the most important questions are not technical at all. Adam Halsey looks at why transparency should come before implementation.

InfoGov Masthead Newsletter 800