DeepMind outlines an “AI Control Roadmap” for agents

Google DeepMind has published what it calls an “AI Control Roadmap” aimed at securing advanced internal AI agents. In the post, the team frames the plan as a defense-in-depth approach for building and managing AI that is deployed inside Google.

The core idea is that model alignment alone isn’t enough. DeepMind says the roadmap adds system-level security on top of multiple safeguards, including sandboxing, endpoint security, and defenses meant to reduce the risk of prompt injection.

That matters because many real-world AI agent failures tend to come from how systems behave once they’re connected to tools, permissions, and workflows—not just from what the model learns during training. By describing controls that sit around the model, DeepMind is signaling an approach that focuses on the whole chain of actions agents take.

A second document for policymakers

Alongside the roadmap, DeepMind says it is publishing a technical framework for policymakers called “Three Layers of Agent Security.” The company does not describe it in detail in the summary available here, but the naming suggests a structured way to think about securing agents at more than one level.

Taken together, the two posts point to a practical message: if AI agents are going to be used, the security work can’t stop at improving the model’s behavior. It also has to cover how those agents are isolated, controlled at the system level, and protected against common ways attacks can be carried out through inputs and access.

For organizations watching how agent use may expand, these documents offer a glimpse into what “agent security” could look like when it’s treated as layered engineering rather than a single fix.

Source: Google DeepMind