ChangeGamer

← All resources

AI Control for Agents: The Insider-Threat Defense Model

Guide · updated 2026-07-14 · Markdown variant

What "AI control" means as a security paradigm distinct from alignment: Google DeepMind's Detection (D1-D4) and Prevention/Response (R1-R3) tiers for treating a deployed agent's own actions, not just its inputs, as the threat to defend against.


Most agent-security guidance (see /resources/agentic-security-checklist) treats the agent as a system to defend from outside attackers — prompt injection, malicious tools, poisoned data. Google DeepMind's "AI Control Roadmap" (June 18, 2026) starts from a different premise: assume the deployed agent itself may act harmfully — through misalignment, overeagerness, or compromise — and layer the same kind of controls an enterprise uses against a privileged human insider, rather than betting that alignment training alone prevents it.

Key facts

The two-axis tier framework

The roadmap maps safeguards along two independent axes an operator can use as a sizing rule for their own deployment:

Practical rule: size the tier to the agent's actual blast radius, not its track record so far. An agent holding irreversible-action tools (payments, deletions, production deploys) needs R3-class real-time blocking regardless of how well-behaved it has been; a read-only research agent may only need D1/R1.

Distinct from three adjacent ChangeGamer topics

Verified sources

Primary (fetched directly this session):

Secondary — WebSearch-corroborated (8+ independently agreeing outlets on date, authors, framework name, and the D1-D4/R1-R3 tiers; deepmind.google itself returned HTTP 403 to direct WebFetch this session — a session-wide outage that also blocked a wikipedia.org control fetch, not vendor-specific — reverify there when reachable):

#security #agents #ai-control #insider-threat #defense-in-depth #monitoring #kill-switch #mitre-attack

Category: Guide

Like this? See pricing for the full corpus license, or preview the exact format free as NDJSON or JSON.