ChangeGamer

← All guides · Agent security operations

How to Enforce Permission Boundaries for AI Agents at Runtime

Part 5 of Agent security operations · 1,749 words · ~8 min read · published 2026-09-14 · updated 2026-09-14 · Markdown variant

How a least-privilege grant actually gets enforced once an agent is running — the eight named checkpoints, five verdicts, and fail-closed contract in Microsoft's draft Agent Control Specification (ACS), and what happens when the policy engine that enforces the boundary breaks.

In short

  • The Agent Control Specification (ACS), a Microsoft draft announced 2 June 2026, names eight runtime checkpoints — agent_startup, input, pre_model_call, post_model_call, pre_tool_call, post_tool_call, output, and agent_shutdown — where a policy dispatcher decides whether a granted permission is actually exercised.
  • A policy dispatcher under ACS returns one of five verdicts at each checkpoint — allow, warn, transform, deny, or escalate — rather than a single binary allow-or-block decision at the moment a tool call is about to fire.
  • ACS specifies that any error in policy evaluation must yield a deny verdict, meaning a correctly scoped least-privilege grant still needs an enforcement layer that refuses action by default the instant its own policy engine breaks.
  • The ACS runtime contract requires that a policy-evaluation engine must not retain mutable state that influences a verdict from one evaluation to the next, so every checkpoint decision has to be reproducible from the current snapshot alone.
  • As of September 2026, ACS is a Draft specification at version 0.3.1-beta inside Microsoft's open-source, MIT-licensed Agent Governance Toolkit, and the spec text itself warns to expect breaking changes before a stable 1.0 release.

Part of the How to Secure AI Agents in Production guide.


What enforces a permission boundary once an agent is granted it?

A permission boundary for an AI agent is enforced by a policy dispatcher checking a live snapshot of the agent's current step against a manifest at a defined point in its execution loop, and returning a verdict the runtime is obligated to act on — not by the grant decision made when the agent was first configured. The agent security operations pillar's permissions section covers that grant-time decision: what scope to hand an agent for a given task, when to prefer read-only, how long a token should live. This piece covers the step after that decision is made — how a granted boundary actually gets checked, call by call, once the agent is running, grounded in the Agent Control Specification (ACS), Microsoft's open, MIT-licensed draft for exactly this enforcement layer, announced 2 June 2026 at Microsoft Build 2026 and currently versioned 0.3.1-beta, status Draft.

ACS does not run an agent and does not decide what to grant it. It defines a stateless, deterministic contract: a host submits a snapshot of the current step plus a policy manifest at a named checkpoint, and the ACS runtime returns a normalized verdict the host must enforce. Framework-specific adapters plug into that contract — LangChain and LangGraph, the OpenAI Agents SDK as middleware, an Anthropic Messages-API adapter, AutoGen, CrewAI, Semantic Kernel natively, MCP tools, and a .NET Microsoft.AgentGovernance package family — so the same eight checkpoints and five verdicts apply regardless of which framework runs the agent underneath. This is a distinct question from who holds a credential across a chain of tool servers and MCP connectors, covered in managing secrets for AI agents; this piece asks a narrower question about that same credential — not who has custody of it, but what it gets waved through for on any single call.

The eight checkpoints where a permission boundary is actually enforced

ACS checks a permission boundary at eight named points across an agent's lifecycle, not only once before a tool call fires, because a boundary that only gets checked at one point in the loop leaves every other point unguarded. Each checkpoint gates a different kind of decision, illustrated below with a concrete least-privilege scenario at each one:

Checkpoint Gates A least-privilege scenario at this checkpoint
agent_startup The run before a single step executes The manifest confirms the agent's declared tool list matches what this task actually needs; a tool outside that scope is never made callable for the run at all
input The external request entering the agent A request asking the agent to act on a resource it doesn't own for the current tenant is denied before it ever reaches the model
pre_model_call What the model is about to see An over-broad tool description or an unrelated credential reference gets transformed out of the context assembled for the model, before the call
post_model_call What the model just proposed The model's proposed next action names a tool outside the task's granted scope; the dispatcher denies it before any tool call fires
pre_tool_call The specific call about to execute A task granted read-only access attempts a delete call; the dispatcher denies the call at the tool boundary itself, not after it runs
post_tool_call The result about to flow back to the agent A lookup tool returns more fields than the task's scope authorizes; the dispatcher transforms the response, trimming it down before the agent sees it
output The final assembled response leaving the agent A response confirming an irreversible action — a transfer, a deletion — is escalated to a host approval workflow rather than released automatically
agent_shutdown The end of the run The dispatcher warns and logs if a credential issued for the task was not confirmed released before the run closed out

Read across the table, the pattern is that "permission" is not one gate but eight, each one closing a specific window where an over-scoped action could otherwise slip through — before the model sees anything, after it proposes something, immediately before a tool executes, and on the way back out.

Five verdicts a policy check can return

A policy check under ACS returns one of five verdicts, not a binary allow-or-block, because a real deployment needs finer-grained responses than "let it happen" or "stop it":

Escalate is the verdict worth naming specifically against the pillar's own grant-time guidance to require explicit confirmation before an irreversible or state-changing call: that confirmation step is exactly what an escalate verdict routes to at runtime, giving the grant-time policy a concrete mechanism to fire through rather than leaving it as an unenforced house rule.

What happens when the policy engine itself fails?

When the ACS policy engine itself fails to evaluate a check — a manifest fails to parse, a dependency the policy logic calls times out, an unhandled exception in the dispatcher — the specification requires the runtime to return deny, not to fall through to whatever verdict the request would otherwise have gotten. This is the detail that closes the actual gap in a least-privilege program: a permission grant can be scoped exactly right at the moment it's issued and still leave an agent over-privileged in practice if the layer checking that grant at call time goes down silently and lets the call through anyway. Failing closed means an enforcement outage looks like a denied action, not an unenforced one.

ChangeGamer's own publish pipeline runs the same posture at a structural level, unrelated to agents: build validation checks that every article's takeaways array holds at least three entries and that no article body contains a markdown H1 outside the page title, and either violation halts the release — no partial site goes out with the one bad page quietly skipped. The mechanism is trivial by comparison to an agent's runtime policy dispatcher, but the posture is the same one ACS specifies for a permission check: when a required verification step can't confirm the thing is fine, the default outcome is refusal, not silent pass-through.

Why must the runtime stay stateless between checks?

The ACS specification states that its runtime "MUST NOT retain mutable state that influences a verdict from one evaluation to the next," because a verdict that depends on accumulated hidden state is not reproducible from the current step alone, which breaks both auditability and the fail-closed guarantee above. If an earlier, unrelated call could quietly change how a later call gets evaluated, a deny at one checkpoint could not be trusted to mean the same thing on a repeat of the identical request — and an incident investigation trying to reconstruct why a specific call was allowed or denied would have no way to do it from the manifest and the snapshot alone. Each checkpoint's verdict has to be a pure function of the current step's snapshot and the policy manifest in force at that moment, nothing carried over from any call before it.

How do you implement the policy logic — Rego, Cedar, or a custom adapter?

Policy logic under ACS is implemented through a pluggable dispatcher, not a fixed rule engine baked into the spec itself: teams can express policy in Open Policy Agent's Rego language, in Cedar, or in a custom host-supplied adapter, and any of the three can sit behind the same eight-checkpoint, five-verdict contract. That pluggability is deliberate — it lets an operator reuse a policy engine and rule set it may already run elsewhere rather than adopting a new policy language solely to gate an agent. It also means the actual rules an organization writes — what counts as an over-scoped delete, which tool names trigger an escalate rather than a deny — live outside the specification entirely, in whichever dispatcher a team chooses to wire up.

As of September 2026, treat ACS itself as a mechanism for enforcing whatever least-privilege rules an operator has already decided on, not as a source of those rules or as a finished, adopted standard. It is community-governed rather than vendor-locked, with adapters shipping for a wide framework list, but the spec text remains Draft at 0.3.1-beta and explicitly warns to expect breaking changes before a stable 1.0 — reasonable to pilot in a non-critical path, not yet something to treat as a settled foundation for a production permission boundary.

A runtime permission-enforcement checklist

A least-privilege grant earns runtime trust once these hold, beyond the grant-time scoping the pillar already covers:

A permission grant scoped exactly right at issuance and never checked again is not a boundary — it is a policy document. What turns it into something an attacker or a bug actually has to get past is an enforcement layer that checks it at every step of the loop, returns a real verdict every time, and refuses action by default the moment it can't confirm the answer.

Frequently asked questions

What are the eight runtime checkpoints in Microsoft's Agent Control Specification?
The eight checkpoints in the Agent Control Specification (ACS) are agent_startup, input, pre_model_call, post_model_call, pre_tool_call, post_tool_call, output, and agent_shutdown, spanning an agent's full lifecycle from before it begins a run to after it ends one, with a policy dispatcher returning a verdict at each point rather than only at the moment a tool actually executes.
What happens when an ACS policy engine fails to evaluate a permission check?
The Agent Control Specification fails closed: any error during policy evaluation at any checkpoint yields a deny verdict rather than defaulting to allow, which means an agent's permission boundary holds even when the engine responsible for enforcing it is the thing that broke, not only when the engine is working correctly.
What is the difference between a warn, transform, and deny verdict in ACS?
A warn verdict permits the action to proceed but logs it, a transform verdict permits the action but replaces its target per the policy manifest before it proceeds, and a deny verdict refuses the action outright — three distinct responses to a policy check, alongside allow (unchanged) and escalate (deferred to the host's own approval workflow), rather than a single pass-or-fail outcome.
Is Microsoft's Agent Control Specification a finished industry standard for agent permissions?
No — as of September 2026 the Agent Control Specification is an open, MIT-licensed Draft at version 0.3.1-beta inside Microsoft's Agent Governance Toolkit, announced 2 June 2026 at Microsoft Build 2026, and its own specification text states plainly to expect breaking changes before a stable 1.0 release, so it should be evaluated as an emerging mechanism rather than adopted as a settled standard.

#agents #security #least-privilege #permissions #policy-engine #runtime-enforcement #acs

Put this corpus inside your own agents

This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.

Agents: this page as Markdown · JSON · offers at /api/pricing.json · payment methods at /api/payment.json · single-resource access via HTTP 402 (how that works)