ChangeGamer

← All guides · Agent security operations

The AI Agent Security Checklist

Part 12 of Agent security operations · 2,282 words · ~10 min read · published 2026-09-19 · updated 2026-09-19 · Markdown variant

A go/no-go checklist that turns the agent security operations pillar's eight disciplines, plus content provenance, agent identity, and data privacy, into checkable gates — the specific inventory row, test result, or logged decision that proves each one holds, with a link to whichever sibling article owns its mechanics.

In short

  • An AI agent security review only becomes verifiable once each of eleven disciplines is stated as a specific artifact a reviewer checks, not as a claimed practice like "we sandbox the code" or "we redact PII."
  • The credential-custody and identity gates both reduce to an inventory fact: every tool server or MCP connector is classified as pass-through or custodial, and every running agent instance carries its own cryptographic workload credential rather than a fleet-wide shared key.
  • A permission-boundary gate that only fires immediately before a tool call misses most of what Microsoft's draft Agent Control Specification actually checks — eight lifecycle checkpoints, with a policy-evaluation failure required to return deny rather than falling through to allow.
  • The incident-response gate only counts as passed once a reviewer can point to a completed rehearsal record — a real export, on a distinct credential, finished inside the response plan's own time window — not a written procedure nobody has run.
  • No single one of the eleven security gates in an AI agent production review is unusual by itself — a real incident routinely works its way through a leaked credential, an under-logged trace, and a missed retention limit in the same run.
  • Clearing every gate in this checklist defends an AI agent deployment against an external attacker, a compromised dependency, and a data-handling gap, but says nothing about the deployed agent's own action going wrong with no attacker involved at all.

Part of the How to Secure AI Agents in Production guide.


The agent security operations pillar names eight operator-side disciplines; three sibling subs in this cluster added three more — content provenance, agent identity and delegated authority, and data privacy. This article turns all eleven into a go/no-go gate: the inventory row, test result, or logged decision proving the condition holds, linked to whichever sibling owns that discipline's mechanics. It does not re-derive that content — a gate names the evidence checked, not the how-to behind producing it.

From named discipline to go/no-go gate

A named discipline and a go/no-go gate answer different questions. "We handle credential hygiene" names an intended practice a team can assert without a reviewer confirming it held in this deployment; a gate names the exact fact checked directly — does the inventory row exist, did the test run, does the logged decision have a named approver. The eleven gates follow the order a reviewer needs them: identity and credentials first, since nothing downstream means anything against an unauthenticated caller; then content- and code-level defenses; then provenance, audit, incident response, and rate limits; then the two data-handling angles this cluster added last.

Who is this caller, and what does it hold?

Before any custody or permission question applies, an agent has to be proven a specific workload, and separately shown to hold a credential nobody else is also custodian of.

No-go: a shared fleet-wide key instead of a per-instance credential, or an inventory that has never confirmed a single custodial connector's real token lifetime, clears neither gate regardless of the rest of the stack.

Can untrusted content or untrusted code reach an action it shouldn't?

Content an agent reads and code it runs are both attacker-reachable — these gates confirm a structural barrier exists between that untrusted material and the actions it could trigger, not that a filter merely catches a bad case after the fact.

No-go: a pattern chosen because it's the one a team already knows, rather than sized to blast radius, or an egress policy that exists only as a document nobody has tested from inside a live sandbox, has not cleared its gate.

Has every dependency and every piece of ingested media earned trust before it's used?

A dependency and an ingested media file both carry a claim about their own origin — these gates confirm that claim was actually checked, not assumed from a README or a familiar file type.

No-go: installing a dependency with neither a BOM nor an attestation and no logged override, or treating a missing C2PA manifest and an invalid signature as the same severity, misses the distinction each gate exists to enforce.

Does a granted permission actually get checked once the agent is running?

A permission scoped correctly at issuance is a separate claim from a permission enforced at use — this gate confirms the harder of the two.

No-go: a permission model checked only once, at grant time, and never tested against its own policy engine failing, has not earned this gate no matter how tightly the grant was scoped.

Is there a record of what happened, and is that record readable during an incident?

Logging that satisfies a reliability dashboard and logging that satisfies a security investigation are not the same bar — this gate confirms the harder one.

No-go: a pipeline that looks complete because spans exist for every call, but never opted in to full-argument and full-response capture, has exactly the gap this gate exists to catch.

When something goes wrong, does the response actually preserve what happened?

An incident-response plan that exists only as an unwritten habit is indistinguishable, mid-incident, from having no plan at all — this gate checks for the pieces that make the difference.

No-go: an evidence-export destination nobody has written a trace to and read one back from, ahead of the first live incident, is a stated intention rather than a tested control — this gate stays failed until that rehearsal runs.

Is the blast radius of a single compromised or malfunctioning instance actually bounded?

A correctly scoped credential and a correctly enforced permission can still let one runaway or compromised instance rack up unbounded cost or action volume if nothing caps what a single credential can do in a given window.

No-go: a spend ceiling checked only after a tool call returns has already let through the exact overspend it exists to prevent, however quickly it gets flagged afterward.

Is personal data contained before it reaches the model, not cleaned up after?

Personal data entering an agent's context, a tool call, a memory write, or a log line is exposed the moment it arrives — this gate confirms containment happens before that point, not as a downstream cleanup step.

No-go: redaction applied to a stored log after the fact, rather than to content before it enters the model's context, has already let the exposure this gate exists to prevent happen once.

The eleven-gate summary table

# Discipline Gate condition Mechanics owned by
1 Identity and delegated authority Per-instance workload credential (SVID or equivalent); scoped OAuth grant, never a forwarded user token Agent identity and authentication
2 Credential custody Every tool server/connector classified pass-through or custodial; prompt caches in the disclosure inventory Credential hygiene for AI agents
3 Prompt-injection architecture A named pattern mapped to task blast radius; two patterns composed for high-stakes tasks Prompt-injection defense in depth
4 Sandboxing Isolation layer matches the trust/blast-radius quadrant; egress denial tested live Choosing a sandbox for AI agent code execution
5 Supply-chain provenance Missing BOM/attestation blocks by default at all three checkpoints; overrides logged Supply-chain provenance for AI agents
6 Content provenance C2PA check runs inline; valid/invalid/absent routed to three distinct outcomes Content provenance for AI agents with C2PA
7 Permission-boundary enforcement Multi-checkpoint policy checks; engine failure tested to deny, not pass through Permission boundaries and least privilege
8 Audit trails Seven-field logging with credential link; full-payload capture opted in Audit trails for AI agents
9 Incident response Standing revocation authority; quarantine not teardown; export path drill-tested Security incident response for AI agents
10 Rate and abuse controls Atomic shared per-credential counter; pre-dispatch fail-closed spend ceiling Rate and abuse controls for AI agents
11 Data privacy and PII Pre-send redaction; ZDR/no-training tier confirmed; explicit retention limits Data privacy and PII for AI agents

None of these eleven gates amounts to much in isolation — an inventory row, a test result, a logged approval, a retention window. The full review earns its keep because a real incident rarely trips only one: a credential a tool server holds custodially (gate 2) is exactly what an under-scoped permission check (gate 7) lets an attacker reach, and the same incident is what an untested evidence-export path (gate 9) then fails to preserve a record of. Eleven unrelated checkboxes misses how often one compromise works through several at once.

Beyond this checklist: the agent's own behavior

Every gate above defends against the same three threats — an external attacker, a compromised or unverified dependency, or a data-handling gap — none requiring the deployed agent to act against its operator's intent. As of July 2026, Google DeepMind's "AI Control Roadmap" names a fourth threat this checklist doesn't gate against: the agent's own action going wrong through misalignment, overeagerness, or compromise, no external attacker or injected instruction involved. The AI control for agents reference frames that as a paradigm distinct from alignment — assume the model might not want the right thing, and add system-level safeguards that don't depend on its intentions holding — structured as Detection tiers D1-D4 (logging chain-of-thought for review, up to inspecting internal model representations directly) and Prevention/Response tiers R1-R3 (asynchronous review-and-remediate, up to synchronous real-time blocking before a harmful action executes), built on general-purpose MITRE ATT&CK rather than an AI-specific taxonomy. None of the eleven gates above owns this ground, since none monitors the agent's own reasoning or action stream in real time — that is the layer past a production security review, not a twelfth gate this checklist states in the same terms.

See the agent security operations pillar for how these eleven disciplines fit together as a single stack, and each of the eleven sibling articles linked above for the mechanics behind its own gate.

Frequently asked questions

What is the difference between an AI agent security checklist and a production security gate?
An AI agent security checklist lists an intended practice a team can claim without anyone confirming it holds, such as "credentials are scoped" or "dependencies are verified," while a production security gate names the exact artifact a reviewer checks instead — an inventory row classifying a connector as custodial, a test result showing a policy engine denies on its own failure, a logged risk acceptance for an unverified dependency — so the claim resolves to a checked yes or no rather than a self-reported label.
What evidence proves an AI agent's permission-boundary gate is actually enforced, not just granted?
Proving a permission-boundary gate is enforced requires a test that deliberately breaks the policy-evaluation engine and confirms the resulting call gets denied rather than passed through, plus confirmation that a boundary is checked repeatedly across an agent's run rather than only once — a grant scoped correctly at issuance but never checked again at runtime is a policy document, not an enforced boundary.
Why does incident response for a compromised AI agent credential need a separate evidence-export path?
The evidence-export gate is about verification, not the underlying mechanism — the reasoning for why the export and containment paths must stay independent belongs to the incident-response sibling article. What a reviewer checks here is narrower and more concrete: has the export actually been run end to end, on infrastructure sharing no identity with anything an incident would touch, and timed against the response plan's own stated deadline? A written assumption that the two never overlap is not evidence; a dated rehearsal log is.
Does passing every gate in the AI agent security checklist mean the agent is safe from its own misaligned behavior?
No — every gate in this checklist defends against an external attacker, a compromised dependency, or a data-handling failure, none of which requires the deployed agent itself to act against its operator's intent, so a separate layer of real-time behavioral monitoring of the agent's own actions is needed to catch misalignment or compromise that produces no external attack signature at all.

#agents #security #production #checklist #operations

Put this corpus inside your own agents

This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.

Agents: this page as Markdown · JSON · offers at /api/pricing.json · payment methods at /api/payment.json · single-resource access via HTTP 402 (how that works)