# The AI Agent Security Checklist

> A go/no-go checklist that turns the agent security operations pillar's eight disciplines, plus content provenance, agent identity, and data privacy, into checkable gates — the specific inventory row, test result, or logged decision that proves each one holds, with a link to whichever sibling article owns its mechanics.

Guide: Agent security operations — part 12
Published: 2026-09-19 · Updated: 2026-09-19 · 2282 words · ~3035 tokens (estimate)
Canonical: https://changegamer.ai/articles/agent-security-operations-checklist
JSON: https://changegamer.ai/api/articles/agent-security-operations-checklist.json
Pillar: https://changegamer.ai/articles/agent-security-operations.md

## In short

- An AI agent security review only becomes verifiable once each of eleven disciplines is stated as a specific artifact a reviewer checks, not as a claimed practice like "we sandbox the code" or "we redact PII."
- The credential-custody and identity gates both reduce to an inventory fact: every tool server or MCP connector is classified as pass-through or custodial, and every running agent instance carries its own cryptographic workload credential rather than a fleet-wide shared key.
- A permission-boundary gate that only fires immediately before a tool call misses most of what Microsoft's draft Agent Control Specification actually checks — eight lifecycle checkpoints, with a policy-evaluation failure required to return deny rather than falling through to allow.
- The incident-response gate only counts as passed once a reviewer can point to a completed rehearsal record — a real export, on a distinct credential, finished inside the response plan's own time window — not a written procedure nobody has run.
- No single one of the eleven security gates in an AI agent production review is unusual by itself — a real incident routinely works its way through a leaked credential, an under-logged trace, and a missed retention limit in the same run.
- Clearing every gate in this checklist defends an AI agent deployment against an external attacker, a compromised dependency, and a data-handling gap, but says nothing about the deployed agent's own action going wrong with no attacker involved at all.

---

The [agent security operations](/articles/agent-security-operations) pillar names eight operator-side disciplines; three sibling subs in this cluster added three more — content provenance, agent identity and delegated authority, and data privacy. This article turns all eleven into a go/no-go gate: the inventory row, test result, or logged decision proving the condition holds, linked to whichever sibling owns that discipline's mechanics. It does not re-derive that content — a gate names the evidence checked, not the how-to behind producing it.

## From named discipline to go/no-go gate

A named discipline and a go/no-go gate answer different questions. "We handle credential hygiene" names an intended practice a team can assert without a reviewer confirming it held in this deployment; a gate names the exact fact checked directly — does the inventory row exist, did the test run, does the logged decision have a named approver. The eleven gates follow the order a reviewer needs them: identity and credentials first, since nothing downstream means anything against an unauthenticated caller; then content- and code-level defenses; then provenance, audit, incident response, and rate limits; then the two data-handling angles this cluster added last.

## Who is this caller, and what does it hold?

Before any custody or permission question applies, an agent has to be proven a specific workload, and separately shown to hold a credential nobody else is also custodian of.

- **Gate 1 — identity and delegated authority.** Evidence: each agent process authenticates with a distinct, short-lived, cryptographically attested workload credential — a SPIFFE/SPIRE SVID or equivalent — instead of one static key copied fleet-wide, and every delegated-authority token comes from a scoped OAuth grant sized to the task, not the user's own forwarded access token. Mechanics: [how AI agents prove identity and delegated authority](/articles/agent-identity-and-authentication-for-ai-agents).
- **Gate 2 — credential custody.** Evidence: an inventory lists every tool server, MCP connector, and plugin the agent calls, sorts each into the pass-through/custodial split by who actually holds the downstream credential, and confirms each custodial dependency's token lifetime and revocation path directly rather than from vendor documentation — with provider-side prompt caches in the same disclosure inventory, since a cache retains whatever populated it regardless of the rendered transcript. Mechanics: [how to manage secrets for AI agents in production](/articles/credential-hygiene-for-ai-agents).

No-go: a shared fleet-wide key instead of a per-instance credential, or an inventory that has never confirmed a single custodial connector's real token lifetime, clears neither gate regardless of the rest of the stack.

## Can untrusted content or untrusted code reach an action it shouldn't?

Content an agent reads and code it runs are both attacker-reachable — these gates confirm a structural barrier exists between that untrusted material and the actions it could trigger, not that a filter merely catches a bad case after the fact.

- **Gate 3 — prompt-injection architecture.** Evidence: each production task has at least one of the six named architectural patterns — Action-Selector, Plan-Then-Execute, LLM Map-Reduce, Dual LLM, Code-Then-Execute, Context-Minimization — mapped to it by that task's actual blast radius, with a high-stakes, authenticated-session task showing two patterns composed rather than one applied uniformly across every agent a team runs. Mechanics: [how to defend an AI agent against prompt injection](/articles/prompt-injection-defense-in-depth-for-agents).
- **Gate 4 — sandboxing.** Evidence: the isolation layer for a given code-execution task matches its quadrant on the trust-in-source-by-blast-radius grid — never a plain container for fully untrusted code reaching a high-value target — egress is denied by default and proven by actually trying to reach a disallowed host from a live sandbox rather than only documented as policy, and no real secret is ever injected into it. Mechanics: [how to choose a sandbox for AI agent code execution](/articles/choosing-a-sandbox-for-ai-agent-code-execution).

No-go: a pattern chosen because it's the one a team already knows, rather than sized to blast radius, or an egress policy that exists only as a document nobody has tested from inside a live sandbox, has not cleared its gate.

## Has every dependency and every piece of ingested media earned trust before it's used?

A dependency and an ingested media file both carry a claim about their own origin — these gates confirm that claim was actually checked, not assumed from a README or a familiar file type.

- **Gate 5 — supply-chain provenance.** Evidence: at each of three checkpoints — package-install time, model-load time, MCP-server-connect time — a dependency missing a bill of materials or a build-provenance attestation is blocked by default, and any override is a named approver's own logged exception, not something untraceable later. Mechanics: [how to verify supply-chain provenance for AI agent dependencies](/articles/supply-chain-provenance-for-ai-agents).
- **Gate 6 — content provenance.** Evidence: a C2PA reader runs inline on every ingested media file before it reaches the model, routing a valid trust-list-verified manifest, an invalid or tampered signature, and an absent manifest to three genuinely different outcomes rather than one pass/fail check — with every invalid-signature result flagged for review specifically, since a broken signature is a stronger warning than a manifest that was never present. Mechanics: [how to verify content provenance for AI agents with C2PA](/articles/content-provenance-for-ai-agents).

No-go: installing a dependency with neither a BOM nor an attestation and no logged override, or treating a missing C2PA manifest and an invalid signature as the same severity, misses the distinction each gate exists to enforce.

## Does a granted permission actually get checked once the agent is running?

A permission scoped correctly at issuance is a separate claim from a permission enforced at use — this gate confirms the harder of the two.

- **Gate 7 — permission-boundary enforcement.** Evidence: policy checks fire at several distinct points across an agent's run, not solely right before a tool call executes; a deliberately broken policy-evaluation engine produces a deny verdict in a test rather than an unenforced pass-through; and an escalate verdict on an irreversible action reaches a human reviewer's queue rather than resolving itself. Mechanics: [how to enforce permission boundaries for AI agents at runtime](/articles/permission-boundaries-and-least-privilege-for-ai-agents).

No-go: a permission model checked only once, at grant time, and never tested against its own policy engine failing, has not earned this gate no matter how tightly the grant was scoped.

## Is there a record of what happened, and is that record readable during an incident?

Logging that satisfies a reliability dashboard and logging that satisfies a security investigation are not the same bar — this gate confirms the harder one.

- **Gate 8 — audit trails.** Evidence: every tool call lands in an append-only store carrying all seven required fields at once — outcome, tool name, timestamp, full response, trace ID, latency, full arguments — durably linked back to the credential or instance that generated it, with full-payload capture deliberately switched on rather than left at OpenTelemetry's off-by-default PII-safety posture, which would otherwise silently withhold the two fields an audit needs most. Mechanics: [how to build an audit trail for an AI agent](/articles/audit-trails-for-ai-agents).

No-go: a pipeline that looks complete because spans exist for every call, but never opted in to full-argument and full-response capture, has exactly the gap this gate exists to catch.

## When something goes wrong, does the response actually preserve what happened?

An incident-response plan that exists only as an unwritten habit is indistinguishable, mid-incident, from having no plan at all — this gate checks for the pieces that make the difference.

- **Gate 9 — incident response.** Evidence: a named on-call security role can revoke any agent credential unilaterally, without the sign-off process ordinary infrastructure changes route through; "stopping" an instance means quarantining its process and network access rather than a destructive teardown that could erase the trace data the next step needs; and the evidence-export path has already handled a live rehearsal on a credential distinct from anything the incident itself would revoke. Mechanics: [how to respond when an AI agent takes a harmful action](/articles/security-incident-response-for-ai-agents).

No-go: an evidence-export destination nobody has written a trace to and read one back from, ahead of the first live incident, is a stated intention rather than a tested control — this gate stays failed until that rehearsal runs.

## Is the blast radius of a single compromised or malfunctioning instance actually bounded?

A correctly scoped credential and a correctly enforced permission can still let one runaway or compromised instance rack up unbounded cost or action volume if nothing caps what a single credential can do in a given window.

- **Gate 10 — rate and abuse controls.** Evidence: a per-credential tool-call counter lives in shared, atomically updated state rather than one worker's own memory, so two calls arriving at once can't both see room below the ceiling and proceed past it; and a per-session spend ceiling is checked before a tool call goes out, aborting the call if dispatching it would cross the ceiling rather than reconciling the overspend afterward. Mechanics: [how to rate-limit and cap spend for your own AI agent](/articles/rate-and-abuse-controls-for-ai-agents).

No-go: a spend ceiling checked only after a tool call returns has already let through the exact overspend it exists to prevent, however quickly it gets flagged afterward.

## Is personal data contained before it reaches the model, not cleaned up after?

Personal data entering an agent's context, a tool call, a memory write, or a log line is exposed the moment it arrives — this gate confirms containment happens before that point, not as a downstream cleanup step.

- **Gate 11 — data privacy and PII.** Evidence: a PII detector runs against user input, retrieved content, and every tool result before it reaches the model's context window; logs and traces get redaction applied at the exporter layer rather than shipped raw; a workload processing regulated personal data has a confirmed Zero Data Retention or no-training tier switched on; and a written, short expiry governs memory stores and vector indices, not only the log store. Mechanics: [how to protect PII and personal data in AI agent pipelines](/articles/data-privacy-and-pii-for-ai-agents).

No-go: redaction applied to a stored log after the fact, rather than to content before it enters the model's context, has already let the exposure this gate exists to prevent happen once.

## The eleven-gate summary table

| # | Discipline | Gate condition | Mechanics owned by |
|---|---|---|---|
| 1 | Identity and delegated authority | Per-instance workload credential (SVID or equivalent); scoped OAuth grant, never a forwarded user token | [Agent identity and authentication](/articles/agent-identity-and-authentication-for-ai-agents) |
| 2 | Credential custody | Every tool server/connector classified pass-through or custodial; prompt caches in the disclosure inventory | [Credential hygiene for AI agents](/articles/credential-hygiene-for-ai-agents) |
| 3 | Prompt-injection architecture | A named pattern mapped to task blast radius; two patterns composed for high-stakes tasks | [Prompt-injection defense in depth](/articles/prompt-injection-defense-in-depth-for-agents) |
| 4 | Sandboxing | Isolation layer matches the trust/blast-radius quadrant; egress denial tested live | [Choosing a sandbox for AI agent code execution](/articles/choosing-a-sandbox-for-ai-agent-code-execution) |
| 5 | Supply-chain provenance | Missing BOM/attestation blocks by default at all three checkpoints; overrides logged | [Supply-chain provenance for AI agents](/articles/supply-chain-provenance-for-ai-agents) |
| 6 | Content provenance | C2PA check runs inline; valid/invalid/absent routed to three distinct outcomes | [Content provenance for AI agents with C2PA](/articles/content-provenance-for-ai-agents) |
| 7 | Permission-boundary enforcement | Multi-checkpoint policy checks; engine failure tested to deny, not pass through | [Permission boundaries and least privilege](/articles/permission-boundaries-and-least-privilege-for-ai-agents) |
| 8 | Audit trails | Seven-field logging with credential link; full-payload capture opted in | [Audit trails for AI agents](/articles/audit-trails-for-ai-agents) |
| 9 | Incident response | Standing revocation authority; quarantine not teardown; export path drill-tested | [Security incident response for AI agents](/articles/security-incident-response-for-ai-agents) |
| 10 | Rate and abuse controls | Atomic shared per-credential counter; pre-dispatch fail-closed spend ceiling | [Rate and abuse controls for AI agents](/articles/rate-and-abuse-controls-for-ai-agents) |
| 11 | Data privacy and PII | Pre-send redaction; ZDR/no-training tier confirmed; explicit retention limits | [Data privacy and PII for AI agents](/articles/data-privacy-and-pii-for-ai-agents) |

None of these eleven gates amounts to much in isolation — an inventory row, a test result, a logged approval, a retention window. The full review earns its keep because a real incident rarely trips only one: a credential a tool server holds custodially (gate 2) is exactly what an under-scoped permission check (gate 7) lets an attacker reach, and the same incident is what an untested evidence-export path (gate 9) then fails to preserve a record of. Eleven unrelated checkboxes misses how often one compromise works through several at once.

## Beyond this checklist: the agent's own behavior

Every gate above defends against the same three threats — an external attacker, a compromised or unverified dependency, or a data-handling gap — none requiring the deployed agent to act against its operator's intent. As of July 2026, Google DeepMind's "AI Control Roadmap" names a fourth threat this checklist doesn't gate against: the agent's own action going wrong through misalignment, overeagerness, or compromise, no external attacker or injected instruction involved. The [AI control for agents](/resources/ai-control-for-agents) reference frames that as a paradigm distinct from alignment — assume the model might not want the right thing, and add system-level safeguards that don't depend on its intentions holding — structured as Detection tiers D1-D4 (logging chain-of-thought for review, up to inspecting internal model representations directly) and Prevention/Response tiers R1-R3 (asynchronous review-and-remediate, up to synchronous real-time blocking before a harmful action executes), built on general-purpose MITRE ATT&CK rather than an AI-specific taxonomy. None of the eleven gates above owns this ground, since none monitors the agent's own reasoning or action stream in real time — that is the layer past a production security review, not a twelfth gate this checklist states in the same terms.

See the [agent security operations](/articles/agent-security-operations) pillar for how these eleven disciplines fit together as a single stack, and each of the eleven sibling articles linked above for the mechanics behind its own gate.

## Frequently asked questions

### What is the difference between an AI agent security checklist and a production security gate?

An AI agent security checklist lists an intended practice a team can claim without anyone confirming it holds, such as "credentials are scoped" or "dependencies are verified," while a production security gate names the exact artifact a reviewer checks instead — an inventory row classifying a connector as custodial, a test result showing a policy engine denies on its own failure, a logged risk acceptance for an unverified dependency — so the claim resolves to a checked yes or no rather than a self-reported label.

### What evidence proves an AI agent's permission-boundary gate is actually enforced, not just granted?

Proving a permission-boundary gate is enforced requires a test that deliberately breaks the policy-evaluation engine and confirms the resulting call gets denied rather than passed through, plus confirmation that a boundary is checked repeatedly across an agent's run rather than only once — a grant scoped correctly at issuance but never checked again at runtime is a policy document, not an enforced boundary.

### Why does incident response for a compromised AI agent credential need a separate evidence-export path?

The evidence-export gate is about verification, not the underlying mechanism — the reasoning for why the export and containment paths must stay independent belongs to the incident-response sibling article. What a reviewer checks here is narrower and more concrete: has the export actually been run end to end, on infrastructure sharing no identity with anything an incident would touch, and timed against the response plan's own stated deadline? A written assumption that the two never overlap is not evidence; a dated rehearsal log is.

### Does passing every gate in the AI agent security checklist mean the agent is safe from its own misaligned behavior?

No — every gate in this checklist defends against an external attacker, a compromised dependency, or a data-handling failure, none of which requires the deployed agent itself to act against its operator's intent, so a separate layer of real-time behavioral monitoring of the agent's own actions is needed to catch misalignment or compromise that produces no external attack signature at all.


---

## The rest of this guide

- [How to Secure AI Agents in Production](https://changegamer.ai/articles/agent-security-operations.md): Credential hygiene, prompt-injection defense in depth, sandboxing choices for code execution, supply-chain provenance, least privilege, audit trails, incident response, and rate/abuse controls — eight operator-side defenses against an adversarial actor or a compromised dependency, not against ordinary load or failure.
- [How to Manage Secrets for AI Agents in Production](https://changegamer.ai/articles/credential-hygiene-for-ai-agents.md): Why single-agent credential issuance is not the whole secrets problem: auditing which tool servers and MCP connectors hold credentials on an agent's behalf, treating provider-side prompt caches as a disclosure surface, and the named frameworks — OWASP's Secrets Management Cheat Sheet, Twelve-Factor config, and the OWASP GenAI project — that govern the rest.
- [How to Defend an AI Agent Against Prompt Injection](https://changegamer.ai/articles/prompt-injection-defense-in-depth-for-agents.md): A decision framework for matching Action-Selector, Plan-Then-Execute, Dual LLM, and the other named architectural patterns to a task's actual blast radius, including when to compose two patterns together and when none of them is worth the overhead.
- [How to Choose a Sandbox for AI Agent Code Execution](https://changegamer.ai/articles/choosing-a-sandbox-for-ai-agent-code-execution.md): A two-axis framework — trust in the code's source crossed with the blast radius of a successful escape — for picking an isolation layer, choosing among six hosted sandbox APIs, and hardening the harness around whichever one you pick.
- [How to Verify Supply-Chain Provenance for AI Agent Dependencies](https://changegamer.ai/articles/supply-chain-provenance-for-ai-agents.md): An operational playbook for three separate trust-boundary gates — package-install time, model-load time, and MCP-server-connect time — that turns SBOM and attestation formats into checks a pipeline can actually run, plus a fail-closed default for the dependency that carries neither.
- [How to Enforce Permission Boundaries for AI Agents at Runtime](https://changegamer.ai/articles/permission-boundaries-and-least-privilege-for-ai-agents.md): How a least-privilege grant actually gets enforced once an agent is running — the eight named checkpoints, five verdicts, and fail-closed contract in Microsoft's draft Agent Control Specification (ACS), and what happens when the policy engine that enforces the boundary breaks.
- [How to Build an Audit Trail for an AI Agent](https://changegamer.ai/articles/audit-trails-for-ai-agents.md): A field-by-field forensic playbook for an AI agent audit log: why each of the checklist's seven required fields matters for reconstruction, a worked incident walkthrough, and the real tension between OpenTelemetry's redact-by-default tracing and full audit logging.
- [How to Respond When an AI Agent Takes a Harmful Action](https://changegamer.ai/articles/security-incident-response-for-ai-agents.md): Why the pillar's cut-credential, stop-instance, export-evidence order is structurally forced rather than a tidy convention, a concrete answer for who holds standing revocation authority, a pre-incident drill for the evidence-export path, and what changes when the compromised credential belongs to a third-party tool server instead of the agent itself.
- [How to Rate-Limit and Cap Spend for Your Own AI Agent](https://changegamer.ai/articles/rate-and-abuse-controls-for-ai-agents.md): Enforcement mechanics for the two ceilings an agent operator should set before production: where a per-credential tool-call counter has to live to stay correct under concurrent calls, where a spend ceiling gets checked in the tool-call loop, and how to reject out-of-scope tool-call arguments with canonicalization rather than a naive prefix match.
- [How to Verify Content Provenance for AI Agents with C2PA](https://changegamer.ai/articles/content-provenance-for-ai-agents.md): A three-state decision procedure — valid manifest, invalid signature, absent manifest — for what an AI agent's ingestion pipeline should do differently with a web image, an email attachment, or a retrieved document, plus a checklist for wiring a C2PA reader library into that pipeline as a gate before content reaches the model.
- [How AI Agents Prove Identity and Delegated Authority](https://changegamer.ai/articles/agent-identity-and-authentication-for-ai-agents.md): The two-layer model an autonomous agent needs to pass before any credential-custody or permission question even applies: a cryptographic workload identity proving what it is (SPIFFE/SPIRE, cloud workload identity federation) and a separate delegated-authority grant proving it may act on a human's or org's behalf (OAuth scopes, RFC 8693 token exchange, RFC 8707 audience binding).
- [How to Protect PII and Personal Data in AI Agent Pipelines](https://changegamer.ai/articles/data-privacy-and-pii-for-ai-agents.md): Why an AI agent expands PII exposure past a bounded API call — large ingested context, external tool calls, persistent memory and logs, provider training risk — and the containment controls, provider data-handling terms, and GDPR/EU AI Act/CCPA compliance boundary that follow from it.

## Reference resources

- https://changegamer.ai/resources/agentic-security-checklist.md
- https://changegamer.ai/resources/secrets-management-for-agents.md
- https://changegamer.ai/resources/agent-identity-authentication.md
- https://changegamer.ai/resources/prompt-injection-design-patterns.md
- https://changegamer.ai/resources/code-execution-sandboxing.md
- https://changegamer.ai/resources/ai-supply-chain-provenance.md
- https://changegamer.ai/resources/c2pa-content-credentials.md
- https://changegamer.ai/resources/agent-control-specification.md
- https://changegamer.ai/resources/agent-observability.md
- https://changegamer.ai/resources/data-privacy-for-agents.md
- https://changegamer.ai/resources/handling-rate-limits-and-retries.md
- https://changegamer.ai/resources/shipping-agents-to-production.md
- https://changegamer.ai/resources/ai-control-for-agents.md

All guides: https://changegamer.ai/api/articles.json · Reference corpus: https://changegamer.ai/llms.txt
Licensing: https://changegamer.ai/api/pricing.json (offer catalog) · https://changegamer.ai/api/payment.json (payment methods, HTTP 402 flow) · access guide: https://changegamer.ai/resources/access-and-pricing.md
