ChangeGamer

← All guides · Agent security operations

How to Choose a Sandbox for AI Agent Code Execution

Part 3 of Agent security operations · 1,600 words · ~7 min read · published 2026-09-13 · updated 2026-09-13 · Markdown variant

A two-axis framework — trust in the code's source crossed with the blast radius of a successful escape — for picking an isolation layer, choosing among six hosted sandbox APIs, and hardening the harness around whichever one you pick.

In short

  • Choosing a sandbox for AI agent code execution rests on two variables working together — where the code being run came from, and what a break-out from that sandbox could actually reach — not on defaulting to the strongest isolation layer available.
  • Fully untrusted, model-generated code destined for a high-value environment holding real secrets calls for a microVM such as Firecracker or Cloud Hypervisor at minimum, never a plain Docker container, regardless of how convenient a plain container is to run.
  • Among the six hosted agent-sandbox APIs compared as of June 2026 — E2B, Modal, Daytona, Cloudflare Sandbox, Northflank, and Vercel Sandbox — only Modal offers GPU access, with A100 and H100 hardware.
  • Blocking network egress by default only counts as hardening once outbound traffic is denied at the sandbox's firewall or proxy layer and re-tested after every allowlist change, not merely documented as an intended policy.
  • No published source in the hosted-sandbox comparison states usage pricing for any of the six vendors, so treat a specific cost comparison between them as unverified rather than filling that gap with an estimate.

Part of the How to Secure AI Agents in Production guide.


Choosing a sandbox for AI agent code execution means matching one layer from the pillar's isolation spectrum to two variables specific to the task at hand — where the code being run came from, and what a break-out from that sandbox could actually reach — rather than defaulting to whichever option is fastest to wire up. The agent security operations pillar already lays out each layer in that spectrum: in-process language sandboxes, OS containers, hardened runtimes such as gVisor and Kata, microVMs such as Firecracker, and full VMs. This piece treats that breakdown as settled ground. What follows is the part a survey of isolation options leaves out: a framework for picking a layer, a vendor-selection pass across the six hosted sandbox APIs the pillar names, and the harness-hardening checklist turned into steps you can implement and verify, not a restated list of six bullet points.

The two variables that decide where you land on the isolation spectrum

Two variables decide where on the isolation spectrum a given sandbox should sit: where the code being run came from, and what a break-out from that sandbox could actually reach. Neither variable alone settles the question — a fully untrusted source in a disposable, low-value environment doesn't need the same layer as a fully untrusted source that can reach production, and a semi-trusted source pointed at a high-value environment still deserves more isolation than the same code running somewhere throwaway.

Trust in the code's source ranges from fully untrusted — code an agent generates freely in response to arbitrary, often adversarial input such as a fetched web page or unconstrained user text — to semi-trusted, where the code passes through some constraint or review first. The Code-Then-Execute pattern covered in defending an agent against prompt injection, where a human or a rule reviews generated commands before execution, is what moves code from the fully untrusted end toward the semi-trusted one.

Blast radius of a successful escape ranges from low-value and ephemeral — a sandbox holding no real secrets, no persistent volume, and nothing beyond disposable output — to high-value and persistent, where the sandbox can reach a live credential, a shared filesystem, or a network path into production. The rule that a secret must never enter anything a model or its output can touch, covered in managing secrets for AI agents, is what should keep even a strongly isolated sandbox from holding a real key at all — isolation and secret-handling are separate controls, and neither substitutes for the other.

Mapping the four quadrants to an isolation layer

Crossing those two variables produces four quadrants, each pointing to a different minimum layer from the pillar's spectrum rather than one fixed answer for every sandbox a team runs.

Quadrant Example task Minimum layer Why
Fully untrusted, low blast radius Coding-assistant sandbox running arbitrary generated snippets, no secrets, filesystem wiped each run Hardened runtime (gVisor) Zero trust in the source still means a plain container's shared kernel is one exploit from every other tenant on the host
Fully untrusted, high blast radius Agent-written code with network reach into production or a live credential MicroVM (Firecracker or Cloud Hypervisor) Both variables are worst-case at once — the exact situation the pillar's "plain container is not a security boundary" warning targets
Semi-trusted, low blast radius Code reviewed under Code-Then-Execute, run in a disposable environment with no real data OS container Review removed most of the source-trust risk, and the environment has nothing worth escaping toward
Semi-trusted, high blast radius Reviewed code that still touches a shared, production-adjacent resource Hardened runtime or microVM Review lowers but doesn't eliminate risk, so blast radius alone justifies real isolation

Layer definitions, startup costs, and the technology behind each one are the pillar's own territory (/articles/agent-security-operations); this framework only supplies which layer a given combination calls for.

How do you choose among the six hosted sandbox APIs?

Choosing among the six hosted agent-sandbox APIs the pillar names — E2B, Modal, Daytona, Cloudflare Sandbox, Northflank, and Vercel Sandbox — comes down to three questions once the isolation layer is settled: does the workload need a GPU, how much do cold start and session length matter, and who should operate the infrastructure underneath it.

Does the workload need a GPU?

Only one of the six, Modal, offers GPU access — A100 and H100 hardware, as of the comparison's June 2026 update. The other five (E2B, Daytona, Cloudflare Sandbox, Northflank, Vercel Sandbox) have no GPU option. For a CPU-bound workload, GPU access isn't a differentiator and the next two questions matter more.

How much do cold start and session length matter?

E2B (Firecracker microVM, ~150 ms) and Daytona (Sysbox containers, under 90 ms via pre-warmed pools) suit a workload spinning up many short-lived sandboxes per minute. Northflank's unlimited sessions and Daytona's persistent, stateful workspace suit the opposite case — a long-running interactive sandbox that shouldn't reset between steps. Vercel Sandbox's fixed 45-minute-to-5-hour cap sits between the two, ruling it out for anything needing to run longer uninterrupted.

Who should operate the isolation infrastructure?

A hosted API means the vendor operates the Firecracker, gVisor, or Kata fleet underneath it, including host-kernel patching and isolation engineering. Self-hosting the same layers gives full control over network topology and compliance boundary at the cost of owning that burden; Northflank's bring-your-own-cloud option, deploying onto a team's own AWS, GCP, or Azure account, is a documented middle path. None of the six vendors publishes usage pricing in this reference — Daytona's $24M Series A in February 2026 is a funding fact, not a cost figure — so treat a cost comparison between them as unverified.

A provider's own built-in tool is a seventh path that skips vendor selection: OpenAI's Code Interpreter and Anthropic's code execution tool run inside the provider's own sandbox. The pillar covers their tool versions and tradeoffs (/articles/agent-security-operations) — the same three questions above still decide whether that built-in isolation is enough or the task needs one of the six dedicated APIs instead.

What does hardening a sandbox harness actually require?

Hardening a sandbox harness requires implementing and verifying six controls mechanically, not just listing them, because an untested policy behaves the same as no policy the first time a sandboxed process tries to exceed it.

A sandbox decision checklist

A sandbox choice for agent code execution clears four checks before it is ready for production. The isolation layer must match the task's trust-and-blast-radius quadrant. The vendor or self-hosted setup must match the workload's GPU, cold-start, and session-length needs. Every harness-hardening control above must be implemented and tested, not assumed. And no real secret should ever sit inside the sandbox, regardless of how strong its isolation layer already is.

This decision sits inside the full eight-discipline stack the agent security operations pillar covers, alongside managing secrets for AI agents and defending an agent against prompt injection.

Frequently asked questions

How do you decide which isolation layer to use for an AI agent's code execution?
Deciding which isolation layer to use starts from two variables crossed against each other — where the code being executed came from (fully untrusted, adversarial model output versus semi-trusted, reviewed code) and what a break-out could reach (a low-value ephemeral environment versus a high-value environment holding real secrets or persistent data) — with fully untrusted code aimed at a high-value target requiring at minimum a microVM such as Firecracker or Cloud Hypervisor, never a plain Docker container.
Which hosted AI agent sandbox APIs offer GPU access?
As of June 2026, only Modal offers GPU access among the six hosted agent-sandbox APIs in the code-execution-sandboxing comparison — E2B, Modal, Daytona, Cloudflare Sandbox, Northflank, and Vercel Sandbox — with A100 and H100 hardware available; the other five have no GPU option, so a workload that needs GPU compute inside the sandbox itself has exactly one match in that comparison.
Should a team self-host sandbox infrastructure or use a hosted agent-sandbox API?
Self-hosting sandbox infrastructure such as Firecracker, gVisor, or Kata Containers gives a team full control over network topology and compliance boundary at the cost of owning the operational burden of patching host kernels and running multi-tenant isolation directly, while a hosted agent-sandbox API shifts that operational burden to the vendor; Northflank's bring-your-own-cloud option, deploying onto a team's own AWS, GCP, or Azure account, sits as a documented middle path, and no usage-pricing data exists in the source comparison to settle the choice on cost alone.
What counts as sufficient network egress hardening for an agent's code sandbox?
Sufficient network egress hardening means outbound traffic is denied by default at the sandbox's firewall or proxy layer with only specific, legitimately needed hosts allowlisted, and that block is confirmed by actually attempting a connection to an unlisted host from inside a running sandbox rather than by checking that a policy document describes the intended behavior.

#agents #security #sandboxing #code-execution #microvm #gvisor #firecracker

Put this corpus inside your own agents

This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.

Agents: this page as Markdown · JSON · offers at /api/pricing.json · payment methods at /api/payment.json · single-resource access via HTTP 402 (how that works)