How to sandbox AI agents that write and run their own code
Updated

Generated code means every agent run is partly untrusted: isolate it in microVM or gVisor sandboxes and fence its network with gateway egress policies.
Every AI agent demo shows the magic moment: the agent writes code and runs it. Every security review asks the follow-up question: runs it where, exactly? Generated code is untrusted code by definition, because no human read it before execution. Sandboxing answers the where with boundaries an agent cannot talk its way past.
The threat in one paragraph
An agent with code execution and network access can do anything its machine can do: read mounted secrets, rewrite shared files, exfiltrate data to an outside server, or pivot to neighboring workloads. Prompt injection makes this worse, because the instruction to misbehave can arrive inside normal-looking tool output. Policy in the prompt is a suggestion. Isolation in the runtime is a fact.
Isolation options, weakest to strongest
Plain containers. Namespaces and cgroups fence processes cheaply, but everyone shares the host kernel. Fine for trusted code, not enough for generated commands.
gVisor-style syscall filtering. A user-space kernel sits between the agent and the host and answers system calls itself. The attack surface shrinks to the filter instead of the full kernel. Good balance of density and safety for most agent fleets.
MicroVMs per agent. Each agent gets its own tiny virtual machine with hardware-enforced boundaries. Startup costs more and memory overhead rises, but a kernel bug in one guest stays in that guest. Early Substrate coverage points at Cloud Hypervisor class runtimes for exactly this tier.
Pick per workload, not per religion. A read-only research agent and a shell-wielding ops agent deserve different boxes.
The network fence: gateway egress policy
Isolation contains the blast. Egress policy shrinks what there is to blast toward. A gateway lists allowed destinations: the ticket API, the documentation server, the internal package mirror. Everything else, including the open internet, is denied and logged.
Write the allowlist from the task, not from convenience. If the triage agent only needs tickets and docs, it gets two destinations, not twenty. Review the list quarterly the way you review firewall rules, because agents accumulate permissions exactly like people do.
Secrets: never in the prompt
Agents love to paste things. Model keys, database passwords and tokens end up in prompts, files and logs unless the platform prevents it. Hold keys in the central Model configuration, inject them at call time, and scope each key to the minimum it needs. An agent that never sees the raw key cannot leak it in a log line.
Worked example: a fictional code-fix agent with teeth
The context below is fictional. Fictional team ParcelTrack (fictional) runs an agent that reads bug reports, writes patches and runs the test suite. One week a report contains hidden instructions telling the agent to upload the repo to an outside server.
In the contained design the agent runs in a per-task microVM, the gateway allows only the Git server and the package mirror, and the model key is injected, never visible. The injected instruction executes, tries the upload, and hits the denied egress rule. The attempt lands in the gateway log, the reviewer sees it, and the patch still gets written. Same agent, same prompt attack, two different endings decided entirely by the fence.
Checklist: a sandbox review that fits on one page
- Isolation boundary named per workload: container, gVisor tier or microVM.
- Egress allowlist with a reason per entry, default deny, all denials logged.
- Secrets held centrally and injected, none visible to the agent process.
- Workspace is disposable: repos re-cloned, no production credentials mounted.
- Kill path tested: suspend and destroy a misbehaving task in under a minute.
Related reading
- The full two-layer picture is in Google AX and Agent Substrate explained.
- Container isolation habits transfer directly from small safe container images.
- Permission discipline for service accounts is the same muscle as permissions broke my service.
Straight answers
Frequently asked questions
Why do AI agents need sandboxing?
They generate and execute code, including terminal commands nobody reviewed. Without isolation one bad command can read secrets, corrupt data or attack neighbors.
What is gVisor sandboxing?
A user-space kernel that intercepts an application's system calls, shrinking what a compromised process can reach on the host.
What is a microVM for agents?
A tiny hardware-isolated virtual machine per agent, often via Cloud Hypervisor class runtimes, giving stronger boundaries than plain containers.
Containers versus microVMs for agent isolation?
Containers share the host kernel, so a kernel escape affects everyone. MicroVMs add a hardware boundary per agent at some startup and memory cost.
What is a gateway egress policy?
A declared list of which outside hosts and APIs an agent may contact. Everything else is denied, so stolen credentials and prompt-injected instructions hit a wall.
Should agents have open internet access?
Default to no. Allowlist the ticket API, the docs server and the package mirror the task needs, and log every allowed call.
Where do model API keys live?
Centrally in the Model configuration, injected at runtime, never pasted into prompts, files or chat logs.
How do I review an agent sandbox setup?
Check four things: the isolation boundary, the egress allowlist, the secret handling and the kill path. If any is missing, the setup is not done.
Do sandboxes slow agents down?
Slightly at startup and on system-call-heavy work. Measure the overhead on your workload; it is usually far smaller than one incident.