Guide
Safeguarding AI Code: The Role of Sandbox Environments
An AI agent that writes and runs its own code can run a bad command as easily as a good one. A sandbox is how you make that safe to find out, one tutorial step at a time.
Safeguarding AI-generated code means giving it a place to run where a mistake cannot reach your real files, your secrets, or your network. For most developers that place starts as a locked-down Docker container: no network, no root user, a read-only filesystem, and hard limits on memory and CPU. For stronger isolation, a microVM platform such as E2B or Modal gives the code its own kernel instead of sharing yours. A newer option, AWS's open source Strands Box, takes a different angle: instead of only walling off where code runs, it watches what an agent does across a session and can block an action based on what it already did. None of these replace reviewing the diff before you trust the result.
Why code an AI agent writes needs its own space to run
When a coding agent can execute its own code, it is not just suggesting a change for you to read. It is running a command, and that command has whatever permissions the agent's process has. If the agent misreads an instruction or copies a broken example from training data, the result is not a red squiggly line. It is a command that actually runs.
On a normal developer machine, that command runs as you. It can overwrite a file outside the project, read a credential sitting in your home folder, or try to reach a server you never intended to contact. None of this requires the agent to be malicious. A model under-specifying a shell command is enough.
A sandbox is the answer to that risk, not a ban on letting the agent run code at all. It gives the agent a space where it can write, install packages, and run tests, while the parts of your machine that matter stay out of reach.
- A command an agent runs has the permissions of the process that runs it.
- A model does not need to be malicious to issue a destructive command, only wrong.
- The goal of a sandbox is to keep the agent useful while keeping the damage contained.
Containers, microVMs, and the layer above both
A Docker container isolates a process using Linux namespaces and control groups, but it still shares the host's kernel. That is fast and cheap, and good enough for code you wrote yourself and trust in general. It is a thinner wall if the code is fully untrusted, because a kernel bug reachable from inside the container reaches the host too.
A microVM, such as the Firecracker technology AWS built and open sourced, gives each sandbox its own dedicated guest kernel instead of a shared one. E2B runs on Firecracker with cold starts under 200 milliseconds and a free Hobby tier capped at one-hour sessions, with a Pro tier at 150 US dollars a month for 24-hour sessions. Modal isolates its Sandboxes product with gVisor instead, Google's user-space kernel intercepting syscalls, and advertises sub-second starts at very high concurrency.
Isolation like this answers one question: what can the code touch? It does not answer a second question that matters just as much for an agent working over a long session: should this particular action be allowed right now, given what the agent already did a minute ago? That second question is what a policy engine is for.
- Docker containers share the host kernel; fast, but a weaker wall for fully untrusted code.
- Firecracker microVMs give each sandbox its own kernel; E2B's Hobby tier is free, Pro is $150/mo.
- Modal's Sandboxes use gVisor instead of a microVM and target very high concurrency.
- Isolation controls where code runs. It does not decide whether a specific action should be allowed.
A Docker sandbox you can set up right now
The default Docker configuration is too permissive for code you did not write yourself. A command like `docker run --network none --read-only --cap-drop ALL --security-opt no-new-privileges:true --user 65534:65534 my-agent-image` removes network access, mounts the filesystem read-only, drops every Linux capability, and runs as an unprivileged user instead of root.
On Linux, Docker already applies the docker-default AppArmor profile to a new container unless you turn it off. That profile denies writes to most of /proc, blocks mount operations, and restricts ptrace to processes under the same profile, even if the container was granted a capability that would otherwise allow it. Pairing that with seccomp, which filters the specific syscalls a process may make, closes most of the gap a bare container leaves open.
Resource limits matter just as much as access control. An agent that writes an infinite loop or a fork bomb can take down the host if nothing stops it. Flags like `--memory 512m`, `--cpus 1.0`, and `--pids-limit 100` cap what a single runaway process can consume before Docker kills it.
- Strip privileges: `--network none --read-only --cap-drop ALL --user 65534:65534`.
- Docker's docker-default AppArmor profile already blocks mount and most of /proc by default on Linux.
- Add seccomp for syscall filtering on top of AppArmor, not instead of it.
- Cap memory, CPU and process count so one runaway script cannot take down the host.
Strands Box: a brand-new way to police what an agent does
On October 7, 2026, the Strands Agents team at AWS released Strands Box, an open source sandbox built specifically for this problem. Instead of only drawing a wall around a filesystem and a network interface, Box routes every action an agent tries to take through a policy engine called Dogwood, also open source and also built by AWS.
Dogwood can reason about a session's history, not just the single action in front of it. A policy can let an agent post to a chat channel, but cap it at three posts in ten minutes, so a loop that would otherwise spam the channel gets throttled without the agent needing to remember the limit itself. A file read through a shell command can also tighten what that same session is allowed to do afterward on the network, because the policy sees both.
Strands Box checks actions routed through its own shell, its Python interpreter, and its Model Context Protocol broker, all against that same Dogwood engine and event history. This is a different kind of safeguard than a container or a microVM. It assumes the agent is going to be let inside the walls, and focuses on what it is allowed to do once it's there.
- Strands Box (AWS, released October 7, 2026) is an open source sandbox with OS-level isolation plus a policy layer.
- Dogwood, its policy engine, can restrict an action based on what the agent already did earlier in the session.
- Example: cap how often an agent can post to a channel, independent of the agent remembering the limit.
- Strands Box checks its own shell, its Python interpreter, and its MCP broker against the same policy engine.
Keeping secrets out of the sandbox entirely
A sandbox that contains a plaintext `.env` file has not solved the credential problem, it has just moved it. An agent running inside the sandbox can still read that file, memorize the value, and print it to a log the moment something goes wrong.
The stronger pattern keeps the secret outside the sandbox altogether. LangSmith Sandboxes, LangChain's own product for this, uses what it calls an Auth Proxy: requests to an external API go through a proxy that injects the real credential into the outgoing request at the network layer. The sandbox itself never holds the value, so there's nothing for the agent's process, a dependency, or a printed debug log to leak.
You do not need LangChain's specific product to copy the idea. Whatever platform you use, prefer a design where the agent's command asks for an action by name and a broker outside the sandbox supplies the credential for that one request, rather than handing the whole secret to the sandboxed process up front.
- A `.env` file inside the sandbox is still readable, printable and leakable by the code running there.
- LangSmith Sandboxes keep credentials out of the sandbox with an Auth Proxy that injects them at the network layer.
- Prefer brokered, per-request credentials over handing a sandboxed process the whole secret up front.
How Forkbench handles it
Forkbench is a desktop app that runs coding agents in a terminal on your Mac. Its Vault keeps secrets in your Keychain, and an agent uses one by name rather than being shown the value, so the secret does not sit in a file or appear in the conversation transcript.
The limit is specific and worth stating plainly. An unpinned Vault key can still be read by the program that was run with it, so a script that deliberately prints its environment will still expose the value in that terminal output. Forkbench also does not stop an agent from reading a plaintext `.env` file you left in the project folder; it protects what you move into the Vault, not files you haven't moved yet.
Forkbench is not a code-execution sandbox in the sense this page has been describing. It does not give the agent a separate kernel the way a container or a microVM does. If you need that stronger boundary for genuinely untrusted code, pair Forkbench's secret handling and supervision with one of the sandboxes above rather than expecting it to replace them.
- Forkbench's Vault keeps secrets in the Keychain; an agent references them by name.
- An unpinned Vault key can still be read by the program that was run with it.
- Plaintext files already in the project folder remain readable; move secrets to the Vault first.
- Forkbench supervises an agent on your Mac; it is not a substitute for a kernel-level sandbox.
A checklist for today
You don't need every tool on this page to start. Fix the open exposures first, then add the stronger boundaries as the work calls for them.
Work through the list in order. The first two remove risk that already exists; the rest stop it from coming back.
- Find every `.env` file, credentials file, and token pasted into a shell profile.
- Rotate anything that has ever sat in a file an agent could open.
- Run untrusted code in a locked-down container at minimum: no network, read-only filesystem, non-root user.
- Add CPU, memory, and process-count limits so a runaway script cannot take down your machine.
- Move secrets into a Keychain, a vault, or a broker, and inject them per command rather than per session.
- Review a session's transcript before sharing it, in case something was printed that shouldn't have been.
Related: How Forkbench handles your data, Forkbench as an E2B alternative, Forkbench vs Docker sandboxes, Download Forkbench
Frequently asked
What is the real difference between a container and a microVM?
A container, such as a Docker container, shares the host's kernel through namespaces and control groups. A microVM, such as Firecracker, gives each sandbox its own dedicated guest kernel. The microVM is the stronger wall for code you don't trust at all.
Is a locked-down Docker container enough to safeguard AI-generated code?
For most day-to-day use, yes, if you strip privileges and add limits: no network, read-only filesystem, non-root user, and caps on memory and CPU. For code you consider genuinely hostile, a microVM or a policy-based sandbox like Strands Box is the stronger choice.
What is Strands Box?
Strands Box is an open source AI agent sandbox AWS released on October 7, 2026. It combines OS-level isolation with a policy engine called Dogwood, which can restrict an action based on what the agent already did earlier in the same session, not just the action in front of it.
How does Forkbench protect my API keys?
Forkbench's Vault stores secrets in your macOS Keychain. An agent references a secret by name, and Forkbench supplies the value to the command at the moment it runs, so it doesn't sit in a file or show up in the conversation. The program that runs can still read the value once it has it.
Can a sandbox stop an agent from misusing a key it's allowed to use?
No. A sandbox controls where code runs and what it can touch. If you've granted an agent a key to deploy or delete something, isolation doesn't stop it from using that key the way you authorized. Scope every key narrowly and keep production credentials out of experimental sessions.