Guide

Designing a Secure Dev Sandbox for AI Agent Execution

Running autonomous code requires a bounded environment. Here is how to build a safe workspace for your AI agents, isolate execution, and protect your secrets.

Quick Answer

Designing a sandbox for an AI coding agent means isolating three things: its file system, its network, and its secrets. For basic hosted execution, OpenAI's own Agents API now ships hosted sandboxes you can use without building anything. For a custom agent or desktop app, you will usually build your own boundary instead, using a container runtime like Docker, a microVM like Firecracker, or a managed platform like E2B that runs Firecracker for you. Whichever you pick, keep secrets in a vault outside the agent's reach and inject them only when a specific command needs them.

Why an AI coding agent needs a sandbox

When an AI agent writes and executes code, it acts with the permissions of the environment it runs in. If you run an agent directly on your laptop without boundaries, a mistake in its generated code can overwrite your personal files, modify your system configuration, or expose your environment variables to a remote server.

A sandbox creates a controlled space where the agent can explore, install dependencies, and run scripts without touching the host operating system. This is critical for developers because autonomous agents learn by trying things, and trying things means executing unknown, unverified code repeatedly until a test passes.

The goal is not to stop the agent from working, but to limit the blast radius when it makes a mistake. A well-designed sandbox gives an agent enough freedom to compile code, run test suites, and call the APIs it actually needs, while strictly blocking access to your private data and the broader internet. Without these boundaries, a simple hallucination could result in a destructive command being executed on your local machine.

  • Agents run with the permissions of their host environment.
  • Untrusted code execution is a normal part of how agents solve problems.
  • A sandbox limits the blast radius of destructive commands.
  • Developers need a workspace where agents can install packages safely.

Choosing the right sandbox architecture

There are several ways to build a sandbox, each with different trade-offs between performance and security. The simplest approach is using a Docker container. Docker provides process isolation, a separate file system, and resource limits, making it a popular choice for running temporary agent workloads locally.

However, Docker shares the host kernel, which can be a security risk if the agent generates highly untrusted or intentionally malicious code that exploits kernel vulnerabilities. For stronger isolation, developers turn to microVMs like Firecracker. Built by AWS and open sourced to power AWS Lambda, Firecracker gives each sandbox its own kernel while still booting in under 125 milliseconds, so the isolation costs almost nothing in startup time.

If you do not want to manage this infrastructure yourself, cloud platforms like E2B provide managed, ephemeral sandboxes built on Firecracker specifically designed for AI agents. These cloud solutions offer SDKs that allow you to spin up a secure environment, execute Python or Node.js code, and retrieve the results via an API, offloading the complexity of container management.

  • Docker offers lightweight process isolation and file system boundaries.
  • Firecracker microVMs provide hardware-level isolation for running untrusted code.
  • Platforms like E2B offer managed ephemeral environments for agents.
  • Cloud sandboxes offload the complexity of local container management.

Using hosted environments and OpenAI sandboxes

Many developers start by looking for a sandbox straight from the model vendor instead of building one. OpenAI's Agents API now ships exactly that: an OpenAI-hosted sandbox gives an agent a Linux workspace at /workspace with Python, Node.js and common command-line tools already installed, and OpenAI provisions and connects it for you.

These hosted sandboxes are convenient because you supply the task and get results back without running any infrastructure yourself. You can still shape the environment somewhat: config lists let you install specific Python, npm or system packages and run setup commands before the agent starts, and files persist across turns for as long as that session's sandbox exists.

The limit is real but narrower than it looks. You cannot swap in a fully custom Docker image, pick your own compute, or reach a private corporate network from an OpenAI-hosted sandbox. OpenAI's own answer for that is a self-hosted sandbox, where you run the executor yourself. If you are building an independent desktop app or a custom agent framework, you are usually in that self-hosted camp anyway, choosing between Docker, Firecracker or a managed platform like E2B rather than the hosted option.

  • OpenAI's Agents API ships hosted sandboxes with Python, Node.js and CLI tools preinstalled.
  • You can still install specific packages and run setup commands before the agent starts.
  • A custom Docker image, custom compute, or a private network needs the self-hosted option.
  • Custom agent frameworks usually end up choosing Docker, Firecracker or E2B instead.

Securing secrets with a vault

A sandbox keeps the agent away from your host files, but the agent still needs API keys to interact with external services like GitHub, AWS, or Stripe. If you place a .env file inside the sandbox or export variables globally, the agent can read them, memorize them, and potentially leak them in the chat transcript.

The fix is a vault that sits outside the sandbox entirely: HashiCorp Vault, AWS Secrets Manager, or the macOS Keychain on a desktop setup, any of which can hold the real token values instead of a file inside the agent's workspace.

The sandbox then injects the secret only when a specific, authorized command runs. If the agent runs a deployment script, the environment fetches the AWS credentials from the vault, hands them to that one script, and drops them immediately after. The key exists in memory for the length of one command and never rests in a file the agent can read at its leisure.

  • Do not store .env files inside the sandbox.
  • Agents can accidentally leak credentials they read into their transcripts.
  • Use a vault to keep secrets outside the agent's direct file system.
  • Inject secrets dynamically only when a specific command runs.

Network boundaries and egress filtering

By default, an isolated Docker container or microVM can still make outbound network requests to the public internet. If your agent is running untrusted code, this open network access allows it to download malicious payloads, participate in botnets, or exfiltrate sensitive data to a remote server.

A secure environment requires strict egress filtering. You can use iptables firewall rules, network namespaces, or a service mesh to block all outbound traffic by default. You then open only specific ports and domains that the agent explicitly needs to function, such as package registries like npmjs.com or pypi.org.

Keep in mind that heavily restricting the network can break agents that rely on web scraping, downloading datasets, or calling external APIs to gather context. You must carefully balance strict security policies with the agent's ability to complete its assigned tasks efficiently.

  • Containers typically allow outbound network requests by default.
  • Untrusted code can download malware or exfiltrate data.
  • Use firewall rules to restrict outbound traffic to necessary domains.
  • Balancing security with functionality is required for capable agents.

Managing dependencies and persistent state

Agents frequently install third-party libraries to finish a task, so a sandbox needs to allow that without permanently changing its base image. Overlay file systems handle this by writing changes to a temporary upper layer that disappears with the sandbox.

Make the sandbox itself disposable: create it fresh for every task and destroy it right after, so a corrupted config or a bad install never carries over to the next run. If an agent genuinely needs state across sessions, such as a cache of downloaded files, mount one isolated volume for that and still discard everything else.

  • Use overlay file systems to allow temporary package installations.
  • Ephemeral environments should be destroyed after every session.
  • Destroying the sandbox prevents state corruption and malware persistence.
  • Use isolated mounted volumes for any data that must be saved.

How Forkbench implements boundaries

Forkbench is a desktop application that runs coding agents in a terminal on your Mac. It uses a Vault mechanism to keep your secrets safely in the macOS Keychain. When an agent needs a token, it references the secret by name, so the command runs without the agent ever reading the underlying value in a file.

It is crucial to understand the limits of this approach. An unpinned Vault key can be read by the program that ran. This means if the agent runs a script that explicitly prints the environment variables, the value will be exposed in the terminal output. Forkbench also includes a folder lock feature to prevent the agent from touching files outside the designated project directory.

However, this folder lock is strictly opt-in and does not restrict the network. The agent can still make outbound internet requests unless you manually configure your own firewall rules outside of Forkbench. The tool is designed primarily to protect your host files and secrets from accidental reads, not to contain deliberately malicious software.

  • Forkbench Vault keeps secrets in the macOS Keychain.
  • Agents use secrets by name without seeing the actual token.
  • An unpinned Vault key can be read by the program that ran.
  • The folder lock is opt-in and does not restrict network access.

Best practices for building your environment

Start by defining the exact permissions your agent actually needs to do its job. Apply the principle of least privilege. If the agent only needs to format text or perform static analysis, it does not need network access. If it needs to build a web application, it needs access to package managers and a compiler, but it probably does not need root privileges.

Regularly audit the tools and commands your agent uses. Monitor what the agent is doing by logging all executed commands and capturing network requests. This gives you a clear audit trail if something goes wrong and helps you refine your security boundaries over time.

Finally, assume that the sandbox will eventually be breached. Defense in depth is required. Combine process isolation, network filtering, secret vaults, and strict monitoring to ensure that a failure in one layer does not lead to a complete system compromise.

  • Apply the principle of least privilege to all agent permissions.
  • Do not run agents as the root user inside the sandbox.
  • Log all executed commands and network activity for auditing.
  • Use defense in depth to protect against eventual sandbox breaches.

Related: Forkbench vs E2B, Forkbench vs Claude Code's sandbox, Codex CLI sandbox modes explained, Forkbench pricing

Frequently asked

  • What is the best sandbox for developers building AI agents?

    For most developers, Docker is the easiest starting point for process isolation. For production systems running highly untrusted code, microVMs like Firecracker or managed platforms like E2B provide stronger security.

  • Does OpenAI provide a sandbox for developers?

    Yes. OpenAI's Agents API ships an OpenAI-hosted sandbox, a Linux workspace with Python, Node.js and CLI tools preinstalled. It does not give you a custom image or a private network; that needs OpenAI's self-hosted sandbox option instead.

  • How does a vault protect my API keys from an agent?

    A vault stores the key outside the agent's file system in a secure location. The agent requests the key by name, and the system injects it directly into the running command, ensuring the token is never saved in a readable file.

  • Can Forkbench stop an agent from accessing the internet?

    No. The Forkbench folder lock is opt-in and restricts file system access, but it does not restrict the network. Agents can still make outbound requests unless you configure external firewall rules.

Keep reading