Guide

Best Practices for Open AI Sandbox Security

When you run an AI agent locally or in the cloud, isolation is never absolute. Here is how to configure a secure sandbox for OpenAI-powered tools and protect your host system.

Quick Answer

OpenAI does not sell a product called 'Open AI Sandbox.' Developers typically mean one of two things: the sandbox that OpenAI's own Codex agent runs commands in, on your machine or in OpenAI's cloud, or a sandbox you build yourself to run code safely next to the OpenAI API. Codex restricts a session to a read-only, workspace-write, or danger-full-access mode and blocks network access by default. A sandbox you build yourself needs the same discipline: strict network egress filtering, dropped Linux container capabilities, and no plain-text secrets in the environment where the agent executes its code.

What OpenAI actually ships for sandboxing

OpenAI does not sell a product named 'Open AI Sandbox.' The closest real things are Codex, OpenAI's own coding agent, and the code_interpreter tool in the Responses API. Both run code in a sandbox, but they are not the same sandbox and neither is configured through a single setting called 'sandbox'.

Codex runs every command inside one of three modes: read-only, workspace-write (the default), or danger-full-access. It enforces this with the operating system's own sandboxing: Seatbelt on macOS, bubblewrap on Linux and WSL2, and Windows Sandbox or a WSL2 Linux sandbox on Windows. An approval policy sits on top, on-request by default, so an action outside the current mode asks you before it runs. Codex Cloud adds a second boundary: a setup phase that can reach the network to install dependencies, then an agent phase that runs offline by default unless you turn on internet access for that environment, with .git and .codex or .agents folders kept read-only throughout.

The code_interpreter tool, available through the Responses API and the Chat Completions API, is a different thing again. OpenAI runs your model's Python in a container it hosts, not one you configure yourself. You pick a memory tier (1 GB by default, up to 64 GB) and the container is torn down after 20 minutes of inactivity. This replaced the Assistants API's own Code Interpreter, which OpenAI retired in 2026, so a guide that still points you at the Assistants API for sandboxed execution is describing a product that no longer exists.

  • Codex sandbox modes: read-only, workspace-write (default), danger-full-access.
  • Codex enforces this with Seatbelt on macOS, bubblewrap on Linux and WSL2, and Windows Sandbox on Windows.
  • Codex Cloud's agent phase runs offline by default; only the setup phase can reach the network.
  • The Responses API's code_interpreter tool runs in an OpenAI-hosted container, not a setting you configure.

Network isolation for agent workspaces

A secure sandbox AI system requires strict network isolation. AI agents often need to download packages via npm, pip, or cargo to test the code they write, but they should never have unrestricted access to the open internet or your local corporate subnet. The most common escape vector for a coding agent is using its execution environment to exfiltrate environment variables, private source code, or internal API keys to an external server. Once an agent has arbitrary execution rights, it can curl or wget any payload it desires if the network is left open.

For local development on macOS, developers can enforce strict egress filtering using Little Snitch, a host-based application firewall. You can configure a profile that denies all outbound connections from your terminal emulator, your code editor, or your Docker daemon, explicitly allowing only the specific domains required for package management (such as registry.npmjs.org or pypi.org). For Linux environments, standard tools like iptables or ufw must be used to drop all outbound traffic from the specific user or container running the agent.

This matters even for agents that still need the OpenAI API. Keep the execution environment offline and let only the orchestration logic talk to the API over the network. That split means that even if the agent attempts a reverse shell or tries to exfiltrate data, the packets get dropped at the host boundary, because the execution space and the API communication space are never on the same network.

  • Isolate the agent from the local subnet and the open internet.
  • Use host-based firewalls like Little Snitch on macOS to control egress traffic.
  • Block outbound traffic with iptables for Linux container environments.

Container isolation and dropped capabilities

If you are executing agent-generated code, you must never run it directly on your host machine or your primary development laptop. Docker is the industry standard for creating isolated environments, but a default Docker container is absolutely not a secure sandbox AI system. By default, Docker containers run as the root user and retain several Linux kernel capabilities that can be exploited by an agent to break out of the container and access the host operating system.

To harden a Docker-based sandbox, you must drop all Linux capabilities using the `--cap-drop=ALL` flag and run the process as an unprivileged, non-root user. You should also enable the `--security-opt no-new-privileges` flag to prevent privilege escalation within the container, and use `--network none` if the task does not require fetching external dependencies. If you require cloud-based isolation, specialized infrastructure providers like E2B offer secure, firewalled sandboxes built on AWS Firecracker microVMs.

These Firecracker microVMs provide hardware-level virtualization, ensuring that even if an agent escapes the container namespace or exploits a zero-day in the container runtime, it remains trapped within a strict, virtualized hardware boundary that prevents access to the underlying hypervisor or neighboring virtual machines.

  • Never run agent-generated code directly on the host machine.
  • Drop all Linux kernel capabilities when using Docker containers.
  • Consider Firecracker microVMs for hardware-level isolation in the cloud.

Secret management and vault configuration

Agents require API keys, database credentials, and cloud tokens to perform meaningful work and interact with external services. However, storing these secrets in plain `.env` files within the sandbox is a critical vulnerability. OWASP's 2026 Top 10 for LLM Applications ranks prompt injection first and excessive agency third, and a coding agent sandbox runs straight into both: a hidden instruction in a codebase or an issue ticket can trick the agent into reading the `.env` file and printing your production database credentials into its output, and an agent holding more permissions than the task needs turns that one trick into a much bigger incident.

The most reliable fix is to keep secrets completely outside the folder the agent works in and hand them over at the exact moment a command runs, so the value never sits in a file the agent can open. Use short-lived, least-privilege tokens whenever possible. For example, AWS STS (Security Token Service) should be used to generate temporary, time-bound credentials rather than providing the agent with long-lived IAM access keys.

A key that exists in memory for the length of one command is a much smaller target than a key that sits in a file on disk for months. By injecting credentials directly into the execution context and rotating them frequently, you dramatically limit the blast radius if an agent ever goes rogue or falls victim to a prompt injection attack.

  • Do not store production secrets in .env files within the agent's workspace.
  • Inject secrets directly into the commands that need them at runtime.
  • Use short-lived, narrowly scoped tokens to limit the blast radius of a leak.

File system boundaries and access control

A critical aspect of a sandbox AI system is tight file system isolation. The agent must only have access to the specific project directory it is modifying, and nothing else. Allowing an autonomous agent broad access to your home directory exposes your SSH keys (`~/.ssh`), cloud credentials (`~/.aws`), browser profiles, and personal documents to unauthorized reading or modification.

On Linux systems, you can use user namespaces and `chroot` jails to restrict the agent's view of the filesystem, ensuring it cannot traverse upwards into the host's sensitive directories. On macOS, the kernel sandbox (known as Seatbelt, the same mechanism Codex and Claude Code's own /sandbox use) enforces file access restrictions on a per-process basis. Without explicit boundaries, an agent has the exact same filesystem access that you have, which reinforces the need to explicitly define limits before starting a session.

Always ensure that the agent operates within a dedicated, restricted workspace folder. You must verify that it cannot read files outside its designated root, especially considering that agents actively scan their environment to understand the context of the codebase they are asked to modify.

  • Restrict the agent's file access to its specific workspace folder.
  • Prevent traversal into sensitive directories like ~/.ssh or ~/.aws.
  • Use kernel-level sandboxing features to enforce file system boundaries.

How Forkbench handles local boundaries

Forkbench is a desktop app that runs coding agents in a terminal on your Mac. Its Vault keeps secrets in your Keychain and lets an agent use a secret by name, so the agent can run a command that needs a token without ever being shown the token itself. However, it is vital to understand the limits of this protection: an unpinned Vault key can still be read by the program that was run. If you grant an agent permission to execute arbitrary scripts, those scripts can read the injected secrets from memory.

A Forkbench Thread can be locked to its folders by the macOS kernel sandbox, meaning `~/.ssh`, other repositories, and the Documents folder stay completely shut off from the agent's view. This folder lock is opt-in and does not restrict the network. Files that you leave in the project folder are still readable by an agent, so you must still practice good hygiene by keeping sensitive files out of the workspace.

Forkbench governs what it holds and the folders you lock, not the whole machine. You are still fully responsible for configuring network firewalls, installing tools like Little Snitch, and managing the secrets you leave lying around in the open. It provides the boundaries, but you must enable and enforce them.

  • Forkbench uses the macOS kernel sandbox to lock agents to specific folders.
  • The folder lock is opt-in and does not restrict outbound network access.
  • An unpinned Vault key can still be read by the program that was executed.

Human-in-the-loop safeguards

No sandbox AI system is perfectly impenetrable. Zero-day vulnerabilities in container runtimes, microVM hypervisors, or the Linux kernel can be and have been exploited. Therefore, the ultimate safeguard for an OpenAI-powered agent is a strict human-in-the-loop (HITL) architecture. You should never fully automate the deployment pipeline when AI agents are generating the code.

The agent should be allowed to draft code, run unit tests within the isolated sandbox, and propose changes via pull requests, but it must never be allowed to deploy to production or merge into the main branch without explicit human approval. Human review remains the most effective way to catch logic errors, hallucinated dependencies, and potential security flaws introduced by the AI.

Implementing a mandatory review process ensures that any malicious or hallucinated code generated by the agent is caught before it affects your wider infrastructure. By treating all agent-generated output as untrusted until verified, you add a final, non-technical layer of defense that cannot be bypassed by a clever prompt injection or a sandbox escape.

  • Require human approval before merging or deploying agent-generated code.
  • Treat all code written by an AI agent as untrusted until reviewed.
  • Use a human-in-the-loop workflow as the final line of defense against sandbox escapes.

Related: Forkbench for OpenAI Codex CLI, Codex CLI sandbox modes, and what they don't cover, E2B alternative: sandbox agents on your Mac, How Forkbench handles your data

Frequently asked

  • What does an OpenAI sandbox system actually do?

    It depends which OpenAI product you mean. Codex restricts a command to read-only, workspace-write, or danger-full-access using the operating system's own sandbox, Seatbelt on macOS or bubblewrap on Linux. The Responses API's code_interpreter tool runs Python in a separate, OpenAI-hosted container instead.

  • Can I run Codex or an OpenAI-powered agent's sandbox fully offline?

    You cannot run OpenAI's models offline; they need OpenAI's servers. But the sandbox around them can be offline. Codex Cloud's agent phase runs with no network access by default, after a setup phase that is allowed online just long enough to install dependencies. On your own machine, you can do the same thing with Docker and --network none while only your orchestration script talks to the API.

  • Why shouldn't I use default Docker containers for my sandbox AI system?

    Default Docker containers run as the root user and retain Linux kernel capabilities that can be exploited for container escape. A secure sandbox must drop all capabilities (--cap-drop=ALL), run as a non-root user, and restrict network access.

  • Is OpenAI's Code Interpreter sandbox still part of the Assistants API?

    No. OpenAI retired the Assistants API in 2026. The same sandboxed Python environment is now the code_interpreter tool in the Responses API and the Chat Completions API, running in an OpenAI-hosted container where you pick the memory tier.

  • Does Forkbench guarantee total network security for agents?

    No. While Forkbench allows you to lock an agent to specific folders using the macOS kernel sandbox, this folder lock is opt-in and does not restrict the network. You must still use a host firewall if you wish to block outbound connections.

Keep reading