Guide
Optimizing Your Workflow with the Open AI Sandbox Framework
Searches for an Open AI Sandbox Framework usually mean one of three real things OpenAI ships: Codex's own sandbox, the Agents SDK's sandbox support, or the hosted Agents API. They solve different problems and get configured differently.
OpenAI does not sell a product named the Open AI Sandbox Framework. What it ships instead is three separate things that get called this in search: Codex's built-in execution sandbox (Seatbelt on macOS, bubblewrap plus seccomp on Linux, an isolated container in the cloud), native sandbox support added to the open-source Agents SDK in April 2026 that plugs into hosted providers such as E2B, Daytona and Cloudflare, and the hosted Agents API, which went to public beta on September 10, 2026 and lets you run the same harness that powers Codex against a sandbox OpenAI manages or one of its partners. A workflow built around one of these needs different settings than a workflow built around another.
Three things people mean by this phrase
Nobody at OpenAI calls anything the Open AI Sandbox Framework. The phrase is a stand-in for three separate pieces of real infrastructure, and mixing them up is the fastest way to misconfigure a workflow.
The first is Codex, OpenAI's own coding agent product. Its sandbox is something you use, not something you build with: it is already wired into the CLI and the cloud product, and your job is to set its policy correctly.
The second is the Agents SDK, the open-source library for building a custom agent of your own. Since April 2026 it has shipped native sandbox support, which means it can hand off code execution to a hosted provider instead of you writing that integration yourself.
The third is the Agents API, a managed service OpenAI opened to public beta on September 10, 2026. It runs the same harness that powers Codex on OpenAI's own infrastructure, so you call an API instead of hosting the SDK and the sandbox integration at all.
- Codex: a product, already sandboxed, you configure its policy.
- Agents SDK: a library you self-host, now with native sandbox support.
- Agents API: a managed service that runs the Codex harness for you.
How Codex's own sandbox works
Codex enforces its sandbox at the operating system level. On macOS it uses Seatbelt policies through sandbox-exec, the same kernel mechanism Apple uses for its own app sandboxing. On Linux it uses bubblewrap plus seccomp. On Windows it uses the Microsoft eXecution Container sandbox, with a legacy fallback, or you can run it inside WSL2 and get the Linux implementation.
Network access is off by default and stays off even in the permissive workspace-write mode. Turning it on for a task means opting into an allowlist: an exact hostname matches only itself, a leading *.example.com matches subdomains but not the bare domain, and **.example.com matches both. Local and private network destinations are blocked unless you name them explicitly, and the allowlist checks what a hostname actually resolves to, so a DNS trick cannot smuggle a request past it.
Codex's cloud product adds a second layer. A task runs in an OpenAI-managed container in two phases: a setup phase that can reach the network to install whatever dependencies you declared, and an agent phase that runs with no internet access by default unless you turn it on for that specific environment. Any secret you configured for the cloud environment is available only during setup and is gone before the agent phase starts, so the code that is actually iterating on your task never had it.
- macOS: Seatbelt. Linux: bubblewrap plus seccomp. Windows: MXC sandbox or WSL2.
- Network access defaults to off, with a domain allowlist syntax when you turn it on.
- Cloud Codex splits setup (networked, has secrets) from the agent phase (offline by default, secrets removed).
The Agents SDK's native sandbox support
If you are building your own agent rather than using Codex directly, the relevant change landed in the Agents SDK in April 2026. Before that, running an agent's generated code somewhere safe meant you wired up a sandbox provider's API by hand. Since then, the SDK ships a sandbox agent primitive that talks to a provider directly.
Seven hosted providers are officially integrated at the time of writing: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop and Vercel. On top of those, the SDK also supports Docker and a local Unix sandbox client, for when you want the same code to run on your own box without a hosted account at all. Python got this first in mid-April 2026; the TypeScript SDK reached the same feature set in version 0.9.1, released in May 2026, so the gap between the two languages has closed.
What you get for picking one of the seven over hand-rolling your own integration is less code, not more security by itself: each provider still has its own isolation model (a microVM, a gVisor container, a full Linux VM) and its own network and persistence defaults, which is the subject worth comparing before you commit to one.
- Native sandbox support in the Agents SDK shipped in April 2026, for Python first.
- Seven hosted providers are officially integrated: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, Vercel.
- Docker and a local Unix client are also supported, for running without a hosted provider.
- TypeScript reached the same feature set in SDK version 0.9.1, released May 2026.
The Agents API, for teams who would rather not self-host
The Agents API is a different answer to the same problem. Instead of running the SDK yourself, you call an API and OpenAI hosts and operates the harness, the same one that drives Codex, on its own infrastructure. You still choose where the agent's code actually executes: an OpenAI-managed sandbox, your own infrastructure, or a partner environment such as Cloudflare, Modal or Vercel.
OpenAI handles session management, orchestration and context compaction for you, which is the part that usually takes the most engineering time to get right in a long-running agent. There is no separate charge for the API itself; you pay standard rates for the model, for any built-in tools you use, and for the sandbox if you pick an OpenAI-managed one.
This is the option worth evaluating first if your enterprise constraint is engineering time rather than infrastructure control. The Agents SDK gives you more control over exactly how execution happens; the Agents API trades some of that control for not having to operate the harness yourself.
- The Agents API reached public beta on September 10, 2026.
- It runs the Codex harness on OpenAI's infrastructure; you still pick where code executes.
- Pricing is standard model and tool rates, plus container rates for an OpenAI-managed sandbox.
What a sandbox does not solve for you
Whichever of the three you pick, the sandbox answers one question: where does untrusted code run so a mistake cannot reach your real systems. It does not answer what you do with the output once the sandbox hands it back to you.
Developers often assume the boundary means the result is also safe to trust. A sandbox stops a bad command from touching your host. It does nothing once you copy the agent's generated script out of the sandbox and run it somewhere else, or paste its suggested code straight into your codebase without reading it.
State loss is the second common surprise. Several of these environments are ephemeral by design, meaning anything the agent wrote disappears when the session ends unless you explicitly export it. Daytona, E2B and Vercel support longer-lived, resumable sessions; others are built to be thrown away after every run. Check which kind you are using before you plan a workflow around keeping anything inside it.
- A sandbox protects your host. It does not vet the code that comes back out.
- Ephemeral environments discard everything at session end unless you export it first.
- Providers differ on persistence, so check the specific one you picked rather than assuming.
Where Forkbench fits, and where it does not
Everything above is for a specific situation: code that might be adversarial or from an unknown source, running at scale, where you need a dedicated kernel or container per task. If that is your enterprise workflow, a hosted sandbox provider is the right tool, and Forkbench is not a substitute for one.
Forkbench solves a different problem. It is a desktop app that runs your coding agent, Codex included, in a real terminal on your own Mac, for the much more common case of one developer supervising an agent they already trust on a project they already own. Instead of isolating code in a remote container, it gives you a folder lock that keeps the agent to the project directory you approved, though that lock is opt-in and does not restrict the network, so a locked-down agent can still make outbound requests.
For secrets, Forkbench's Vault keeps API keys in your Mac's Keychain and lets the agent use one by name instead of reading the value. That stops a key from sitting in a file the agent can open or in its context, but the protection ends at the command boundary: a program the agent is allowed to run can still read any key it was handed, pinned or not. If your workflow needs to contain a key from the process using it, not just from the agent requesting it, that is a job for the enterprise token-scoping described above, not for a desktop Vault.
- Forkbench supervises a trusted agent on your own Mac; it does not replace a multi-tenant hosted sandbox.
- The folder lock is opt-in and does not restrict the network.
- The Vault hides a key's value from the agent, but a program the agent runs can still read it.
Related: Codex CLI sandbox modes, explained, Forkbench for Codex, How Forkbench handles your data, Download Forkbench
Frequently asked
Is there really a product called the Open AI Sandbox Framework?
No. OpenAI ships three separate things people describe this way: Codex's own built-in sandbox, native sandbox support added to the open-source Agents SDK in April 2026, and the hosted Agents API that reached public beta in September 2026.
What sandbox technology does Codex use?
Locally, Seatbelt on macOS, bubblewrap plus seccomp on Linux, and the Microsoft eXecution Container sandbox on Windows (or WSL2 for the Linux implementation). In the cloud, Codex runs tasks in isolated, OpenAI-managed containers with network access off by default during the agent phase.
Which hosted sandboxes work with the OpenAI Agents SDK?
Seven providers are officially integrated: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop and Vercel. The SDK also supports Docker and a local Unix sandbox client for running without a hosted provider.
What is the difference between the Agents SDK and the Agents API?
The SDK is an open-source library you host and run yourself. The Agents API, in public beta since September 2026, runs the same underlying harness as a managed service on OpenAI's infrastructure, so you call an API instead of operating the harness and the sandbox integration.
Can Forkbench replace a hosted sandbox provider for enterprise AI code execution?
No. Forkbench supervises a trusted agent working in a terminal on your own Mac; it is not built to contain adversarial, multi-tenant code the way E2B, Daytona or a similar hosted provider is. Use a hosted sandbox for that job and Forkbench for day-to-day supervision of an agent you already trust.