Guide
Evaluating the Best Sandbox for OpenAI Models
Three different things already compete for the title of best OpenAI sandbox: Codex's own sandbox, a hosted provider behind the Agents SDK, and OpenAI's own Agents API. Which one wins depends on what you're actually building.
There is no single best sandbox for OpenAI, only a best fit for what you're building. If you're using Codex, its own built-in sandbox, Seatbelt on macOS, bubblewrap and seccomp on Linux, the Microsoft eXecution Container sandbox or WSL2 on Windows, is already the right answer and needs no extra setup. If you're building a custom agent on the open-source Agents SDK, you choose a hosted provider through its native sandbox support: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, or Vercel, each with a different isolation model and persistence story, or you run Docker or a local Unix sandbox yourself. If you'd rather not host any of that, the Agents API, in public beta since September 10, 2026, runs the same harness as a managed service and lets you pick where the code executes. None of the three is universally best; each wins a different workflow.
There's no single best OpenAI sandbox, only a best fit
Three real options all claim to be the best way to sandbox code around an OpenAI model, and they are not interchangeable. Codex's own sandbox, a hosted provider wired in through the Agents SDK, and the Agents API all solve the same underlying problem, keeping untrusted code away from your real systems, but they solve it for different people.
The right question is not which one wins outright. It's which of four things matters most to you: how much you want to set up, whether state needs to survive between runs, who is on the hook for operating the sandbox, and whether your constraint is infrastructure control or engineering time.
- Setup effort: none for Codex's own sandbox, more for a self-hosted Agents SDK integration, least infrastructure work for the Agents API.
- Persistence: some providers are ephemeral by design, others support resumable sessions.
- Who operates it: you, for Codex and a self-hosted provider; OpenAI, for the Agents API.
- Cost: bundled with your existing Codex or API usage, or a separate per-sandbox bill.
If you're already using Codex, the sandbox is already there
If your workflow is Codex itself, you don't need to evaluate anything. Codex's sandbox runs locally through Seatbelt on macOS or bubblewrap plus seccomp on Linux, and through the Microsoft eXecution Container sandbox or WSL2 on Windows. Codex's cloud product adds a second layer: a networked setup phase to install dependencies, then an agent phase with no internet access by default.
The only decision left is how permissive to make it. Network access defaults to off and stays off even in the more permissive workspace-write mode until you turn it on and name an allowlist.
For the full breakdown of Codex's modes and what each one covers, see our guide on optimizing a workflow around OpenAI's real sandbox tools.
- macOS: Seatbelt. Linux: bubblewrap plus seccomp. Windows: MXC sandbox or WSL2.
- Network access is off by default, on only if you name an allowlist.
- Cloud Codex splits a networked setup phase from an offline-by-default agent phase.
If you're building your own agent, pick a hosted provider through the Agents SDK
Once you're not just using Codex but writing your own agent on the open-source Agents SDK, the sandbox becomes a choice. Native sandbox support, shipped in April 2026, lets the SDK hand code execution to a hosted provider instead of you wiring up that integration by hand.
Seven providers are officially integrated: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, and Vercel. They aren't the same underneath. Some run a microVM per task, some a container, and they differ on whether a session survives after it ends. Docker and a local Unix client are also supported, for running the same code on your own infrastructure without a hosted account.
Evaluating which is best for you means comparing on persistence, not just on the isolation model. Daytona, E2B, and Vercel support longer-lived, resumable sessions; others are built to be thrown away after every run. We cover E2B specifically in a separate comparison.
- Seven officially integrated providers: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, Vercel.
- Docker and a local Unix sandbox client cover the no-hosted-account case.
- Compare on persistence, not just isolation: some sessions are resumable, others are ephemeral by design.
If you'd rather not host any of it, the Agents API
The Agents API is OpenAI's answer for teams who don't want to operate the harness or the sandbox integration themselves. It reached public beta on September 10, 2026, and it runs the same underlying harness that drives Codex, as a managed service.
You still choose where the code actually executes: an OpenAI-managed sandbox, your own infrastructure, or a partner environment such as Cloudflare, Modal, or Vercel. What you give up is operating that execution layer yourself; what you get back is session management, orchestration, and context compaction handled for you.
This is the option worth evaluating first if engineering time, not infrastructure control, is your actual constraint.
- Public beta: September 10, 2026.
- Runs the Codex harness as a managed service; you still choose where code executes.
- Best fit when engineering time, not infrastructure control, is the limiting factor.
Offline and private: what that can and can't mean here
A search like sandbox open ai offline private usually conflates two different things. You cannot run an OpenAI model itself offline; every option above still needs a call to OpenAI's servers to get a response from the model.
What can be offline is the execution environment around that response. Codex Cloud's agent phase runs with no network access by default, after a setup phase that is online just long enough to install what you declared. On your own infrastructure, the same pattern works with a self-hosted sandbox: let only your orchestration script talk to the API, and run the code itself with no network access at all.
Private, in practice, means the code never leaves a boundary you control, not that OpenAI is out of the loop entirely. A page promising a fully offline OpenAI sandbox with no model calls at all is describing a different, OpenAI-free setup, not a sandbox for OpenAI's models.
- The model call itself always needs OpenAI's servers; nothing here changes that.
- The execution sandbox around it can be fully offline, in Codex Cloud or in a self-hosted setup.
- Private means code stays inside a boundary you control, not that no network call ever happens.
A short way to decide
If you're already in Codex, stop here and just set its sandbox mode correctly. If you're building a custom agent and want the least code to write, pick one of the seven integrated providers based on whether you need resumable state. If you're building a custom agent and want the least infrastructure to operate, evaluate the Agents API first.
Whichever you pick, a sandbox answers one question only: where the code runs. It doesn't vet what the code does once you copy it out, and it doesn't replace reading the diff before you merge it.
- Already on Codex: configure its mode, you're done.
- Custom agent, minimal code: pick a hosted provider by persistence need.
- Custom agent, minimal infrastructure: evaluate the Agents API first.
- Any option: still read the code before you merge it.
Where Forkbench fits in this comparison, and where it doesn't
Everything above assumes you might be running code you don't fully trust, at some scale, where a dedicated container or microVM per task matters. Forkbench is not an entry in that comparison. It supervises an agent you already trust, including Codex, in a real terminal on your own Mac.
Instead of a per-task sandbox, it gives a Thread a folder lock, enforced by the macOS kernel sandbox, that keeps the agent to the project directory you approved. That lock is opt-in and does not restrict the network, so a locked agent can still make outbound requests if the command it runs is allowed to.
For secrets, the Vault keeps API keys in your Mac's Keychain and lets a command use one by name. That stops a key from sitting in a file the agent can read, but a program the agent is allowed to run can still read the key it was handed; pinned or not, the protection ends at that boundary.
- Forkbench supervises a trusted agent on your own Mac; it isn't a multi-tenant sandbox provider.
- The folder lock is opt-in and does not restrict the network.
- The Vault hides a key's value from the agent, but a program it runs can still read it.
Related: Optimizing your workflow with OpenAI's sandbox tools, Best practices for OpenAI sandbox security, E2B alternative: sandbox agents on your Mac, Forkbench for OpenAI Codex CLI
Frequently asked
What's the best sandbox for running OpenAI's Codex?
Codex's own, which is already built in. It uses Seatbelt on macOS, bubblewrap plus seccomp on Linux, and the Microsoft eXecution Container sandbox or WSL2 on Windows, so there's nothing extra to set up.
What's the best sandbox for a custom agent on the OpenAI Agents SDK?
It depends on whether you need state to survive between runs. Daytona, E2B, and Vercel support resumable sessions; the others are built to be thrown away after each run. All seven, Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, and Vercel, are officially integrated.
Can an OpenAI sandbox run completely offline?
The execution environment can. The model call cannot, since every option still needs to reach OpenAI's servers to get a response.
Is the OpenAI Agents API itself a sandbox?
No. It's a managed service that runs the Codex harness for you; you still choose whether code executes in an OpenAI-managed sandbox, your own infrastructure, or a partner environment.
Does Forkbench replace a hosted OpenAI sandbox provider?
No. Forkbench supervises a trusted agent working in a terminal on your own Mac. For adversarial or multi-tenant code at scale, use a hosted provider like the ones listed above instead.