Guide

The Top Autonomous Coding Agents With Sandbox Support, Reviewed

Some coding agents ship an operating system boundary around every command they run. Others only ask nicely. Here is which is which, checked against each one's own documentation.

Quick Answer

Among autonomous coding agents with sandbox support, four actually enforce the boundary at the operating system level: Claude Code and Codex CLI both build on Apple's Seatbelt on macOS and on a Linux sandbox tool, Cursor CLI has its own sandbox.json boundary, and Gemini CLI can wrap itself in Seatbelt or a container. GitHub Copilot CLI and OpenCode only prompt for permission before a risky action, which is a request, not a boundary a command can be forced into. Aider has neither, and runs with whatever access your user account already has. None of these sandboxes block network access by default except Claude Code's, so check that separately from the file boundary before you trust any of them with code you did not write yourself.

What 'sandbox support' means for an autonomous coding agent

An autonomous coding agent reads files, edits files and runs shell commands without asking you first every time. Sandbox support means something specific: an operating system boundary around those actions that the agent cannot argue its way past. A system prompt that says 'ask before deleting files' is not that. The agent can forget the instruction, or a later message can override it.

A real sandbox answers three questions for every command the agent runs. Where can it write? What can it read? Where can it connect? Most of the tools in this review answer the first question well and the third one badly, which matters more than it sounds, because a command that cannot write outside your project folder can usually still read a credential file and send it somewhere.

This review checks each agent's own documentation for how it answers those three questions, not its marketing page.

  • Writes: can the agent only touch the project folder and a temp directory?
  • Reads: can it still open ~/.ssh or a cloud credentials file?
  • Network: can a sandboxed command reach the internet, and does that need a separate allow list?

Claude Code: a sandbox that's off until you turn it on

Claude Code ships a built-in Bash sandbox, built on Apple's Seatbelt framework on macOS and on bubblewrap on Linux and WSL2. It is off by default. You turn it on by running /sandbox in a session, or by setting sandbox.enabled to true in a settings file.

Once it is on, a sandboxed command can write to the working directory, a per-user temp folder, and any directory you added, and nothing else. Network access goes through a local proxy that checks each host against an allow list that starts empty, so a command cannot reach a new domain without you approving it first.

The part worth knowing: reads are not locked down by that same default. Claude Code's own documentation says a sandboxed command can still read most of the machine, including files such as ~/.ssh and ~/.aws/credentials, unless you add them to a deny list or use the separate credentials setting. The sandbox also wraps shell commands only. Claude's file tools, MCP servers and hooks run outside it, with your full access.

  • Off by default. Turn it on with /sandbox or sandbox.enabled in settings.
  • Writes and network are locked down by default; reads are mostly not.
  • File tools, MCP servers and hooks run outside the Bash sandbox entirely.
  • Linux and WSL2 need bubblewrap and socat installed; macOS needs nothing extra.

Codex CLI: three modes, and a default that still writes

OpenAI's Codex CLI has three sandbox modes. Read-only lets it look at files but not change anything or run a command without your approval. Workspace-write, the default, lets it read, edit inside the project, and run routine local commands inside that boundary. Danger-full-access removes the filesystem and network boundary entirely, and OpenAI's own documentation advises against it outside an environment that is already isolated.

The enforcement is Seatbelt on macOS and bubblewrap on Linux and WSL2, the same primitive Claude Code uses, so the two tools fail in similar ways when a command needs something outside the box.

Approval policy is a second, separate dial. On-request means Codex works inside the sandbox and asks only when it needs to go further. Never means it does not stop to ask at all, a mode you want paired with a sandbox, not with danger-full-access.

  • Default mode is workspace-write, not read-only. Check which mode you are actually in.
  • Seatbelt on macOS, bubblewrap on Linux and WSL2.
  • danger-full-access plus a never approval policy removes every boundary at once.

Cursor CLI and Gemini CLI: a boundary, built differently

Cursor's agent separates two concerns. permissions.json decides which calls run automatically and which get reviewed. sandbox.json decides what a sandboxed shell command can reach, including network domains and extra readable or writable paths. Cursor's own documentation recommends Auto-review as the safest useful setup: it runs known-safe calls, sandboxes shell commands when it can, and sends anything else to a classifier for review. In headless mode without the force flag, any command not on the allow list is simply denied, with no prompt, a stricter failure mode than most agents use.

Gemini CLI takes a third approach. You turn sandboxing on with the -s flag or the GEMINI_SANDBOX environment variable, and you pick the mechanism: macOS Seatbelt through sandbox-exec, or a Docker or Podman container for a cross platform boundary. The macOS profile is set by SEATBELT_PROFILE, and the default, permissive-open, only restricts writes outside the project. Google's own documentation describes it as allowing most other operations, which includes network access unless you choose a stricter profile.

  • Cursor: sandbox.json for the boundary, permissions.json for what still needs review.
  • Cursor's headless mode denies anything off the allow list with no prompt.
  • Gemini CLI: -s or GEMINI_SANDBOX, then Seatbelt or a container.
  • Gemini's default Seatbelt profile is permissive-open: it blocks stray writes, not network access.

Copilot CLI, OpenCode and Aider: ask, don't enforce

GitHub Copilot CLI automatically allows read-only actions and asks before anything that modifies your system, such as a destructive shell command or a file edit. You can grant an approval for one action or the rest of the session, and the --allow-all-tools or --yolo flag skips every prompt at once. None of this is an operating system boundary. GitHub's own documentation recommends pairing wide permissions with a local or cloud sandbox you set up yourself, since the CLI does not restrict network access by default even when terminal file access is sandboxed.

OpenCode, the open source agent from SST, runs on a permission system with three states per rule: allow, ask or deny, matched against tool and path patterns. Most permissions default to allow, with doom-loop detection, access outside the project, and .env files defaulting to ask or deny instead. It is a prompt layer, not operating system level isolation, by its own design.

Aider has neither. It commits every change to git with attribution, which gives you a clean audit trail, but nothing stops it from reading or writing anything your user account can touch. If you want a boundary around Aider, you have to build one yourself, with a container or a tool that wraps the process.

  • Copilot CLI: read-only auto-allowed, writes need approval, --yolo skips all of it.
  • OpenCode: allow, ask or deny per tool and path, no OS-level sandbox.
  • Aider: no sandbox and no permission layer. Git history is your only record.

Where Forkbench fits, and where it doesn't

Forkbench is a Mac app built to run your coding agents in real, local terminals rather than a hosted sandbox. Its folder lock uses the macOS kernel sandbox to shut a Thread out of ~/.ssh, your other repositories and Documents, and it works the same way whichever agent you run inside it, including Aider, which has no boundary of its own to begin with.

The Vault handles the other half: a key stays in your Keychain, and a command uses it by name without the value ever reaching the prompt, the command line or the transcript.

Know what it does not cover. The folder lock is opt-in and does not restrict the network, so a locked agent can still send out whatever it was allowed to read. An unpinned Vault key can still be read by the program it was handed to. Forkbench governs the folders you lock and the keys you put in the Vault, not the rest of the machine.

  • Folder lock: opt-in, applies the same way to any agent you run in a Thread.
  • Vault: a key is used by name, not shown, and stays in the Keychain.
  • Not covered: network access, and anything outside the folders you locked.

How to pick one

Match the sandbox to the agent you are actually running, not the one with the best documentation. If you run Claude Code or Codex, turn their sandbox on, since it is free and already built. If you run Aider or a tool with no boundary of its own, add one, with a container for risky code or Forkbench's folder lock for everyday work on your own Mac.

Then check network access separately from writes. A command that cannot touch your home folder can still leak a file it was allowed to read, so decide what each agent may connect to and do not assume the write boundary covers it.

None of this replaces reading the diff. A sandbox limits the damage from a bad command. It has no opinion on whether the code is correct.

  • Turn on the sandbox your agent already ships, if it has one.
  • Add a boundary yourself for any agent that ships none.
  • Set network access on purpose. Most of these defaults are about files, not connections.
  • Keep reviewing diffs. A sandbox is a limit on damage, not a code reviewer.

Related: How to sandbox AI coding agents on macOS, Codex CLI sandbox modes, and the part the sandbox does not cover, Forkbench vs. the Claude Code sandbox, Download Forkbench

Frequently asked

  • Which coding agents sandbox commands by default, with no setup?

    None of them. Claude Code's Bash sandbox and Codex CLI's sandbox are both off or partial until you choose a mode, Cursor and Gemini CLI need a flag or a run mode, and Copilot CLI and OpenCode only prompt for permission. Every one of them needs you to turn something on.

  • Does a sandboxed coding agent still need code review?

    Yes. A sandbox limits what a bad command can reach. It says nothing about whether the change the agent made is correct, so review the diff the same way you would for a human's pull request.

  • Can I sandbox Aider?

    Not with anything built into Aider itself. Run it inside a container, wrap it with your own Seatbelt profile, or run it inside a Forkbench Thread with the folder lock turned on.

  • Do these sandboxes stop an agent from leaking an API key over the network?

    Only Claude Code's sandbox blocks network access by default, through a proxy with an empty allow list. The others either allow it by default, like Gemini's permissive-open profile, or don't address it, like Copilot CLI, so check your tool's network behavior separately from its file boundary.

  • Does Forkbench add a sandbox to agents that don't have one?

    Its folder lock does, for the folders a Thread can touch, using the macOS kernel sandbox. It's opt-in and doesn't restrict the network, so pair it with the same network thinking you'd apply to any other tool on this page.

Keep reading