Guide
Local AI Agent Security: Enforcing Permissions and Boundaries
A permission mode is a setting your agent agrees to respect. A sandbox is a wall the operating system enforces whether the agent agrees or not. Local security needs both, in that order.
Enforcing permissions and boundaries for a local AI coding agent means two separate things. The first is the agent's own permission settings: Claude Code has six modes, from Manual, which only auto-approves reads, up to bypassPermissions, which Anthropic's own documentation says is for isolated containers and VMs only. Codex CLI splits the same idea into two independent knobs, a sandbox_mode (read-only, workspace-write, or danger-full-access) and a separate approval policy for how often it pauses to ask. The second, and the one people skip, is that these settings are honored by the agent's process, not enforced by the operating system below it, except where a real sandbox sits underneath, such as Claude Code's Bash sandbox, which runs on macOS, Linux, and WSL2 but not on native Windows. Pick the strictest mode you can tolerate for the work in front of you, and add a kernel-level boundary, such as Forkbench's folder lock on a Mac, underneath whichever mode you choose.
Two different things people mean by 'permission'
When someone asks how to manage AI agents locally with strict permissions, they usually mean one of two different controls, and conflating them is where setups go wrong.
The first is a policy the agent's own program checks before it acts: can it edit this file, run this command, reach this URL. The agent decides to honor that policy. The second is a boundary the operating system enforces on the process no matter what the agent's code decides: it is simply unable to write outside a folder, or reach the network, because the kernel will not let the system call through.
This guide is mostly about the first kind, because that is what Claude Code's permission modes and Codex CLI's approval settings actually are. The second kind, OS-level sandboxing with Seatbelt, containers, or VMs, is covered in more depth in the sandboxing guide linked below. A serious local setup uses both, with the OS-level boundary doing the real work and the permission mode controlling day-to-day friction.
- Permission mode: a policy the agent's process agrees to follow.
- OS sandbox: a boundary the kernel enforces regardless of what the agent decides.
- Treat the permission mode as convenience, and the sandbox as the actual wall.
Claude Code's permission modes, named correctly
Claude Code ships six permission modes, and the names matter because the wrong one by accident is a common source of surprise. Manual mode, whose config value is 'default', auto-approves reads only and asks before every edit, every shell command, and every network reach. acceptEdits auto-approves file edits and common filesystem commands like mkdir and mv inside your working directory, but still asks for anything else. plan mode lets Claude research and propose a plan without touching your source at all.
Auto mode, the newer addition, removes routine prompts by sending each action to a separate classifier model that blocks things like production deploys, force pushes, and sending data to an unrecognized destination, while still asking you directly for anything the classifier cannot resolve. dontAsk denies anything that would normally prompt, which suits a locked-down script. bypassPermissions skips checks entirely, and Anthropic's own documentation says plainly it is best for isolated containers and VMs only, not your regular Mac or Windows desktop.
You switch between these with Shift+Tab during a session, or start in one directly with a flag such as claude --permission-mode plan. The --dangerously-skip-permissions flag is the command-line route into bypassPermissions, and the name is accurate: it removes the one thing standing between an agent's mistake and your filesystem.
- default (Manual): reads only auto-approved. Best for sensitive or unfamiliar work.
- acceptEdits: file edits and safe filesystem commands run without asking.
- plan: research and propose, no edits until you approve.
- auto: a classifier model reviews actions instead of prompting you for each one.
- bypassPermissions: everything runs. Anthropic says to use it only in a container or VM.
Codex CLI's two independent knobs
Codex CLI separates the same problem into two settings instead of one blended mode. sandbox_mode decides what a command can physically touch: read-only, workspace-write for edits confined to your project folder, or danger-full-access for no restriction at all. A separate approval policy decides how often Codex pauses to ask you before it runs something, independent of the sandbox level.
That separation is useful once you understand it: you can run a strict read-only sandbox with a loose approval policy for quick exploration, or a permissive workspace-write sandbox with an approval policy that still checks in before anything irreversible. Codex CLI installs and runs directly on Windows, not only through WSL, which is a real difference from Claude Code's own sandbox, covered next.
Either tool's settings only matter if you actually set them somewhere other than the most permissive option out of convenience. The default worth aiming for is read-only or workspace-write plus an approval policy that still asks for anything outside your project, loosened only once you have a reason to trust the specific task.
- sandbox_mode: read-only, workspace-write, or danger-full-access.
- Approval policy: a separate setting for how often you get asked.
- The two combine, so a strict sandbox does not require a strict approval policy too.
What's actually enforced versus merely asked
Claude Code also has a true OS-level sandbox for its Bash tool, separate from permission modes, that restricts which files and network hosts a shell command can reach, enforced by the operating system rather than by Claude's own judgment. The important limit: it runs on macOS, Linux, and WSL2. On native Windows, Claude Code's own documentation says commands run unsandboxed, full stop. If you want that sandbox on a Windows machine, you run Claude Code inside a WSL2 distribution, not directly on Windows.
That single fact changes the local security picture for a lot of Windows developers: a permission mode like Manual still asks you before most actions, but nothing underneath enforces a boundary if you say yes to something you should not have. The permission mode is doing all of the work on native Windows, with no sandbox backstop.
On macOS and Linux, and on Windows through WSL2, you get both layers: the permission mode controlling what gets asked, and the sandbox controlling what is even possible if something slips through.
- Claude Code's Bash sandbox: real OS enforcement, on macOS, Linux, and WSL2.
- Native Windows: no sandbox backstop for Claude Code today. The permission mode is the only control.
- Codex CLI's sandbox_mode is enforced per-command wherever it runs, including native Windows.
A sane default for a local setup
Start stricter than feels necessary. For unfamiliar code or anything touching production, that means Claude Code's Manual mode or Codex's read-only sandbox, with every edit and command reviewed by you. Loosen one step at a time as a specific task earns it: acceptEdits for a refactor you are already watching in git diff, auto mode for a long task you are willing to let the classifier police instead of you.
Reserve bypassPermissions and danger-full-access for exactly the case their own documentation describes: an isolated container or VM you would not mind an agent trashing completely, never your main development machine. If you run many agents locally at once, give every one of them the same conservative starting mode by default, in a shared settings file, rather than trusting each session to remember to set it.
- Default to the strictest mode that does not slow you down into frustration.
- Loosen per task, not permanently, and tighten back up when the task is done.
- bypassPermissions and danger-full-access belong in a container or VM, not your desktop.
Where Forkbench adds a layer under any of this
Forkbench runs coding agents in real terminals on a Mac and can lock a Thread to its folders using the macOS kernel sandbox, so paths like ~/.ssh, other repositories, and Documents stay out of reach no matter what the agent's own permission mode allows. That boundary works underneath Claude Code, Codex, or anything else you run in its terminals, which matters most exactly when you have chosen a looser permission mode on purpose.
It has real limits, the same as any boundary. The folder lock is opt-in, so you still have to turn it on. It does not restrict the network, so an agent that can read an allowed folder can still send what it reads somewhere else. And a Vault key that is not pinned to a specific command can still be read by whatever program you handed it to. Forkbench governs what an agent holds and the folders you lock, not the whole machine.
Each tab also carries a live pulse that reflects CPU and output activity, and a red flag when an agent is blocked waiting on you. That is a supervision signal, not a permission control, and it is most useful exactly when you have relaxed the permission mode and want to know something is still happening.
- Folder lock: opt-in, enforced by the macOS kernel sandbox, works under any agent.
- Limit: does not restrict the network, and only covers folders you chose to lock.
- The live pulse tells you an agent is active. It is not a permission check.
A checklist for running agents locally with real boundaries
Put this in place once and it holds for every session after, rather than deciding permission settings fresh each time you open a terminal.
- Set a conservative permissions.defaultMode in your shared Claude Code settings file.
- Check which Codex sandbox_mode and approval policy you actually run with, not the defaults you assume.
- Confirm whether your OS-level sandbox is active for the session you are about to start, especially on native Windows.
- Reserve bypassPermissions and danger-full-access for a container or VM only.
- Add a folder lock or equivalent OS boundary underneath whichever permission mode you pick.
- Review loosened permissions at the end of a task instead of leaving them loose by default.
Related: How to sandbox AI coding agents on macOS, Claude Code yolo mode: what it actually skips, Securing agent API keys on Windows, Download Forkbench
Frequently asked
Is a permission mode the same thing as a sandbox?
No. A permission mode is a policy the agent's own process agrees to follow, and it can only ask before acting. A sandbox is a boundary the operating system enforces, so the agent cannot get past it even if it tries.
What is Claude Code's bypassPermissions mode, and should I use it?
It skips every permission check so the agent can run unattended. Anthropic's own documentation says it is best for isolated containers and VMs only, not a regular desktop machine.
Does Claude Code's own sandbox work on native Windows?
No. It runs on macOS, Linux, and WSL2. On native Windows, Claude Code's shell commands run unsandboxed, so running Claude Code inside WSL2 is the way to get the sandbox on a Windows machine.
What's the difference between Codex CLI's sandbox_mode and its approval policy?
sandbox_mode decides what a command can physically touch, from read-only up to full access. The approval policy is a separate setting for how often Codex pauses to ask you first. They work independently of each other.
Does Forkbench's folder lock replace my agent's own permission settings?
No, it sits underneath them. The permission mode still decides what the agent asks to do. The folder lock is an extra, opt-in boundary on top, and it does not restrict network access.