Guide

Why Your AI Software Needs a Secure Sandbox Environment

A coding agent runs shell commands and reads files with your own permissions. A sandbox is what stops that automation from touching more than it should.

Quick Answer

A coding agent executes shell commands and reads files using your local permissions, so without a boundary it can touch anything you can touch, on purpose or by accident. A sandbox environment for AI software draws that boundary with containers, microVMs, or operating-system restrictions, so the agent can still do its job inside the lines you drew, but can't silently alter production credentials, rewrite system files, or send data somewhere you didn't approve.

What a 'pi dev agent sandbox framework' search is really asking

Pi, published at pi.dev, is a real open-source coding agent harness built by Mario Zechner, the developer behind libGDX. Its design is deliberately minimal: four tools (read, write, edit, bash) and a system prompt kept under 1,000 tokens, with everything else, including sandboxing, left as an optional extension you add yourself.

Pi's own documentation is direct about this: its tools and extensions run with the permissions of the Pi process, and project trust controls which files it loads, but it does not sandbox tool calls. In practice, that means Pi can read or write anything your user account can reach unless you put a boundary around it from the outside.

That outside boundary is what people mean when they search for a sandbox framework for a tool like Pi. The two common ways to build one are an operating-system restriction, like macOS's sandbox-exec, or a container, like Docker, that only mounts the folder the agent needs.

  • Pi is a real, minimal coding agent harness with four built-in tools: read, write, edit, bash.
  • Pi's own docs state plainly that it does not sandbox tool calls; isolation is left to you.
  • A Docker volume mount like -v $(pwd):/app gives an agent only your project folder, not your whole account.

Building the file-system boundary

Running an agent like Pi inside Docker is a common first step. A command such as docker run -it -v $(pwd):/app -w /app node:22 bash mounts only your current project directory into the container. The agent can read and edit that folder, but it has no path to your global ~/.ssh or ~/.aws directories, because nothing ever mounted them.

On macOS, Codex CLI shows the same idea applied without a container at all. Codex builds its sandbox with Apple's Seatbelt technology, the sandbox-exec mechanism macOS has shipped since Mac OS X Leopard in 2007. It generates a profile on the fly that denies every file-system and network action by default, then opens exactly the folders that specific run needs, with your .git directory carved out as read-only so the agent can't rewrite its own history.

You can borrow the same approach for a harness that doesn't sandbox itself, like Pi: write a custom sandbox-exec profile that lists only the directories the agent should touch, and run the agent under it.

  • docker run -v $(pwd):/app:ro -w /app node:22 bash mounts your project read-only, so the agent can read code but not change it.
  • Codex CLI's macOS sandbox is built on Seatbelt (sandbox-exec), the same technology that has shipped with macOS since 2007.
  • A custom sandbox-exec profile can give any CLI agent, including Pi, the same file-system boundary Codex builds for itself.

Building the network boundary

A file-system boundary alone still leaves the network open. If a dependency an agent installs turns out to be malicious, an open network lets it call home. Docker's --network none flag removes network access from a container entirely, which is the simplest fix when a task doesn't need to reach the internet.

For a stronger boundary than a container provides, AWS Firecracker runs each workload in its own microVM with an independent kernel, booting in around 125 milliseconds and using under 5 MiB of memory overhead per VM. Firecracker is what runs AWS Lambda and Fargate today, and its jailer process adds cgroup limits, namespace isolation, and seccomp filtering on top of the VM boundary itself.

The lesson that carries over to a desktop setup: pair whatever file-system restriction you use with an explicit decision about network access, because the two are separate boundaries and a sandbox that only does one of them is half a sandbox.

  • docker run --network none removes all network access from a container.
  • AWS Firecracker boots a microVM in about 125 milliseconds with under 5 MiB of memory overhead, and powers Lambda and Fargate in production.
  • File-system isolation and network isolation are separate controls; closing one does not close the other.

What a 'pulse' for a sandboxed agent actually means

There is no product named 'pulse sandbox' or 'sandbox AI pulse.' People searching that phrase are usually looking for a way to confirm a sandboxed agent is still alive and working, since isolating an agent also makes it harder to see what it's doing.

Forkbench answers that directly for local agents. Every terminal tab carries a live pulse, a visual indicator that quickens with the agent's CPU use and output rate. It is not a token meter and does not read what the agent sends to a model. A red 'needs you' flag appears with a count whenever the agent is stuck waiting on input, which matters more once an agent is sandboxed and can't just ask over an open network connection.

  • 'Pulse sandbox' is not a real product name; the underlying need is visibility into a sandboxed agent's activity.
  • Forkbench's live pulse tracks CPU and output rate, never tokens.
  • A sandboxed agent that gets stuck is easier to miss, which is why a visible activity signal matters more, not less.

Where secrets fit inside the sandbox

A sandbox with a plain-text API key sitting inside it is still a leak waiting to happen. If the isolated directory contains a credential, the agent will eventually read it during normal work, sandbox or not. The fix is to keep secrets outside the project folder entirely and hand them to a command only at the moment it runs.

Forkbench's Vault keeps secrets in your macOS Keychain. An agent references a secret by name and a command can use it without the value ever appearing in the agent's prompt or transcript. That protection has a real limit: an unpinned Vault key can still be read by the program that the command actually runs, so it guards the agent's context, not the executed code itself.

  • Isolating the file system does not protect a secret that is already inside the isolated folder.
  • Forkbench's Vault hides a secret from the agent's context, not from the program the command runs.

Cloud sandboxes versus a local one, in 2026

The debate searchers mean by 'sandbox AI dev 2026' is really cloud isolation versus local isolation. GitHub Codespaces and similar cloud environments spin up a disposable VM off your hardware entirely, billed per minute of compute (roughly $0.18 an hour for a 2-core machine, scaling up from there), plus storage. That keeps execution off your laptop but requires a live connection and a running bill.

A local sandbox avoids both of those costs by using controls your operating system already has. Forkbench's folder lock uses the macOS kernel sandbox to confine a Thread to specific directories, keeping ~/.ssh, other repositories, and your Documents folder shut. The lock is opt-in and bounds the file system only. It does not restrict the network, so a locked-down agent can still make outbound calls unless you add a firewall rule or run it inside a container with --network none as well.

  • GitHub Codespaces bills per minute of compute, starting around $0.18/hour for a 2-core machine, plus storage.
  • A local sandbox has no per-minute bill but needs you to configure the OS-level controls yourself.
  • Forkbench's folder lock restricts the file system, not the network; pair it with a firewall rule or --network none for full isolation.

A layered setup you can build today

No single tool here is a complete answer by itself. Combine a file-system boundary (a read-only or project-scoped mount, or a sandbox-exec profile), a network decision (on, off, or filtered), and a place for secrets to live outside the sandboxed folder.

Start with the agent's own built-in sandbox if it has one, like Codex's Seatbelt profile. If it doesn't, like Pi, wrap it in Docker or a custom sandbox-exec profile yourself. Either way, add a live view of what the agent is doing, so a stuck or runaway process gets caught while it's happening, not after.

  • Prefer a tool's built-in sandbox (Codex's Seatbelt profile) when one exists.
  • Wrap a harness with no built-in sandbox, like Pi, in Docker or a custom sandbox-exec profile.
  • Keep secrets in a keychain or vault outside the sandboxed folder, never inside it.
  • Add a live activity signal on top of the sandbox so a stuck agent is caught quickly.

Related: Forkbench vs running agents in raw Docker sandboxes, How Codex CLI's sandbox modes work, How the macOS Seatbelt sandbox works, How Forkbench handles your data

Frequently asked

  • Does Pi (pi.dev) sandbox itself?

    No. Pi's own documentation says its tools run with the permissions of the Pi process and that it does not sandbox tool calls. Isolation has to be added around it, for example with Docker or sandbox-exec.

  • What is 'pulse sandbox AI' or 'sandbox AI pulse'?

    Neither is a real product. People searching that phrase usually want a way to see that a sandboxed agent is still active. Forkbench's live pulse, based on CPU and output rate, answers that for local agents.

  • Does a Docker container block network access by default?

    No, a default Docker container can reach the network. Pass --network none to remove network access entirely, separate from whatever file-system mount you've set up.

  • How does Codex CLI sandbox itself on macOS?

    It generates a Seatbelt (sandbox-exec) profile on the fly that denies file and network access by default and opens only the folders the current run needs, with .git kept read-only.

  • Can an agent still read a secret inside a sandboxed folder?

    Yes. Sandboxing restricts what the agent can reach outside the folder, but a plain-text key inside the folder is still readable. Move secrets to a keychain or vault instead.

Keep reading