Blog

Sandbox vs. container vs. microVM vs. VM

Every way to confine a program on Unix sits somewhere on the same spectrum: a wall around a user, a wall around one process, a wall around a group of processes sharing a kernel, or an entirely separate kernel. Picking the right one is a question of what you need isolated and how much of your machine you are willing to give up to get it.

Quick Answer

Isolation on Unix comes in four layers, each tighter and costlier than the last. Unix permissions wall off a user, not a program: every process you run shares the same access. OS sandboxes, macOS Seatbelt, Linux Landlock and seccomp, wrap a single process in a kernel policy with no new kernel and almost no overhead. Containers, Linux namespaces plus cgroups, wrap a group of processes in their own filesystem and network view, but every container still shares the host's one kernel. MicroVMs and full VMs each run their own kernel, so a flaw in one does not reach another, at the cost of booting and feeding a second operating system.

The isolation spectrum, in one line

Every mechanism below answers the same question: how much of the machine does the thing you are running get to touch. They differ in where the wall goes up and what it costs you to put it there.

  • Unix permissions: a wall around a user. Free, and already running; every program you start gets the same access you do.
  • An OS sandbox (Seatbelt, Landlock, seccomp): a wall around one process, enforced by the kernel, with no new kernel and close to no overhead.
  • A container (namespaces plus cgroups): a wall around a group of processes with their own filesystem and network view, sharing the host's one kernel underneath.
  • A microVM (Firecracker) or a full VM: a wall around an entire kernel, each running its own, at the cost of booting and feeding a second operating system.

Unix users and permissions: the oldest boundary

Every process on a Unix system runs as some user ID, and every file carries an owner, a group, and read, write and execute bits for each. That is the original access control, and it still does real work: your shell cannot open another account's files, and a process you start cannot act as root unless something explicitly grants it that.

What it does not do is separate two programs running as you. A coding agent, your editor, and a terminal you opened by hand all run under the same user ID, so the permission system that stops you from reading another person's files does nothing to stop one of your own processes from reading all of yours. Every mechanism past this point exists to put a wall around a program, not a person.

OS sandboxes: a kernel policy around one process

An OS sandbox confines a single process, and whatever it spawns, to a policy enforced inside the kernel, with no second operating system underneath it. On macOS that is Seatbelt; how it actually confines a coding agent, and its two real limits, is covered in full here.

Linux has two separate primitives that do a similar job from different angles. Landlock lets an unprivileged process restrict itself, and any children it starts afterward, to a set of filesystem paths and, since ABI version 4, specific network ports, no root and no system-wide policy required. seccomp works one level lower: it does not decide which files a process may touch, it decides which system calls the process is allowed to make at all, from a strict read, write and exit-only mode to a custom allow-list built with BPF. Browsers, container runtimes, and most hardened server processes run under a seccomp filter by default.

The shared property across all three: no new kernel boots, so the cost stays close to zero, and the policy is attached at the kernel level, so a process running inside it cannot widen its own grant by asking nicely.

Containers: namespaces and cgroups around a shared kernel

A container is built from two ordinary Linux kernel features, not a boundary of its own. Namespaces give a group of processes their own view of a global resource. Linux ships eight kinds, PID, network, mount, UTS (hostname), IPC, cgroup, time and user, and a process inside a PID or network namespace sees its own process-ID space or network stack, isolated from the host's. Cgroups separately group processes to meter and cap what they are allowed to use of CPU, memory and other resources, regardless of what they can see.

Put those two together and you get what every container runtime actually is: a process, or a small tree of them, running in its own namespaces, metered by a cgroup, started from its own filesystem image. The honest limit is in that sentence: there is still exactly one kernel underneath every container on the machine. A vulnerability in the kernel itself, not in anything a namespace or cgroup governs, reaches every container running on top of it.

Same primitives, different defaults: where bubblewrap sits

The line between a sandbox and a container is softer than the names suggest, because both are frequently built from the same kernel primitives. Bubblewrap, the sandbox underneath Flatpak, uses the same Linux user and mount namespaces a container runtime does, and nothing more privileged, to let an ordinary, unprivileged user build a restricted environment for one program.

What makes it read as a sandbox rather than a container is the default it ships with: deny almost everything, and hand back only the paths and devices explicitly bound in, versus a typical container's default of a full filesystem image scoped to its own namespace. It also turns on PR_SET_NO_NEW_PRIVS, the same kernel switch that stops a setuid binary from using its bit to climb back out, a step classic chroot-based jails never had. The mechanism is shared; the posture, what is allowed unless stated, against what is denied unless stated, is the actual difference.

MicroVMs: a tiny kernel per workload

A microVM runs its own, minimal kernel, separate from the host's, under a hypervisor built to start it far faster than a general-purpose VM boots. Firecracker, open-sourced by AWS and built on Linux's KVM, is the reference example: its own design documentation describes a minimalist device model, a handful of VirtIO network, block and vsock devices and little else, built to exclude unnecessary devices and guest-facing functionality specifically to shrink both the memory footprint and the attack surface a guest can reach, which its maintainers credit with decreasing startup time.

That is what buys microVMs their real property over a container: a kernel bug, the one thing a namespace or a cgroup cannot contain, stays inside the single microVM it was triggered in, because every microVM is a separate kernel instance. AWS runs Lambda and Fargate on exactly this property, placing workloads from different customers on the same physical host. The cost is running many kernels instead of one shared kernel, real, but far smaller than a traditional VM's.

Full VMs: a different computer underneath

A full VM goes further still. A hypervisor presents an entire virtual computer, CPU, memory, disk, network card, to a guest operating system that has no idea it is not running on real hardware. Nothing about the host's kernel, filesystem or process table is visible to the guest at all, which is the strongest isolation on this page.

It is also the most expensive. A full VM boots a complete operating system, not a few hundred kilobytes of minimal kernel, and the virtualized hardware sitting between the guest and your actual disk and network adds real latency. What that costs specifically, for a coding agent running on a Mac, is covered here: your Keychain, Touch ID and native toolchains stop being reachable from inside the VM, and you end up maintaining your setup twice.

How this works in Forkbench

Forkbench's folder lock sits at the OS-sandbox layer, not the container or VM layer. Locking a Thread to a set of folders starts every shell in it under a macOS Seatbelt profile that denies reads and writes under /Users, /Volumes and /Network outside the folders you allowed, enforced by the kernel and inherited by everything the shell spawns. No second kernel boots, nothing is rebuilt, and the cost stays close to the near-zero an OS sandbox always costs.

Stated exactly as the limit is published: the lock is on folders only, so a locked Thread can still reach the network, and a profile can only be applied as a shell starts, so a terminal already running cannot be sealed after the fact. That is macOS today; Windows and Linux builds are in development, with no date attached to either.

Related: macOS Seatbelt in depth: SBPL, inheritance and the non-nesting rule, How to sandbox Claude Code on macOS, How Forkbench locks a Thread to its folders

Frequently asked

Keep reading

Sources