Blog

Least privilege and capability-based security for AI agents

Giving an agent every permission you have is the easy way to wire it up, and it is also the one choice a fifty-year-old security principle says not to make.

Quick Answer

The principle of least privilege, stated by Jerome Saltzer and Michael Schroeder in 1975, says every program should run with the smallest set of permissions it needs, not the largest set its owner holds. Access control lists and capabilities are the two classic ways to enforce that: a list sits on the resource and names who may touch it, while a capability is an unforgeable ticket naming the resource and the right together. For an AI agent, least privilege means scoping filesystem access, network egress, credentials and their lifetime to the task at hand, rather than running it as you, with everything you can reach, for the whole session.

What the principle of least privilege says

Jerome Saltzer and Michael Schroeder stated it plainly in their 1975 paper The Protection of Information in Computer Systems, published in Proceedings of the IEEE: "Every program and every user of the system should operate using the least set of privileges necessary to complete the job." It is one of eight design principles in the paper, alongside fail-safe defaults and complete mediation, and it is the one that has aged best, because it says nothing about how trustworthy anything is.

That is the part worth sitting with. Least privilege is not a judgment about whether a program is good. It is a bet about what happens when a program is wrong, whether through a bug, a bad input, or a successful attack: the damage is bounded by what the program was allowed to touch, not by how well it was written. A program that can only reach the one file it needs can only damage that one file.

The principle also implies its own cost. Scoping something down means some legitimate request will eventually be denied, because the scope did not anticipate it. Saltzer and Schroeder treat that as the price of the guarantee, not a flaw in it.

Access control lists versus capabilities

The same paper lays out the two classic mechanisms for enforcing any of this, and they ask different questions. An access control list sits on the resource and names who may do what to it. It answers "who can access this resource," and it makes revocation simple: removing a name from the list immediately precludes every future attempt by that name, because the check happens at the resource every time.

A capability is the mirror image. It is an unforgeable ticket held by the program, naming the resource and the rights to it in one object, and it answers a different question: "what can this program do." The check happens wherever the ticket is presented, not back at a central list. Saltzer and Schroeder note the capability model's real weakness in the same breath: once someone has handed a copy of a capability to someone else, there is no way to reach into that copy and disable it directly. Revocation has to be designed in on purpose, usually by making every capability expire or pass through something that can be told to stop honoring it.

Neither model is simply the right answer. An access list is easier to audit and revoke; a capability is easier to hand out narrowly, for one use, without the resource needing to maintain a list of names at all. What matters for an agent is which failure mode you would rather have: a grant that is hard to take back, or a grant that is hard to review.

Ambient authority, again

The two mechanisms differ in one more way that matters more than it first appears: how authority arrives. An access list checked by identity gives a program whatever that identity is entitled to, automatically, every time it acts, which is ambient authority: standing, broad, and attached to who the program is rather than to the one thing it is currently doing.

A capability, used correctly, arrives the opposite way: it is designated, handed to the program for the task that needs it, naming the one resource the task concerns. That is not a minor implementation detail. It is the reason capabilities compose better with least privilege in practice: a program holding only the capabilities it was actually given cannot exceed its scope by definition, where a program running under an identity's ambient authority can always reach anything that identity can reach, whether or not the current task has anything to do with it.

What least privilege means concretely for an agent

Stated as a principle, least privilege is easy to agree with and easy to skip in practice, because wiring an agent up with everything is faster than scoping it down. Concretely, for a coding agent, it breaks into four separate things to narrow, and they rarely get narrowed together.

  • Filesystem scope: the folders the current task actually touches, not the user's whole home directory.
  • Network egress: the hosts the current task actually calls, not the open internet.
  • Credentials: the one key or token the task needs, not every secret the account holds.
  • Time-boxing: a grant that expires when the task or the session ends, not one that quietly outlives the work that justified it.

Scoped and short-lived tokens

The time-boxing piece has the clearest existing playbook, because cloud platforms solved a version of this problem for service-to-service access years before agents showed up. AWS's Security Token Service issues temporary credentials that last from a few minutes to several hours and stop working outright once they expire, instead of the long-lived access keys that work until someone remembers to rotate them. The documentation states the point directly: you do not have to distribute or embed long-term credentials, and because the credentials expire on their own, you do not have to explicitly revoke them when the work is done.

That property matters more for an agent than it ever did for a human. A person who leaves a terminal open overnight is still the same person who will notice something odd about their own session. An agent that leaves a scoped, five-minute token sitting in a transcript or a log has left something with a short, self-limiting lifespan. An agent that leaves an unscoped, non-expiring API key sitting in the same place has left something an attacker can use indefinitely, against anything that key can reach. The scope decides the blast radius; the expiry decides how long the blast radius stays open.

Why "the agent runs as you" breaks least privilege

The simplest way to wire up a coding agent is to run it in your own shell, under your own account, and that is also how most people actually do it. The agent inherits every file your account can read, every host your network stack can reach, and every credential sitting in your environment or your keychain, for as long as the session runs. Nothing about that setup is scoped to the task you gave the agent five minutes ago.

That is maximum ambient authority, which is the precise opposite of least privilege, and it holds regardless of how well-behaved the agent is. The question least privilege asks is never "do I trust this program," it is "what can this program reach if something, anything, goes wrong with it." Running an agent as you answers that question with "everything you can reach," every single time, whether or not the agent ever does anything you would object to.

Narrowing that has a real cost, the same cost Saltzer and Schroeder accepted in 1975: a task will occasionally fail because the agent was correctly denied something it happened to need, and someone has to widen the scope on purpose. That is the trade the principle asks for, in exchange for a mistake or a successful prompt injection landing somewhere bounded instead of somewhere unlimited.

How this works in Forkbench

Forkbench is a desktop app for running coding agents, and it applies this kind of scoping in three places rather than as one blanket switch. A Thread, its working unit for one piece of work, can be locked to a set of folders: every shell in it starts under a macOS Seatbelt profile that denies reads and writes outside the folders you allowed, applied by the kernel and inherited by everything the shell spawns. See locking a Thread to its folders.

Keys follow the same logic rather than a separate one. A Vault secret bound to a destination is substituted in by a local proxy only on its way to the one host it belongs to, and the command that uses it runs in a child process denied every outbound connection except that proxy, so the scope is enforced at the network layer, not left to the command's cooperation. See where a key can go.

The same shape shows up on the board, not just on the filesystem or the network. Every agent starts at the bottom authority tier, contributor, and the only way to raise one is a person clicking a control in the Thread overview; no tool call an agent can make reaches that function, so an agent cannot promote itself or a peer, not even by registering under a new name. See how the board's authority tiers work. None of these three is a general capability system an agent can mint or pass around on its own. They are three separate scopes, folders, one process's network, and board standing, each narrowed by a person rather than claimed by the agent.

Related: The confused deputy problem, explained, Locking a Thread to its folders, How the board's authority tiers work

Frequently asked

Keep reading

Sources