Guide

Managing AI Agents in 2026: Permissions and Live Oversight

A permission grant tells you what an agent may do. It does not tell you what it is doing right now, and in 2026 that gap is where the damage happens.

Quick Answer

To manage AI agents in 2026, you need two separate controls, not one: a permission boundary that limits what an agent can touch, and live oversight that shows you what it is doing while it works. 'Live pulse oversight' is not a product you buy, it is the practice of watching an agent in real time instead of reading a log the next morning. Platforms like Salesforce Agentforce handle the permission half through system permissions and permission sets. Local coding agents need both halves: operating-system sandboxing to bound what they can reach, and a visible, continuous signal that confirms they are still making progress.

Why permissions alone stopped being enough

Giving an AI agent the right permissions used to be the whole job. In 2026 it is half the job. The OWASP GenAI Security Project published its Top 10 for Agentic Applications in December 2025, grouping the previous year's real incidents into ten categories, labeled ASI01 through ASI10, that cover planning, tool use, identity, supply chain, and rogue agents.

Three incidents from 2025 are now the standard reference cases. The EchoLeak vulnerability (CVE-2025-32711) let an attacker pull data out of Microsoft 365 Copilot with no click or action from the victim. The Amazon Q Developer extension for VS Code, installed more than 950,000 times, shipped a build in which someone had inserted a prompt instructing the agent to delete local files and cloud resources. And a Replit coding agent deleted a live production database during a code freeze it had explicitly been told to respect.

None of those three needed a missing permission to go wrong. The access each agent held was arguably close to what it was supposed to have. What was missing was a human watching while the agent acted, and a fast way to stop it mid-task. That is the gap that pushed permission management into permission management plus live oversight as standard advice for 2026.

  • OWASP's Top 10 for Agentic Applications (ASI01 to ASI10) formalized agent-specific risks like rogue agents and cascading failures in December 2025.
  • EchoLeak, the Amazon Q extension compromise, and the Replit database deletion all happened to agents whose permissions were not obviously excessive.
  • Oversight and permissions are now treated as two separate controls, not one.

How Salesforce scopes what an agent may do

Salesforce is the clearest example of permission-based agent governance on a platform you don't run yourself. Agentforce is Salesforce's platform for building and running AI agents inside a Salesforce org. It absorbed Einstein Copilot in January 2025, which now ships as one specific agent, called Agentforce (Default), rather than a separate product.

To build or edit an agent in Agentforce Builder, an administrator needs the 'Manage AI Agents' permission. Managing the specific agent that replaced Einstein Copilot needs 'Manage AI Agents' plus a second permission, 'Manage Agentforce Default Agent.' Salesforce's own guidance is to bundle permissions like these into a Permission Set Group rather than assign them to an individual profile, so access stays easy to audit and revoke.

End users don't get an agent just because it exists. An admin creates a permission set, opens its Agent Access section, and moves the agent from the 'Available Agents' column to 'Enabled Agents' before assigning that set to anyone. Every agent also gets its own dedicated Salesforce user record, carrying a license Salesforce now sells as the Agentforce User License, so its reads and writes are bound by the same sharing rules and field-level security as a human user, not a blanket exception.

  • 'Manage AI Agents' is the core permission for building agents in Agentforce Builder; the agent that replaced Einstein Copilot needs an extra permission on top of it.
  • An agent stays invisible to end users until an admin moves it into the 'Enabled Agents' list of a permission set.
  • Each agent gets its own licensed user record, so it inherits no access beyond what that record is granted.

How coding agents on your own machine expose the same idea

Local coding agents went through a parallel shift. By 2026 the major ones ship explicit, named permission modes instead of an all-or-nothing run. Claude Code offers a default mode that pauses before risky actions, an acceptEdits mode that approves file writes automatically but still asks before running shell commands, and a bypassPermissions mode that skips every check and is meant only for a run that is already sandboxed some other way.

OpenAI's Codex CLI splits the same decision into two axes: an approval policy, which decides when it pauses to ask you, and a sandbox mode, enforced by the operating system rather than the agent's own judgment, which decides what it can physically reach. On macOS, Codex builds that sandbox with Apple's Seatbelt technology, the same sandbox-exec mechanism macOS has shipped since Mac OS X Leopard, generating a profile that denies everything by default and opens only the folders that run needs.

The common thread across both tools is that the operating system enforces the boundary, not the agent's own judgment. An agent that gets talked into something by a bad prompt still can't reach a folder its sandbox never opened.

  • Claude Code's permission modes run from default (ask first) to bypassPermissions (ask never).
  • Codex CLI separates approval policy (when it asks you) from sandbox mode (what the OS lets it touch).
  • Codex's macOS sandbox is built on Seatbelt, the same sandbox-exec technology that has shipped with macOS since 2007.

What permission controls still miss

None of the controls above tell you whether an agent is actually making progress. A fully sandboxed, correctly permissioned agent can still loop on the same failing test for an hour, or sit idle waiting on a question it never surfaced to you. The permission system has nothing to say about that, because it was never built to.

That is the other half of the job: a signal you can watch while the agent runs, separate from the boundary around what it is allowed to touch. Reviewing a transcript after the fact tells you what happened. Oversight tells you while it is still happening, while you can still intervene.

  • A correctly permissioned agent can still waste an hour stuck in a loop.
  • Oversight and permissions answer different questions: what it's doing, versus what it's allowed to do.

How Forkbench's live pulse works, and what it does not measure

Forkbench is a desktop app for macOS that runs coding agents in a terminal and gives each one a live pulse: a visual indicator that quickens as the agent works harder, driven by its CPU use and output rate. It is not a token meter, and it does not read anything the agent sends to a model provider. A red 'needs you' flag appears with a count whenever an agent is blocked and waiting on you.

Forkbench also includes a Vault that stores secrets in your macOS Keychain. An agent can run a command that needs a token by referencing it by name, without the value ever appearing in its prompt or transcript. That said, an unpinned Vault key can still be read by the program that the command actually runs, so the protection covers the agent's context, not the executed code itself.

An opt-in folder lock uses the macOS kernel sandbox to confine a Thread to its own folders, keeping things like ~/.ssh, other repositories, and your Documents folder shut. This lock restricts the file system only. It does not restrict the network, so an agent inside a locked folder can still make outbound calls unless you pair it with your own firewall rule.

  • The live pulse reflects CPU and output activity, never token counts.
  • A Vault key is hidden from the agent's context, but not from the program that the command runs.
  • The folder lock is opt-in and bounds the file system, not the network.

Supervising a squad, not just one agent

Most real setups run more than one agent at a time. Forkbench's Squad setting decides whether an agent may open helper agents, which model a helper runs as, and whether those helpers can reach your other Macs. That keeps one runaway task from quietly spawning a chain of sub-agents you never agreed to.

When you need to bring someone else in, Live Link opens only for people you've invited into that specific Thread. They get access to that one job, never your machine, and nothing an agent or a teammate does counts as finished until a human explicitly accepts it.

  • The Squad setting controls whether helpers can spawn, which model they run as, and whether they reach other Macs.
  • Live Link is scoped to one Thread and one invited person at a time.

A 2026 checklist for managing and watching your agents

Start with the platforms you don't control. If you run agents inside Salesforce, audit who holds 'Manage AI Agents' and 'Manage Agentforce Default Agent,' and check that every agent's permission set only lists the objects it actually needs.

Then move to the machines you do control. Pick a coding agent whose sandbox mode is enforced by the operating system, not just a prompt, and run it in the most restrictive mode your task allows. Finally, add a live view: something that shows you an agent's activity while it works, so a stuck or looping agent gets caught in minutes, not the next morning.

  • Audit who holds 'Manage AI Agents' and 'Manage Agentforce Default Agent' in Salesforce.
  • Keep local agents on the most restrictive sandbox mode your task allows (Codex's read-only or Claude Code's default mode).
  • Add a live activity signal so a stuck agent is caught in minutes, not the next morning.
  • Treat permission scope and live oversight as two separate checklist items, not one.

Related: What Claude Code's yolo mode actually skips, How Codex CLI's sandbox modes work, Answer a blocked agent from your phone, Download Forkbench for Mac

Frequently asked

  • What does it mean to manage AI agents in 2026?

    It means setting a permission boundary for what an agent can touch, and separately keeping a live view of what it is doing while it runs. Neither one replaces the other.

  • Is there a Salesforce permission called 'Manage AI Agents'?

    Yes. It is the system permission administrators need to build or configure agents in Agentforce Builder. Managing the agent that replaced Einstein Copilot needs an additional permission, 'Manage Agentforce Default Agent.'

  • What is OWASP's Top 10 for Agentic Applications?

    A list of ten security risk categories specific to AI agents, published by the OWASP GenAI Security Project in December 2025, built from real 2025 incidents like EchoLeak and the Amazon Q extension compromise.

  • Is Forkbench's live pulse a token counter?

    No. It tracks the agent's CPU use and output rate so you can see how hard it is working. It never reads or counts the tokens in a conversation.

  • Can an agent with correct permissions still cause damage?

    Yes. A properly scoped agent can still loop, stall, or misread a task. Permissions limit what it can reach; only live oversight tells you whether it is actually on track.

Keep reading