Guide
The Future of Coworking with Live Supervised AI Agents
Developers are moving from treating AI as a background process to treating it as a live coworker. The defining feature of that shift is supervision: a human still has to approve the steps that matter.
A live AI software agent is a persistent session that reads your project, writes files, runs commands, and keeps going instead of answering one prompt and stopping. Workplaces that run these agents at scale, whether that is a LangGraph workflow or a Salesforce Agentforce deployment, treat live supervision as mandatory rather than optional: the agent pauses at defined checkpoints and waits for a human to approve, reject, or correct the action before anything consequential happens, such as deploying code, changing a database, or issuing a refund. The goal is to keep the agent's speed without losing the human's final say.
From background scripts to active teammates
Automation used to mean writing a script, scheduling a cron job, and hoping it succeeded quietly. Early AI tools were used the same way a search engine is used: you asked a chat window for a function, copied it, and pasted it into your editor yourself. A live AI agent session works differently. It is not a one-shot reply, it is a persistent process that reads your project folder, writes files, runs your test suite, and asks for help when it gets stuck.
That shift changes the job. Instead of writing every line, the developer becomes the person who sets direction and reviews the result. The agent proposes a change, and the human decides whether it is correct, safe, and worth merging. A live agent workplace treats the AI less like an autocomplete tool and more like a fast, junior teammate whose pull requests still need a second pair of eyes before they merge.
The risk scales with the speed. An agent working autonomously can delete the wrong branch, push a broken migration, or commit a secret just as quickly as it can ship a real feature. That is why live supervision, not just logging after the fact, has become the baseline expectation for running agents at work rather than a nice-to-have.
- A live agent session is a persistent process, not a single chat reply.
- Developers move from writing every line to reviewing and approving the agent's output.
- The same speed that makes an agent useful also makes its mistakes land faster.
How LangGraph pauses an agent for a human
Teams building their own agent workflows on LangGraph use checkpointers and interrupts to put a human in the loop. A checkpointer, such as the `AsyncPostgresSaver` class from the `langgraph-checkpoint-postgres` package, saves the full state of the workflow to Postgres after every step, so the agent's progress survives a server restart instead of starting over.
The pause itself comes from an interrupt. The classic approach compiles the graph with `interrupt_before=["sensitive_tool_node"]`, which stops execution right before that node runs. The newer, more flexible approach calls the `interrupt()` function from inside a node's own code, so the pause can depend on runtime conditions rather than being fixed at compile time. Either way, when the agent reaches the pause, say it is about to run `aws s3 sync` or drop a table, it stops and waits for input instead of continuing on its own.
A human reviews the proposed action and its arguments, then sends a `Command(resume="approved")` back into the graph to let it continue, or edits the state first if the action needs correcting. The agent only acts once that signal arrives, which is the actual mechanism behind the phrase human-in-the-loop.
- `AsyncPostgresSaver` persists the workflow's state to Postgres after each step.
- `interrupt_before` sets a fixed pause point; the newer `interrupt()` function can pause based on runtime conditions.
- A human resumes the graph with `Command(resume="approved")`, or edits the state before letting it continue.
How Salesforce Agentforce hands off to a human
Organizations running agents inside a CRM or customer-service stack usually reach for Salesforce Agentforce instead of building their own graph. Agentforce ships a built-in Escalation topic, a permission the agent gets that custom topics do not, which lets it hand a conversation to a human when it hits the edge of what it is allowed to decide, such as a refund above its limit or a request outside its scope.
Once escalation triggers, the conversation is routed through Salesforce's Omni-Channel system using a PendingServiceRouting record, the same object Omni-Channel has used for years to route a work item to the next available agent, now extended with fields that record which bot or copilot handled the conversation before the handoff. The human agent picks up the case with the prior conversation already attached to the record, so they are not starting from nothing.
This keeps the automation's speed for the routine ninety percent of cases while reserving judgment calls, refunds, legal language, and anything involving an upset customer, for a person. It is escalation by policy, not a panic button the agent presses when confused.
- The built-in Escalation topic is what gives an Agentforce agent permission to hand off at all.
- Handoff runs through Omni-Channel's PendingServiceRouting object, the same routing mechanism used for human agents.
- The human picks up the case with the existing conversation attached, rather than starting cold.
Coworking safely with Forkbench
Running an agent locally with Forkbench puts supervision directly in your terminal instead of inside someone else's workflow engine. Forkbench's Vault lets an agent use a secret, an AWS key for a deploy, say, by name, without the value ever reaching the chat transcript. The agent can run the command that needs the key; it cannot read the key itself from the conversation.
That protection has a specific edge, and it is worth knowing it: an unpinned Vault key can still be read by the program that actually runs. If the agent executes a script that deliberately prints its own environment, that script sees whatever the shell session can see, the same as it would for you. The Vault stops the secret from quietly entering the agent's context and getting sent to a model provider. It is not a sandbox around the process itself.
Forkbench also offers a folder lock that restricts which directories a Thread's agent can read or write, using the macOS kernel sandbox. This keeps an agent from wandering out of the project and into your other repositories or home folder. But the lock is opt-in, and it does not restrict the network: a locked agent can still make outbound requests unless you add your own firewall rule. Supervising an agent in Forkbench means knowing these two edges and still reviewing its `git diff` before you let it commit.
- Forkbench's Vault lets an agent use a secret by name without the value reaching the model or the transcript.
- An unpinned Vault key can still be read by the program that ran, since the Vault protects the conversation, not the process.
- The folder lock is opt-in and does not restrict the network, so outbound requests still need a firewall rule if you want to block them.
What a daily supervision habit actually looks like
Coworking with an agent calls for a few concrete habits, not just good intentions. The first is scoping the task narrowly. Asking an agent to refactor the entire billing module invites a sprawling, hard-to-review diff. Asking it to extract one function into its own file and run the existing tests gives you something you can actually check in a few minutes.
The second is reading the diff every time, even when the tests pass. Agents can pass a test suite while quietly deleting a compatibility shim they decided was unnecessary, or handling an edge case in a way that happens to satisfy the assertions without being correct. `git diff` and a normal code review are still the backstop, not a formality you can skip because the agent said it worked.
The third is managing the session itself. A long-running agent accumulates thousands of lines of terminal output and file contents in its context, and past a certain point it starts forgetting earlier instructions or repeating work. Restarting with a clean, focused context and only the files that matter is often faster than trying to talk a confused session back on track.
- Give the agent a narrow, checkable task instead of an open-ended one.
- Read the diff yourself every time, independent of whether the tests passed.
- Restart a long session with a clean context rather than nursing one that has drifted.
What this means for how teams hire and train
Running agents as live coworkers only pays off if the organization invests in the supervision side, not just the automation side. A team that lets an agent run unsupervised will eventually hit a mistake expensive enough to erase whatever time it saved, whether that is a bad migration or a secret in a public commit.
That means training people to configure checkpointers and interrupts in LangGraph, to set escalation rules correctly in Agentforce, and to use a desktop tool's Vault and folder lock the way they are actually designed to be used, not the way marketing implies they work. The skill that matters most is not typing fast. It is knowing when to stop an agent and look.
Done well, this does not replace the developer. It moves them up a level: less time typing boilerplate, more time on the architecture decisions and the judgment calls an agent cannot make for itself.
- Unsupervised agents eventually produce a mistake that erases the time they saved.
- Teams need to train people on checkpointers, escalation rules, and vault boundaries, not just on prompting.
- The valuable skill is knowing when to pause an agent, not how fast you can type.
Related: How Forkbench handles your data, Limit what folders an agent can touch, What --dangerously-skip-permissions actually skips, When an unsupervised agent deleted a production database
Frequently asked
What does human-in-the-loop mean for an AI coding agent?
It means the agent cannot complete certain actions on its own. It pauses at a defined point and waits for a human to approve, reject, or edit the step, such as deploying code or changing a database, before it continues.
How does LangGraph actually pause an agent for approval?
You compile the workflow with
interrupt_beforeset on a sensitive node, or call the newerinterrupt()function from inside that node's code. Either way, the graph stops before the action runs and resumes only when a human sendsCommand(resume="approved").Does Forkbench's folder lock stop an agent from reaching the network?
No. The folder lock restricts which directories the agent can read or write using the macOS kernel sandbox, but it is opt-in and it does not restrict outbound network requests.
How does Salesforce Agentforce hand a conversation to a human?
Through its built-in Escalation topic, which routes the conversation into Omni-Channel as a PendingServiceRouting record, the same object used to route work to human agents, so the person who picks it up already has the conversation attached.
Can a Vault key still be exposed if I use Forkbench?
Yes, in one specific way: an unpinned Vault key can still be read by the program that the agent runs. The Vault keeps the value out of the chat transcript and the model's context, but it does not sandbox what the executed process itself can see.