Guide

How to Build an Offline Multi-Agent Framework With Real Safeguards

An offline multi-agent framework needs a model you host yourself, not just code that happens to run on your laptop. The safeguards on top are a separate job, and none of the three popular Python libraries do that job for you.

Quick Answer

An offline multi-agent framework is a set of agents built with an orchestration library such as CrewAI, AutoGen or LangGraph, pointed at a model you run yourself instead of a cloud API, so no prompt or completion leaves your machine. None of those three libraries sandboxes the code an agent executes, and none of them limits what a credential can reach once an agent holds it. That is left to you. A build with real safeguards adds four things on top of the orchestration code: a local model runtime so nothing calls out by default, an operating system or container boundary around every command an agent runs, one scoped credential per agent instead of one shared key, and a human checkpoint before an agent's output is used. Add the safeguard layer while you build the orchestration layer, not after a key has already leaked.

What 'offline' actually has to mean

The word offline covers two different claims, and the queries that lead here usually mean the stronger one without saying so. The weak claim is that the orchestration code runs on your own laptop. The strong claim is that the model itself runs on your own hardware too, so a request never leaves the building.

Only the strong claim is actually offline. A Python script that imports CrewAI or LangGraph and runs on your machine still calls an API over the network every time an agent needs the model to think, unless you have pointed it at something local. That call is the one that carries your data out.

Check which one you have before you call a system private. If every agent in it still dials out to a cloud model for every step, the only thing that is local is the part that matters least.

  • Offline by architecture: the model runs on your own hardware, usually served through Ollama.
  • Offline by habit: the orchestration code runs locally, but every agent call still leaves the machine.
  • If you are not sure which one you built, assume the second and verify.

Three orchestration libraries, three different gaps

CrewAI describes itself as an open source Python framework released under the MIT license for building production multi-agent workflows. Its own setup instructions have you put model and tool keys in a .env file, the same way a single script would, and the open source core has no vault or sandbox of its own. CrewAI sells a separate paid platform, CrewAI AMP, that adds what it calls advanced security and compliance measures, which is itself an admission that the free library does not have them.

AutoGen is the multi-agent conversation framework Microsoft Research published in 2023. The paper describes agents that combine large language models, human input and tools in different modes, and that human-input mode is the academic root of the oversight people now ask for by name. The paper does not describe a sandbox for the tools an agent calls, and the library does not ship one either.

LangGraph is the open source, MIT-licensed orchestration library from LangChain. It says it works with any model provider, which makes pointing it at a model you run yourself a configuration change rather than a rewrite. LangChain's paid product, LangSmith, separately advertises sandboxes that run agent-generated code safely, which again means the free library you are actually using does not include one.

  • CrewAI, AutoGen and LangGraph all leave secrets management to a .env file or an environment variable by default.
  • None of the three sandboxes the commands or code an agent executes.
  • Paid add-ons from the same vendors exist precisely because the open source core lacks these two things.

Run the model yourself if offline has to be literal

Ollama is the most common way developers run an open model on their own hardware. It serves local models through an interface other tools can call, and its own description promises prompts that are never stored or trained on, because the request never leaves your machine in the first place.

CrewAI's documentation covers pointing an agent at a local model through Ollama directly. AutoGen and LangGraph both accept an OpenAI-compatible base URL for the model endpoint, and Ollama exposes exactly that kind of endpoint, so redirecting either of them to a local model is a configuration change, not new code.

A smaller local model plans worse than a large cloud one on multi-step tasks. Test the actual workflow your agents run, not a single demo prompt, before you trust an offline model to carry the whole crew.

  • Ollama runs the model on your hardware and does not store or train on your prompts.
  • Point CrewAI, AutoGen or LangGraph at Ollama's endpoint instead of a cloud API.
  • Re-test every multi-step task once you swap in a smaller local model.

Give every agent its own key, not one shared key

A crew where every agent shares one API key or one cloud credential turns a single leak into an incident that touches everything any agent in that run could reach. The fix is a scoped credential per agent, each limited to the one service that agent's job actually needs.

HashiCorp Vault is the name people reach for first, and it is worth knowing it is not open source anymore. Since version 1.15, Vault has shipped under the Business Source License rather than an OSI license, a change HashiCorp made in 2023 and that IBM, which completed its acquisition of HashiCorp on February 27, 2025, has kept. It still runs fully self-hosted and offline, you just cannot resell it as a competing product.

Infisical is the clearly open source alternative: MIT licensed and fully self-hostable. It also ships an Agent Vault feature built for exactly this problem, where credentials are attached at the network boundary so, in its own words, the agent makes authenticated calls without ever holding the credential, and a prompt injection cannot exfiltrate a value it never had.

  • One credential per agent, scoped to that agent's own job.
  • HashiCorp Vault: self-hostable, but Business Source License since v1.15, not open source.
  • Infisical: MIT licensed, self-hostable, with an Agent Vault mode agents never see the value of.

Add a human checkpoint before output ships

None of CrewAI, AutoGen or LangGraph pauses by default before an agent's output is committed, sent to someone, or executed somewhere else. That gate has to be written into the crew or the graph yourself, as an explicit step an agent cannot skip.

This borrows directly from the human-input mode in the original AutoGen paper, which treats a person as one more participant a conversation can route through, not an afterthought bolted on at the end. Put that participant before anything irreversible, not after.

  • Write the approval step into the graph or crew definition, not into a comment telling the agent to ask first.
  • Put the checkpoint before anything irreversible: a commit, a send, a delete, a payment.

Where Forkbench fits, and where it does not

Forkbench is a desktop app for the Mac that runs coding agents in real terminals. It does not run or host a CrewAI crew, an AutoGen graph or a LangGraph app itself, and it is not a multi-agent orchestration framework.

Where it helps is the step before any of that exists. While you are writing or debugging the orchestration code with a coding agent such as Claude Code or Codex, you can lock that Thread to the project folder with the macOS kernel sandbox, and put the model or cloud key your code needs into the Vault, so the coding agent editing your Python never reads the key itself.

Know the limit. The folder lock is opt-in and it does not restrict the network, so a locked agent can still send out anything it is allowed to read. An unpinned Vault key can still be read by the program that was run with it. Forkbench governs the folders you lock and the keys you vault while you build, not the multi-agent system you deploy afterward.

  • Forkbench supervises the coding agent that writes your framework, not the framework itself.
  • Folder lock: opt-in, does not restrict the network.
  • Vault key: still readable by the program it was handed to once unpinned.

A build order that holds up

Decide the orchestration library and the model layer at the same time, because the choice of model decides whether offline is even possible. Then add the four safeguards before you call the system done, not as a cleanup pass afterward.

Finish by cutting the network on purpose and watching what happens. A system with real safeguards fails safe when the model it expected is unreachable, instead of quietly falling back to a cloud call you never approved.

  • Pick the framework, then point it at a local model through Ollama if offline has to be literal.
  • Sandbox every command an agent runs, with a container or an OS-level sandbox.
  • Give each agent its own scoped credential in a self-hosted vault, never a shared .env file.
  • Add a human checkpoint before anything irreversible ships.
  • Cut the network and confirm the system fails safe rather than silently calling out.

Related: Evaluating the best vault systems for AI agent workplaces, What Microsoft actually ships for a private multi-agent framework, The AI coding agent sandbox blueprint, Download Forkbench

Frequently asked

  • Can CrewAI run fully offline?

    Yes, if you point it at a model you run yourself. CrewAI's own documentation covers configuring an agent to use a local model through Ollama, and at that point the request never leaves your machine.

  • Do AutoGen or LangGraph include a sandbox for the code agents run?

    No. Neither open source library ships one. LangChain sells a separate paid product, LangSmith, that adds a sandbox for agent-generated code, which the free library does not have.

  • Is HashiCorp Vault still open source?

    No. Since version 1.15, Vault ships under the Business Source License rather than an OSI-approved license. It still runs fully self-hosted, but you cannot resell it as a competing product.

  • What safeguard is a typical multi-agent setup most likely to be missing?

    Usually two things: one shared API key used by every agent instead of a scoped key per agent, and no sandbox around the commands an agent is allowed to run.

  • Does Forkbench run multi-agent frameworks like CrewAI or AutoGen?

    No. Forkbench runs coding agents such as Claude Code or Codex in terminals on your Mac. It is not an orchestration framework and it does not host a crew or a graph.

Keep reading