Guide

What Offline Actually Means for a Supervised AI Agent

Offline gets used for three different setups, and only one of them actually removes the network. Here is how to tell which one you need before you build around it.

Quick Answer

An offline supervised AI agent usually means one of three things. The first is a coding agent running on a fully local model, through something like Ollama or LM Studio, with no API call to any cloud provider once the model is downloaded. The second is a sandboxed agent, such as Claude Code with its bash sandbox turned on, where shell commands are blocked from the network by default but the agent's own reasoning still has to reach its model provider over the one connection the sandbox lets through. The third is supervision that keeps working locally, watching CPU use and output from a process on your own machine, even when the agent being watched has gone quiet because it lost its own connection. Pick the one you actually need before you pick tools for it, because the fix for each is different.

Offline means three different setups, not one

People search for an offline supervised AI agent for three different reasons, and the fix is different for each one.

The first reason is privacy: you want a coding agent that never sends your code anywhere, full stop. The second is a locked-down network: you are on a flight, at a client site, or on a machine with a firewall that blocks most outbound traffic, and you want the agent to still work within whatever it is allowed to reach. The third is resilience: you want to know the agent is still being watched correctly even when your Wi-Fi drops for a minute.

Only the first one is actually offline in the strict sense. The other two still need a network connection somewhere, just a narrower one than usual, or a connection to something other than the agent itself.

  • Full privacy: a local model, no cloud call at all.
  • A locked-down network: the agent's shell commands are blocked from reaching most hosts, but the model call still goes out.
  • Resilience: whatever is watching the agent keeps working even if the agent's own connection drops.

Running the agent itself with no cloud at all

A genuinely offline coding agent needs a model running on your own hardware. Ollama and LM Studio both do this: you pull a model once, and after that it serves requests from your machine without sending anything out. Ollama runs on macOS, Windows and Linux, and its own site describes local operation as complete privacy, whether you run it alone or alongside cloud models.

You then need an agent harness that points at that local model instead of a cloud API. CrewAI, the open source Python framework for multi-agent crews, documents this directly: install Ollama, run a model, then point CrewAI's LLM class at the local server address instead of an API key. CrewAI's own docs describe this as removing the cloud dependency entirely.

The catch is quality. A model you can run on a laptop is smaller than Claude or GPT, and it shows in coding work: more wrong answers, more need for you to check what it did. Offline buys you privacy, not the same level of capability.

  • Pull a model with Ollama or LM Studio once; after that, no download is needed to run it again.
  • Point your agent harness at the local model's address instead of a cloud API key.
  • Expect a smaller model to make more mistakes, so review its output more, not less.

Why Claude Code and Codex cannot go fully offline

If your agent is Claude Code, Codex, or another tool built on a frontier model, offline in the strict sense is not available, and no setting changes that. Anthropic's own documentation for Claude Code is direct about it: isolation does not change what is sent to the model. Your prompts and the files Claude reads are transmitted to the Anthropic API, or whatever provider you have configured, with or without a sandbox around the session.

That holds even inside the strongest isolation Claude Code offers, a dedicated virtual machine or a cloud session. The box gets more locked down. The one connection the model needs does not go away, because the model itself is not on your machine.

If a locked-down network is the actual constraint, which is common on a corporate machine, the fix is not trying to go fully offline. It is working out exactly which host the agent needs and allowing only that one.

  • A sandbox, a container, or a VM changes what the agent can reach. It does not change what the model call sends.
  • Ask what the agent needs to reach, not whether it can go fully dark.

The sandbox that blocks the network without blocking the model

This is the setup behind a search like pulse supervised ai agent offline: a sandbox that denies network access by default and opens only a narrow exception for the model call itself.

Anthropic's sandbox runtime for Claude Code works exactly this way. By default it denies all network access and confines writes to a small set of built-in paths. You then explicitly allow the domains the session needs, starting with api.anthropic.com, or your configured provider's endpoint if you are not using Anthropic's own API. Everything else a shell command tries to reach stays blocked.

That is a real, useful kind of offline. A command the agent runs cannot phone home to a random server, upload a file, or fetch a script from the internet. The one exception is deliberate and narrow: the single connection the agent needs to think at all.

  • Network denied by default; you allow specific domains, not a blanket connection.
  • The model provider's endpoint is the one exception that has to stay open for the agent to function.
  • Everything else a shell command tries to reach is blocked, which is the actual privacy win here.

What keeps working when your own connection drops

Supervision is a different question from whether the agent can think. Watching a process for CPU use and output does not itself need a network connection, because it reads from your own machine, not from anywhere else.

So if your Wi-Fi drops mid-session, the thing watching the agent does not go blind. What you will see instead is the agent itself going quiet: no new output, no CPU activity, because it cannot reach its model to get a next step. A supervision tool doing its job correctly shows you that silence honestly, rather than pretending the agent is still working.

Forkbench's pulse works this way. It is driven by a tab's CPU use and output rate, not by a count of tokens spent, so it is a local measurement of local activity. If the agent behind a tab loses its connection, the pulse goes quiet because there is genuinely nothing happening, which is the correct signal to notice.

  • A process monitor reading CPU and output needs no network of its own.
  • A quiet pulse during a dropped connection is accurate, not broken.
  • The pulse measures activity, never tokens, so it will never tell you how much of your quota an offline retry burned.

Where Forkbench fits, and where it does not

Forkbench does not give you a local model, and it does not make Claude Code or Codex work without a connection. What it does is run whatever you start, local model or cloud agent, in a real terminal on your Mac, with the same folder lock and Vault either way.

If you are running a fully local setup with Ollama, Forkbench's folder lock still confines that terminal's Thread to its own folders, so the experiment stays boxed in even though the model itself is local. The Vault still lets a command use a secret by name without the value landing in the prompt or the transcript.

Know the limits before you rely on them. The folder lock is opt-in and does not restrict the network, so a locked session with network access can still send out whatever it is allowed to read. An unpinned Vault key can still be read by the program it was handed to. Neither of these changes depending on whether the model behind the Thread is local or in the cloud.

  • Folder lock and Vault apply the same way to a local model and a cloud one.
  • The folder lock does not restrict the network on its own.
  • An unpinned Vault key is still readable by the program that received it.

A setup you can actually finish

Work out which of the three offline meanings you actually need before you pick tools for it. Privacy, a narrow network, and resilient supervision are solved differently, and trying to solve all three with one setting usually solves none of them well.

Once you know which one you need, the rest is mechanical.

  • Want full privacy: install Ollama or LM Studio, pick a model that fits your hardware, and point your agent harness at the local address.
  • Want a locked-down network: turn on your agent's sandbox and allow only the model host it needs.
  • Want supervision that survives a dropped connection: use a monitor that reads local process activity, not one that depends on the same connection the agent uses.
  • Either way, keep secrets in the Keychain or a secrets manager, not in a file the agent can read.

Related: How to run AI agents offline on your work desktop, Managing Claude AI agents offline, Does my code get sent to the model provider, Download Forkbench

Frequently asked

  • Can an AI coding agent run completely offline?

    Only if it uses a model running on your own machine, through something like Ollama or LM Studio. A cloud-model agent such as Claude Code or Codex always needs one connection open to its model provider, no matter how locked down the rest of the session is.

  • Does Claude Code need an internet connection?

    Yes. Anthropic's own documentation says isolation does not change what is sent to the model: your prompts and the files Claude reads are transmitted to the Anthropic API, or your configured provider, with or without a sandbox.

  • What does offline mean if a sandbox still allows the model call?

    It means the agent's shell commands are blocked from the network by default, with one narrow exception open for the connection the model itself needs. Everything else a command tries to reach stays closed.

  • Does supervision keep working if my network drops?

    A monitor that reads a process's CPU use and output locally keeps working, because it never needed the network itself. What changes is the agent: it goes quiet because it cannot reach its model, and a correct monitor shows you that honestly.

  • Does Forkbench run agents without the internet?

    Forkbench does not supply a local model. It runs whatever you start, including an Ollama-based setup, in a terminal with the same folder lock and Vault either way, but the agent's own need for a network connection depends on the model behind it, not on Forkbench.

Keep reading