Guide
How to Run AI Agents Offline on Your Work Desktop
Running an AI agent fully offline keeps your code and credentials on your own machine. Here is the current stack, what you trade away, and where it breaks down.
To run an AI coding agent offline on your work desktop, you need a local model runner such as Ollama or LM Studio to host a model on your own hardware, and an agent interface such as the Cline extension for VS Code or the Aider terminal tool to read files, write code and run commands. Download a current coding model that fits your machine, for example Qwen3-Coder-30B or Qwen2.5-Coder-7B, point the agent at the local API endpoint Ollama exposes on localhost:11434, and the entire task runs without a network connection. Your code and any credentials you use never leave the machine, which is why this setup is the standard answer for regulated or air-gapped work. The cost is reasoning power: a model you can run on a laptop is smaller than a hosted flagship model, so it handles small, well-scoped changes better than a sweeping rewrite.
Why developers take an agent off the cloud
The usual reason to run an agent locally is data handling, not preference. A cloud-based agent sends your code, error messages and sometimes environment variables to a third party server to be processed. In a regulated industry, or on a project under an NDA, that is not always something you are allowed to do.
Running locally keeps the whole loop on your machine: the agent reads your files, decides what to change, and writes the result, all without a request leaving your network. For codebases under strict data residency rules, or for an engineer in a facility without general internet access, that is the only option that is compliant at all.
A second, smaller reason is resilience. A local setup keeps working through a flaky connection or an API outage, because it never depended on either in the first place.
- A local agent never sends your code, logs or environment variables off the machine.
- This is often the only option that satisfies data residency or air-gapped requirements.
- Local setups are unaffected by an internet outage or a provider's API being down.
What you give up by staying local
The trade is reasoning power for privacy. The largest hosted models right now, Anthropic's current Claude models and OpenAI's GPT-5 series among them, run on data center hardware with far more memory than a laptop, and tend to handle large, multi-file changes better.
A model you can realistically run locally is smaller. Qwen3-Coder-30B-A3B is a mixture-of-experts model with 30 billion total parameters but only about 3.3 billion active per token, a 262,000-token context window, and it needs roughly 17 to 20 gigabytes of memory at a common quantization level. Mistral's Devstral-24B was built specifically for agent-style tasks, with a 128,000-token context window and a roughly 16 gigabyte footprint, though every Devstral release is now marked deprecated or past retirement on Mistral's own model lifecycle page; the open weights stay downloadable through Ollama, Mistral just is not actively maintaining the line. Qwen2.5-Coder-7B is the lightweight option, under 5 gigabytes, with a 32,000-token context window, and fits comfortably on an 8 gigabyte GPU.
In practice, a local model does well on a contained change: fix this function, write this test, refactor this one file. It struggles more with a change that needs to hold a large, unfamiliar codebase in mind at once. Many teams route the sensitive, narrow work to a local model and send the larger, non-sensitive refactors to a hosted one.
- Hosted flagship models still lead on large, multi-file changes that need a big context window.
- Qwen3-Coder-30B-A3B needs about 17 to 20GB of memory and offers a 262,000-token context window.
- Devstral-24B targets agent tasks and runs in around 16GB, though Mistral now lists every Devstral release as deprecated or retired; Qwen2.5-Coder-7B fits under 8GB.
Setting up the model runner and the agent
Ollama is the standard local model runner on macOS and Linux. Pulling and running a model is one command, for example ollama run qwen3-coder, and Ollama exposes an OpenAI-compatible API at localhost:11434 that any agent built against the OpenAI client library can point at without changing its own code.
LM Studio does the same job through a graphical interface if you would rather not live in the terminal. For the agent side, the Cline extension for VS Code is a common choice: open its API configuration, set the provider to Ollama or to an OpenAI-compatible endpoint, choose your local model, and raise the context window setting to match what the model supports.
If you came here searching something like ai agents office pulse local expecting a specific product called Pulse built for this, that is not a term tied to any offline coding agent tool. The closest real matches are enterprise monitoring and incident platforms built for infrastructure teams, a different job entirely. What you actually want for offline coding is the runner-plus-agent pairing covered here.
- ollama run qwen3-coder pulls and runs a model with one command; Ollama serves it at localhost:11434.
- LM Studio offers the same local hosting through a graphical interface.
- Cline's API configuration lets you point the agent at Ollama and set the model's context window.
Rolling this out for a whole engineering team
Standardizing local agents across a team starts with the hardware. A machine needs enough unified memory, or GPU VRAM, to hold the model plus its context window with room left over for everything else running. Apple's current M5 Max MacBook Pro supports up to 128 gigabytes of unified memory, which comfortably covers even the larger local coding models with headroom to spare.
Most IT teams that go this route distribute a short list of approved model weights internally rather than letting every developer pull whatever they find, so the models in use have actually been reviewed. The agent extension is pre-configured to point at the approved local endpoint, so a new machine is ready without a developer having to set any of this up themselves.
The result is that every developer gets the same AI assistance under the same constraints, instead of a mix of ad hoc local setups with different models and different levels of review.
- Provision machines with enough unified memory or VRAM to hold the model and its context comfortably.
- An M5 Max MacBook Pro's 128GB of unified memory covers even the larger local coding models.
- Standardize on a short, reviewed list of model weights rather than letting everyone pick their own.
Staying in the loop even when the agent is offline
Offline does not mean unsupervised. A local agent can still run a command that deletes the wrong directory or breaks a local database, so the same human-in-the-loop habits apply as with a hosted agent.
Tools like Cline ask for explicit approval before running a terminal command or writing a file by default, which gives you a chance to catch a bad command before it runs rather than after. Read what it is proposing, especially around dependency changes or anything that touches more than the file you expected.
Running the project inside a container is worth doing even for a fully offline setup. If a command does go wrong, the damage is contained to the container instead of reaching your actual files.
- Require explicit approval for terminal commands and file writes, the same as with a hosted agent.
- Read a proposed command before approving it, especially for dependency or multi-file changes.
- A container around the project limits the damage if a local command does go wrong.
How Forkbench handles offline work, and what stays true either way
Forkbench is a desktop app that runs coding agents in an isolated terminal on your Mac. Point it at a local model through Ollama and the agent operates entirely offline. Forkbench's Vault also keeps secrets, like a local database password, in your macOS Keychain, so a local agent can use one without it sitting in a .env file or appearing in the chat transcript.
Two things stay true regardless of whether the model is local or hosted. Forkbench's folder lock is opt-in and does not restrict the network by itself, so if you later connect the same Thread to a cloud model, it will reach the internet. And an unpinned Vault key can still be read by the program that ran, so the Vault protects what the agent sees, not what a command does with a secret once it has it.
For a team that wants both full offline privacy and organized secret handling, that combination covers the common case. If all you need is code completion without an agent running commands on your behalf, a simpler offline editor extension is probably a better fit than a full terminal environment.
- Forkbench runs agents in an isolated terminal and works with a local model through Ollama.
- Its Vault keeps secrets in the Keychain instead of a .env file or the chat transcript.
- The folder lock is opt-in and does not restrict the network; an unpinned key is still readable by the program that runs.
Related: Does my code get sent to the model provider, How Forkbench handles your data, How to sandbox Claude Code on macOS, Download Forkbench
Frequently asked
Can I run an AI coding agent completely offline?
Yes. A local model runner like Ollama, paired with an agent interface like Cline or Aider, does the full read, write and run loop on your own machine with no network connection needed.
What hardware do I need for a local AI agent?
It depends on the model. A small model like Qwen2.5-Coder-7B fits in under 8GB of memory, while Mistral's Devstral-24B, now deprecated on Mistral's own lifecycle page but still downloadable through Ollama, needs around 16GB, and Qwen3-Coder-30B-A3B needs roughly 17 to 20GB.
Are local coding models as capable as hosted ones?
Not on large, multi-file changes. They do well on contained tasks, like one function or one test file, and the gap shows up most on changes that need a big model to hold a lot of context at once.
How does Forkbench change local agent security?
Its Vault keeps secrets in the macOS Keychain instead of a plain file, but its folder lock is opt-in and does not block network access, and an unpinned key can still be read by whatever program ran.