Guide

Local vs Cloud: How Multi-Agent Framework Architecture Actually Splits

CrewAI, Microsoft's Agent Framework and Google's ADK all run as a local package first. The local versus cloud question is three questions wearing one name.

Quick Answer

Local versus cloud for a multi-agent framework is not one choice, it is three. First, where does the orchestrator process run: on your own machine or on a managed platform. Second, where does the model inference happen: a cloud API like Anthropic's or OpenAI's, or a model you host yourself with something like Ollama. Third, where do you deploy the finished system for other people to use: your own server, or a vendor's cloud runtime. CrewAI, Microsoft's Agent Framework and Google's Agent Development Kit are all open source packages you install and run locally by default, and each one also offers a managed cloud platform for the deployment question. None of that changes where the actual model call goes unless you deliberately point it at a local model instead.

Local vs cloud is really three separate questions

Ask someone whether their multi-agent setup is local or cloud and you get a different answer depending on which layer they mean. There are three layers, and they move independently.

The orchestrator is the code that decides which agent runs next, passes messages between them, and keeps track of state. It is almost always a script or a service you can run on your own laptop, because that is how CrewAI, Microsoft's Agent Framework, and Google's ADK all ship: as an installable package, not a product you can only reach through someone else's cloud.

The model is a separate question. Even an orchestrator running entirely on your machine calls out to a cloud API for every reasoning step, unless you have specifically wired in a model you host yourself.

Deployment is the third question: once the thing works, where does it run for other people. Each of the three frameworks below has a local-only mode and a managed cloud platform for exactly this question.

  • Orchestrator: almost always local by default, a package you install and run.
  • Model: cloud by default, local only if you deliberately add something like Ollama.
  • Deployment: your own server or the framework's managed cloud platform, your choice.

Microsoft's framework: one package, local by default

Microsoft's current answer to this is Agent Framework, which Microsoft Learn describes as the direct successor to both AutoGen and Semantic Kernel, built by the same teams. It combines AutoGen's simple multi-agent abstractions with Semantic Kernel's enterprise features, such as session-based state management and telemetry, and adds graph-based workflows for explicit control over how agents hand work to each other.

You install it locally, as a Python, .NET, or Go preview package, and it runs as your own process. A small but telling detail from Microsoft's own quickstart: the framework does not automatically load a .env file, so you call load_dotenv yourself or set environment variables directly, which means an Agent Framework script has the same .env habits as any other local project, good and bad.

For model providers, Agent Framework explicitly supports Microsoft Foundry, Anthropic, Azure OpenAI, OpenAI and Ollama, among others, so you can run the orchestration locally and still choose a cloud model, a different cloud model, or a fully local one, independently of the framework itself.

  • Agent Framework is the successor to AutoGen and Semantic Kernel combined, not a third separate thing.
  • Runs locally as a Python, .NET or Go package; deploys to Azure when you choose to.
  • Supports Anthropic and Ollama as model providers, not just Microsoft's own.

CrewAI: an open source crew, with a separate cloud platform

CrewAI is an open source Python framework for orchestrating a crew of agents, and it runs entirely as your own process by default. Its documentation configures API keys through a .env file, a YAML file, or directly in code, the same three options any Python project has, with no built-in vault.

CrewAI also fully supports a local model through Ollama. Point its LLM class at a local server address instead of an API key, and CrewAI's own documentation describes this as eliminating the cloud dependency entirely, which makes it one of the few setups here where fully local, model included, is a documented, supported path.

CrewAI AMP is the separate piece: a managed platform for deploying, monitoring and scaling crews in production, with a REST API, execution monitoring, and a no-code builder called Crew Studio. It sits on top of the same open source framework rather than replacing it, so moving a crew from your laptop to AMP is a deployment decision, not a rewrite.

  • Open source CrewAI: your process, your .env, your choice of model provider.
  • A documented local path: CrewAI plus Ollama removes the cloud API entirely.
  • CrewAI AMP: the managed cloud platform for running crews in production, separate from the framework.

Google's Agent Development Kit: local CLI, optional cloud deploy

Google's Agent Development Kit, ADK, is an open source framework available for Python, TypeScript, Go, Java and Kotlin. It ships its own CLI for scaffolding, building, testing and evaluating an agent locally, which is the part people searching for a local ADK tutorial usually want.

When you are ready to deploy, ADK is built to go to Google Cloud without rewriting the agent: Cloud Run, GKE, or Google's managed agent runtime are the options, giving you managed infrastructure, authentication and observability instead of running the deployment yourself. Google's own framing for this is deploy anywhere, meaning you can also containerize it yourself if you do not want Google Cloud specifically.

The shape is the same as Microsoft's and CrewAI's: build and test locally with an open package, deploy to a managed cloud platform only when you decide to, and the model provider is a separate configuration from either of those choices.

  • ADK CLI: scaffold, build, test and evaluate an agent locally, in minutes.
  • Deploy to Cloud Run, GKE or Google's managed agent runtime without changing the agent's code.
  • Open source across five languages, not locked to one cloud by default.

Where Claude fits: a model provider, and a different kind of multi-agent

Claude shows up in this picture two ways. The first is as a model provider you plug into any of the three frameworks above, the same way you would plug in OpenAI or a local Ollama model.

The second is Claude Code's own subagents, which are not the same architecture as CrewAI or Agent Framework. A subagent runs in its own isolated context window with its own tool permissions, and subagents run in parallel in the background by default, up to a configurable limit, nested up to three levels deep. That gets you real parallel work inside one coding session, but it is built into one product rather than being a general orchestration library you would use to build a different kind of application.

If you want to orchestrate Claude the way CrewAI orchestrates a crew, programmatically, across your own agents and tools, that is a job for the Claude Agent SDK or direct API calls, not for subagents, which are specifically a Claude Code feature.

  • As a model provider: Claude plugs into CrewAI, Agent Framework, and similar tools like any other model.
  • As Claude Code subagents: isolated context per subagent, parallel by default, but scoped to one coding session, not a general framework.

The security question none of this local vs cloud talk answers

Here is the part that actually matters for a secure multi-agent framework search: running the orchestrator locally does not make the setup private if the model behind it is a cloud API. Your prompts, your file contents, and anything the agents pass to each other through a cloud model all leave your machine for that provider's servers, whether the orchestrator script lives on your laptop or on a managed cloud platform.

The only setup here that is actually private end to end is a local orchestrator paired with a local model, which CrewAI plus Ollama demonstrates is a real, supported combination, not a theoretical one. Microsoft's Agent Framework and Google's ADK can both be pointed at a local model too, but check your specific agent definitions, since a default quickstart example usually points at a cloud model first.

So when you are deciding what counts as secure, ask where the model call goes, not where the orchestrator script happens to run.

  • A local orchestrator with a cloud model is still sending your data to that cloud, every reasoning step.
  • Only a local orchestrator paired with a local model keeps everything on your machine.
  • CrewAI plus Ollama is a documented example of that fully local pairing. Check before assuming the same is true of a framework's default quickstart.

Where Forkbench fits, and where it stops

Forkbench does not orchestrate a multi-agent framework's internal logic. CrewAI's crew, Agent Framework's workflow graph, and ADK's agent tree all run their own decisions about which agent goes next; that logic lives inside the process you start, not in Forkbench.

What Forkbench supervises is the process itself. If you launch a CrewAI script, an Agent Framework app, or an ADK agent from a terminal inside a Forkbench Thread, the Thread's folder lock confines what that whole process, and anything it starts, can read on your Mac. The Vault can hand the script's environment a model provider's API key by name, so it is not sitting in a .env file in the repository.

The same limits apply here as everywhere else. The folder lock is opt-in and does not restrict the network, so a locked Thread running a cloud-model crew can still send out whatever it reads. An unpinned Vault key is still readable by the script it was handed to. Forkbench governs the folders you lock and the keys you put in the Vault, not the orchestration logic running inside the process.

  • Forkbench supervises the process you start, not the framework's internal agent logic.
  • Folder lock and Vault apply to a CrewAI, Agent Framework, or ADK process the same way they apply to any other command.
  • The folder lock does not restrict the network; an unpinned Vault key is still readable by the program using it.

Related: What Microsoft actually ships for a private multi-agent framework, Vault management for multi-agent systems on GitHub, Does my code get sent to the model provider, Download Forkbench

Frequently asked

  • Can CrewAI run without an internet connection?

    Yes, if you pair it with a local model through Ollama. CrewAI's own documentation describes pointing its LLM class at a local Ollama server as eliminating the cloud dependency entirely. With a cloud model provider instead, it still needs a connection for every call.

  • Is Microsoft's AutoGen still a separate product from Semantic Kernel?

    Not as the current recommendation. Microsoft Learn describes Agent Framework as the direct successor to both, built by the same teams, combining AutoGen's agent abstractions with Semantic Kernel's enterprise features.

  • Does Google's Agent Development Kit require Google Cloud?

    No. ADK is open source and has a CLI for building and testing agents locally. Google Cloud, through Cloud Run, GKE, or its managed agent runtime, is an optional deployment target, not a requirement to build or run an agent.

  • Are Claude Code subagents the same as a multi-agent framework like CrewAI?

    No. Subagents run in isolated context windows in parallel inside one Claude Code session, which is a feature of that product. CrewAI, Agent Framework and ADK are general orchestration libraries you use to build your own separate application.

  • Does running a multi-agent framework locally keep my data private?

    Only if the model is local too. An orchestrator running on your own machine still sends your data to a cloud model provider for every call, unless you have specifically configured a local model such as one served by Ollama.

Keep reading