Blog
The Agent Client Protocol (ACP): Architecture, Wire Mechanics, and Multi-Agent Limits
The Language Server Protocol untangled programming languages from text editors a decade ago. The Agent Client Protocol does the same for coding agents, and understanding its wire mechanics explains where its boundaries sit when scaling to squads.
The Agent Client Protocol (ACP) is an open JSON-RPC 2.0 communication standard created by Zed and adopted by JetBrains that decouples autonomous AI coding agents from code editors and IDEs. Similar to how the Language Server Protocol (LSP) solved the N languages to M editors integration matrix, ACP solves the N agents to M editors matrix by standardizing session lifecycles, streaming thought updates, and editor capability delegation over stdio, HTTP, and WebSockets. Unlike LSP, ACP inverts UI control: the agent drives conversation turns and reasoning loops, while the client editor provides UI surfaces and executes host operations like filesystem edits, terminal execution, and permission prompts. ACP sits between the Model Context Protocol (MCP) for tool access and horizontal agent-to-agent protocols, but its single-session, in-memory architecture requires external durable backlogs and credential vaults when coordinating multi-agent teams.
The N x M problem: why coding agents need an LSP
Before 2016, adding language support to a code editor required a custom plugin for every language in every editor. Supporting ten languages across five editors meant building and maintaining fifty distinct plugins. Microsoft solved that integration matrix with the Language Server Protocol, giving compilers a single standardized interface to any text editor.
AI coding agents encountered the identical bottleneck by 2025. Every agent project, including Claude Code, Codex CLI, Gemini CLI, Aider, and Goose, required custom integrations or bespoke extensions to work inside Zed, JetBrains IDEs, VS Code, or Forkbench. Agent teams spent engineering effort maintaining editor UI glue, while editor teams spent engineering effort wrapping vendor-specific agent APIs.
The Agent Client Protocol (ACP) was created by Zed and joined by JetBrains to eliminate that matrix. ACP defines a vendor-neutral protocol between any coding agent and any development environment. An agent team implements ACP once to run inside every participating editor, and an editor implements ACP once to host any compatible agent.
Architecture: inversion of UI control across the boundary
LSP and ACP look similar on paper because both connect an editor to an external subprocess. Their architectural direction is completely reversed.
In LSP, the editor acts as the master. The user types a character, and the editor asks the language server for diagnostics or completions. The language server is passive, stateless between file syncs, and has no concept of a conversational turn or multi-step execution loop.
In ACP, control of the user interface is inverted. The agent controls the turn loop, the reasoning cycle, and the session state. The client editor acts as a host environment providing UI rendering and controlled access to machine capabilities.
The division of responsibility is precise. The agent decides what to think, what files to inspect, what terminal commands to invoke, and when a task is finished. The client renders the conversation view, displays streaming thoughts, visualizes diffs in editor tabs, and presents permission checkpoints to the developer.
Wire mechanics: JSON-RPC 2.0, capability negotiation, and bidirectional RPC
ACP communicates using JSON-RPC 2.0. When running locally on the same workstation, the client spawns the agent as a subprocess and exchanges newline-delimited JSON-RPC messages over standard input and output (stdio). For remote sandboxes and cloud instances, Streamable HTTP and WebSocket transports exist as drafts; stdio is the only transport the spec fully defines today.
The connection begins with an initialize exchange. The client and agent trade version numbers and declare capabilities before processing any user work. A client declares whether it supports terminal management or filesystem primitives; the agent declares whether it supports thought streaming, structured plans, or tool approvals.
Once initialized, conversational state lives inside sessions created via session/new. The client sends user instructions using session/prompt, which starts an active agent turn. If the user clicks stop or closes the window, the client interrupts execution with session/cancel.
While processing a prompt, the agent streams real-time updates back to the editor via session/update notifications. These notifications stream token chunks, internal reasoning traces, plan checklist updates, and tool call progress indicators directly into editor UI components.
What makes ACP unique among editor protocols is bidirectional client requests. The agent does not touch the workstation directly; instead, it issues JSON-RPC requests back to the client editor. The editor provides file operations through fs/read_text_file and fs/write_text_file, executes commands via terminal/create and terminal/output, and pauses for human confirmation through session/request_permission.
initialize: Client and agent negotiate protocol versions and advertise supported feature sets.session/newandsession/prompt: Client instantiates a conversation context and submits user instructions.session/update: Agent streams thoughts, text chunks, plan checklists, and tool progress to the UI.- fs/ and terminal/ RPCs: Agent requests workspace edits and terminal execution from the host editor.
session/request_permission: Agent asks the editor to prompt the human before running dangerous operations.
The protocol triad: MCP, ACP, and A2A
As agent tooling matures, three specialized protocols have emerged to handle different communication vectors. Confusing them leads to bloated architectures and misallocated security boundaries.
The Model Context Protocol (MCP), stewarded by Anthropic, is a vertical downward protocol. It connects an agent to external data sources, enterprise databases, GitHub repositories, Slack channels, and specialized tool servers.
The Agent Client Protocol (ACP), created by Zed and JetBrains, is a vertical upward protocol. It connects an agent to the developer's editing interface, workspace buffers, syntax trees, diff reviewers, and interactive confirmation dialogs.
Agent-to-Agent protocols (A2A) are horizontal protocols. They connect autonomous agents to each other for task delegation, consensus generation, peer review, and cooperative squad execution.
The boundaries between them are clean. MCP has no awareness of editor tabs or diff gutters. ACP has no awareness of Postgres databases or external SaaS APIs. A2A has no awareness of terminal panes or window layouts. A modern coding agent uses ACP to talk to its editor, MCP to query its tools, and A2A to collaborate with peer agents.
The solo-agent boundary: where ACP stops and multi-agent systems begin
ACP was engineered for a specific interaction model: one developer chatting with one agent in one editor window. That design succeeds at eliminating bespoke IDE plugins, but its fundamental assumptions break down when scaling to parallel multi-agent squads.
The first breaking point is session memory. ACP tracks task checklists inside transient session state. If the socket drops, the process crashes, or the editor restarts, that plan vanishes. Multi-agent squads require a durable backlog stored in a database or filesystem that persists across process lifecycles.
The second breaking point is self-graded completion. In an ACP turn, the agent decides on its own when a prompt is satisfied. When several agents work concurrently on related features, self-grading causes silent drift and incomplete deliveries. Multi-agent workflows require independent verification gates, automated test passes, and peer review before marking tasks complete.
The third breaking point is socket failure handling. If an ACP connection terminates mid-turn, work in progress is stranded with no recovery mechanism. Multi-agent orchestration requires transactional task leases with heartbeats and automatic lease release so orphaned work returns to the queue.
The fourth breaking point is interactive authorization. ACP relies on session/request_permission, stopping execution until a human clicks an approval dialog. Running ten background agents with interactive confirmation modals causes immediate alert fatigue and deadlock. Parallel operations require policy-based automation and least-privilege scoping instead of continuous manual prompting.
The final breaking point is credential isolation. ACP forwards commands into the client's terminal where environment variables and local secret files are exposed. In a multi-agent environment, agents must operate against sealed credential vaults that inject authorization at network boundaries without exposing plaintext keys to agent contexts.
How this works in Forkbench
Forkbench is a desktop application for managing parallel AI coding agents across multiple vendors, including Claude Code, Codex CLI, and ACP-compliant runtimes.
Forkbench honors ACP where it excels: rendering clean conversation streams, exposing terminal panes, and delegating workspace file edits without custom IDE wrappers. It then provides the multi-agent coordination layer that ACP leaves out.
In Forkbench, multi-agent work is organized into Threads. Each Thread provides git worktree isolation so parallel agents never overwrite each other's checkouts. Tasks live on a durable Teamwork board outside ephemeral agent sessions, claimed through time-bounded leases with automatic recovery if an agent stalls.
Sensitive credentials stay inside Forkbench Vault. An agent running build commands or deploying artifacts uses authorization through authenticated proxies without the secret value ever touching the agent's prompt, command arguments, or transcript logs. The protocol connects the editor; the surrounding system safeguards the fleet.
Related: Model Context Protocol (MCP) explained, Language Server Protocol (LSP) and agent architectures, Agent-to-Agent Protocol (A2A) for multi-agent systems, Durable task leases and teamwork boards, Forkbench Vault security and credential isolation
Frequently asked
What is the difference between ACP and LSP?
The Language Server Protocol (LSP) standardizes how code editors query compiler language servers for syntax trees, type definitions, and diagnostics. The Agent Client Protocol (ACP) standardizes how code editors host autonomous AI coding agents. While LSP keeps the editor in control as the master querying a passive server, ACP inverts control so the agent drives conversation turns, reasoning steps, and tool execution requests. LSP transfers language analysis; ACP transfers session states, streaming thoughts, tool authorizations, and terminal actions.
How does ACP relate to Anthropic's Model Context Protocol (MCP)?
They operate in opposite directions. MCP is a vertical downward protocol that connects an agent to tools, databases, and third-party web services. ACP is a vertical upward protocol that connects an agent to the code editor's user interface, editor buffers, and terminal instances. An agent uses MCP to fetch database schemas or query Jira tickets, and uses ACP to render diffs, stream thoughts, and prompt the user for permission inside the IDE.
Can an ACP agent execute terminal commands and edit files directly?
Only by delegating those actions back to the host editor. An ACP agent sends JSON-RPC 2.0 requests such as
fs/write_text_fileorterminal/createover the connection. The editor inspects the request, prompts the user throughsession/request_permissionif required by policy, and performs the operation in its own managed workspace. The agent process never bypasses the client editor to alter files or run commands silently.Why can't ACP alone coordinate a multi-agent squad?
ACP is strictly a one-to-one protocol between a single editor client and a single agent process. It maintains task lists in transient session memory, relies on agents to self-evaluate completion, and uses modal confirmation dialogs for permission checks. Running multiple agents requires external infrastructure: a durable teamwork board that survives socket disconnections, lease claiming to prevent duplicate work, and sealed credential vaults that keep API keys out of agent transcripts.