Blog

Leases and heartbeats: how distributed systems hand out work safely

A lock assumes whoever holds it will eventually let go. A lease assumes they might not, which is the one assumption that actually matches how real workers fail.

Quick Answer

A lease is a lock with an expiry: a grant of exclusive use that lasts until a clock runs out, not until the holder says it is done. Gray and Cheriton proposed the idea in 1989 for distributed file caches, and the same shape now sits under Amazon SQS's visibility timeout and Kubernetes' Lease objects. A heartbeat renews the lease while its holder is alive; silence past the timeout returns the work to the pool. What a timeout alone cannot stop is a worker that resumes after its lease expired and writes anyway, which is the job a fencing token, a number that only ever increases, is built for.

What a lock assumes that doesn't hold in a distributed system

A lock, in the ordinary sense, is a promise: whoever acquires it releases it when done, and nothing else touches the resource until then. That works inside one process, where the thing holding the lock and the thing checking it share a fate. It stops working once holder and resource are on different machines, because now there is a third possibility besides finished and working: gone, which looks identical to working for however long it takes something else to notice.

A process that crashes while holding a lock does not release it, and neither does one merely slow or stuck on a network call. A lock has no opinion about either case; it just stays held, and the resource stays unavailable until a human steps in.

Locks vs. leases

Cary Gray and David Cheriton named the fix in a 1989 paper presented at SOSP, written for keeping distributed file caches consistent: a lease. A lease grants exclusive use for a bounded amount of time, decided by the grantor up front, not by however long the holder decides to keep it. When the time is up, so is the grant, whether or not the holder ever says so.

That single change moves the failure mode from held forever to held for at most this long. A crashed worker still stops working, but whatever it held becomes available again on a schedule nobody has to notice and intervene on by hand.

Heartbeats and timeouts

A fixed lease on its own is a bad trade. Generous enough that legitimate slow work survives it, and a crash takes that long to recover from. Short enough to recover fast, and slow-but-alive work keeps losing its claim.

A heartbeat resolves that trade without picking a side. While a worker is alive it renews the lease on a short interval well inside the timeout, so genuine work almost never comes close to losing its claim. The timeout becomes a backstop for the one case a heartbeat cannot rule out: a process still running but no longer responding, wedged or deadlocked, with nothing left to announce it.

Too short, and live work gets reclaimed from a worker simply waiting on something slow. Too long, and a real crash sits unrecovered for the whole window. There is no value that removes the trade-off, only one that fits how long work legitimately takes to go quiet.

Fencing tokens: the gap a timeout alone leaves open

A timeout answers when work comes back up for grabs. It does not answer what happens if the worker that lost its lease wakes up later and acts anyway, which is exactly what a paused process does: it has no idea time passed, and finishes what it was doing the moment it resumes, including writing out a result.

Martin Kleppmann's explanation of this is worth reading closely: picture a storage server behind the lock, and a client that acquires a lease, gets paused past its lifetime, and resumes after a second client has already acquired the same lease and started writing. Left alone, the first client's delayed write can land after the second's and silently overwrite it; neither client did anything wrong, the pause did it.

A fencing token closes that gap by making the resource, not the lock, refuse stale writes. Every time the lease changes hands, the number handed out with it goes up by one, and the thing being written to rejects any write carrying a lower token than the highest it has already seen. The paused client's write arrives eventually carrying its now-stale token, and gets rejected rather than applied.

At-least-once processing and why idempotency is the other half

A lease on a unit of work, a queued message, a claimed task, almost always comes with a promise that reads delivered at least once, never exactly once. A worker can finish the work and die before reporting success, and the only reasonable thing the system can do is hand that work to someone else, who may well redo something that already happened.

That makes the lease half the story. The other half is making the work safe to redo: a write that sets a value rather than increments it, a dedupe key that recognizes a job already processed. A lease stops two workers colliding in flight; idempotency is what lets the work survive landing twice anyway.

Two real systems built this way

Amazon SQS's visibility timeout is a lease on a single message. Receiving a message does not remove it from the queue; it hides the message from other consumers for the visibility timeout, 30 seconds by default and up to 12 hours, and unless it is deleted before that window closes it becomes visible again for the next consumer to pick up. AWS documents the consequence directly: because delivery is at-least-once, nothing stops the same message being delivered more than once inside that window, so a consumer that is not idempotent can double-process a message that was never actually lost.

Kubernetes' Lease objects, in the coordination.k8s.io API group, run the same idea cluster-wide. Every node's kubelet renews a Lease by writing a fresh renewTime on an interval, and the control plane marks a node unhealthy once that stops arriving: a heartbeat and a timeout doing the job above. The same object backs leader election for kube-controller-manager and kube-scheduler: several replicas run, only the Lease holder acts, and a replica that stops renewing hands off with no human paging anyone.

Applying this to several coding agents sharing one backlog

Point several coding agents at one shared task list and the same failure shows up, because an agent running as a subprocess in a terminal tab has all three properties the lock-vs-lease argument is about: it can be killed outright, it can hang without saying so, and its own word that the work is done is a claim, not something to take at face value.

The fix is the same shape, not a new one. Claiming a task has to be a lease, not a flag only the claimant can clear, or a crashed agent's claimed task sits as dead weight until a person notices and frees it by hand. A hard exit should release the claim immediately; a hang or a wedge needs a timeout as the backstop. And a claim has to be checked against who actually holds it, the same job a fencing token does, or an agent that resumes after losing its claim can write over whoever picked the task up next.

How this works in Forkbench

Forkbench's teamwork board, the shared goal and task backlog every Thread carries, treats a claim as exactly this: a lease the board grants and can reclaim, not a promise an agent keeps. Two separate paths end one. A hard signal, the pane closing or the agent's process exiting while the tab stays open, expels the claim immediately, no timer involved. A 120-second lease is the backstop for everything else, renewed by any real interaction with the board and by a background check roughly every 30 seconds for a session still alive, so a working agent never comes close to it.

The board does the fencing job too, by identity rather than a counter: a claim binds a task to the specific agent holding it, so an agent that wakes up after its lease has lapsed and writes to a task it still thinks it owns gets rejected outright, instead of silently overwriting whoever claimed it next. Forkbench states the honest cost of a lease rather than hiding it: 120 seconds is a guess, not a guarantee, so a slow-but-alive agent can occasionally lose a claim it did not deserve and do a little wasted work until its own write gets rejected, a smaller failure than a lock nothing can ever reclaim without a human noticing. The full breakdown walks through the trade.

Related: How the teamwork board survives an unreliable agent, Hand off tasks between AI coding agents, Share a coding agent backlog with your team

Frequently asked

Keep reading

Sources