Blog
Leases and heartbeats: how distributed systems hand out work safely
A lock assumes whoever holds it will eventually let go. A lease assumes they might not, which is the one assumption that actually matches how real workers fail.
A lease is a lock with an expiry: a grant of exclusive use that lasts until a clock runs out, not until the holder says it is done. Gray and Cheriton proposed the idea in 1989 for distributed file caches, and the same shape now sits under Amazon SQS's visibility timeout and Kubernetes' Lease objects. A heartbeat renews the lease while its holder is alive; silence past the timeout returns the work to the pool. What a timeout alone cannot stop is a worker that resumes after its lease expired and writes anyway, which is the job a fencing token, a number that only ever increases, is built for.
What a lock assumes that doesn't hold in a distributed system
A lock, in the ordinary sense, is a promise: whoever acquires it releases it when done, and nothing else touches the resource until then. That works inside one process, where the thing holding the lock and the thing checking it share a fate. It stops working once holder and resource are on different machines, because now there is a third possibility besides finished and working: gone, which looks identical to working for however long it takes something else to notice.
A process that crashes while holding a lock does not release it, and neither does one merely slow or stuck on a network call. A lock has no opinion about either case; it just stays held, and the resource stays unavailable until a human steps in.
Locks vs. leases
Cary Gray and David Cheriton named the fix in a 1989 paper presented at SOSP, written for keeping distributed file caches consistent: a lease. A lease grants exclusive use for a bounded amount of time, decided by the grantor up front, not by however long the holder decides to keep it. When the time is up, so is the grant, whether or not the holder ever says so.
That single change moves the failure mode from held forever to held for at most this long. A crashed worker still stops working, but whatever it held becomes available again on a schedule nobody has to notice and intervene on by hand.
Heartbeats and timeouts
A fixed lease on its own is a bad trade. Generous enough that legitimate slow work survives it, and a crash takes that long to recover from. Short enough to recover fast, and slow-but-alive work keeps losing its claim.
A heartbeat resolves that trade without picking a side. While a worker is alive it renews the lease on a short interval well inside the timeout, so genuine work almost never comes close to losing its claim. The timeout becomes a backstop for the one case a heartbeat cannot rule out: a process still running but no longer responding, wedged or deadlocked, with nothing left to announce it.
Too short, and live work gets reclaimed from a worker simply waiting on something slow. Too long, and a real crash sits unrecovered for the whole window. There is no value that removes the trade-off, only one that fits how long work legitimately takes to go quiet.
Fencing tokens: the gap a timeout alone leaves open
A timeout answers when work comes back up for grabs. It does not answer what happens if the worker that lost its lease wakes up later and acts anyway, which is exactly what a paused process does: it has no idea time passed, and finishes what it was doing the moment it resumes, including writing out a result.
Martin Kleppmann's explanation of this is worth reading closely: picture a storage server behind the lock, and a client that acquires a lease, gets paused past its lifetime, and resumes after a second client has already acquired the same lease and started writing. Left alone, the first client's delayed write can land after the second's and silently overwrite it; neither client did anything wrong, the pause did it.
A fencing token closes that gap by making the resource, not the lock, refuse stale writes. Every time the lease changes hands, the number handed out with it goes up by one, and the thing being written to rejects any write carrying a lower token than the highest it has already seen. The paused client's write arrives eventually carrying its now-stale token, and gets rejected rather than applied.
At-least-once processing and why idempotency is the other half
A lease on a unit of work, a queued message, a claimed task, almost always comes with a promise that reads delivered at least once, never exactly once. A worker can finish the work and die before reporting success, and the only reasonable thing the system can do is hand that work to someone else, who may well redo something that already happened.
That makes the lease half the story. The other half is making the work safe to redo: a write that sets a value rather than increments it, a dedupe key that recognizes a job already processed. A lease stops two workers colliding in flight; idempotency is what lets the work survive landing twice anyway.
Two real systems built this way
Amazon SQS's visibility timeout is a lease on a single message. Receiving a message does not remove it from the queue; it hides the message from other consumers for the visibility timeout, 30 seconds by default and up to 12 hours, and unless it is deleted before that window closes it becomes visible again for the next consumer to pick up. AWS documents the consequence directly: because delivery is at-least-once, nothing stops the same message being delivered more than once inside that window, so a consumer that is not idempotent can double-process a message that was never actually lost.
Kubernetes' Lease objects, in the coordination.k8s.io API group, run the same idea cluster-wide. Every node's kubelet renews a Lease by writing a fresh renewTime on an interval, and the control plane marks a node unhealthy once that stops arriving: a heartbeat and a timeout doing the job above. The same object backs leader election for kube-controller-manager and kube-scheduler: several replicas run, only the Lease holder acts, and a replica that stops renewing hands off with no human paging anyone.
Applying this to several coding agents sharing one backlog
Point several coding agents at one shared task list and the same failure shows up, because an agent running as a subprocess in a terminal tab has all three properties the lock-vs-lease argument is about: it can be killed outright, it can hang without saying so, and its own word that the work is done is a claim, not something to take at face value.
The fix is the same shape, not a new one. Claiming a task has to be a lease, not a flag only the claimant can clear, or a crashed agent's claimed task sits as dead weight until a person notices and frees it by hand. A hard exit should release the claim immediately; a hang or a wedge needs a timeout as the backstop. And a claim has to be checked against who actually holds it, the same job a fencing token does, or an agent that resumes after losing its claim can write over whoever picked the task up next.
How this works in Forkbench
Forkbench's teamwork board, the shared goal and task backlog every Thread carries, treats a claim as exactly this: a lease the board grants and can reclaim, not a promise an agent keeps. Two separate paths end one. A hard signal, the pane closing or the agent's process exiting while the tab stays open, expels the claim immediately, no timer involved. A 120-second lease is the backstop for everything else, renewed by any real interaction with the board and by a background check roughly every 30 seconds for a session still alive, so a working agent never comes close to it.
The board does the fencing job too, by identity rather than a counter: a claim binds a task to the specific agent holding it, so an agent that wakes up after its lease has lapsed and writes to a task it still thinks it owns gets rejected outright, instead of silently overwriting whoever claimed it next. Forkbench states the honest cost of a lease rather than hiding it: 120 seconds is a guess, not a guarantee, so a slow-but-alive agent can occasionally lose a claim it did not deserve and do a little wasted work until its own write gets rejected, a smaller failure than a lock nothing can ever reclaim without a human noticing. The full breakdown walks through the trade.
Related: How the teamwork board survives an unreliable agent, Hand off tasks between AI coding agents, Share a coding agent backlog with your team
Frequently asked
What's the difference between a lock and a lease?
A lock is held until whoever has it explicitly releases it. A lease is held until a clock the grantor set runs out, released or not. That matters anywhere the holder can disappear without telling anyone: a lock gives no way back in, while a lease expires on a schedule nobody has to notice and intervene on by hand. Gray and Cheriton proposed leases in 1989 for this exact failure mode in distributed caches.
What is a fencing token?
A number that only ever increases, issued fresh every time a lease changes hands, and checked by the resource being protected rather than by the lock itself. If a worker's lease expires and someone else acquires it and starts writing, the first worker's eventual write arrives carrying a stale token and gets rejected. Martin Kleppmann's writing on distributed locking is the standard explanation of why a timeout alone needs one.
Why does Amazon SQS hide a message instead of deleting it right away?
So a crashed or slow consumer does not lose it. The visibility timeout hides a received message from other consumers for a window, 30 seconds by default and up to 12 hours, and if it is not deleted before that window closes it becomes visible again for someone else to pick up. Delivery is at-least-once, so the same message can still arrive twice, which is why a consumer needs to be safe to run twice.
How long should a lease or heartbeat timeout be?
Long enough that legitimate work, renewed by a heartbeat well inside the window, almost never comes close to it, and short enough that a real crash does not sit unrecovered for the whole window. There is no universal number: Kubernetes' node heartbeat and SQS's 30-second default timeout solve different problems at different scales. Pick a value from how long work legitimately goes quiet, then widen it if live claims start getting reclaimed.
Can two workers end up thinking they both own the same task?
Briefly, yes, if one's lease lapses before it notices; that is the real cost of a lease over a lock nothing can ever reclaim. What must not happen is both writes landing. The resource, or the board, needs to check a claim against its specific holder, or a fencing token, before accepting a write, and reject the one from whoever's lease already expired.