Distributed Locks With Redis: What Actually Fails
Two workers both believe they hold the lock, and both process the same payout, because a single Redis key with an expiry isn't actually mutual exclusion once a client pauses or a network hiccups at the wrong moment.
Most teams reach for a distributed lock to solve exactly one problem: making sure only one worker acts on a given resource at a time. Getting there safely means understanding what a lock protects and what it doesn't.
The Naive SETNX Lock and Why It Isn't Enough
The common pattern is a single SET with NX and an expiry: whoever sets the key first holds the lock until it expires. That's fine as a first pass, but it has a real gap. If the client holding the lock pauses for garbage collection, a slow disk write, or a network delay longer than the lock's expiry, the lock releases while that client still believes it's holding it, and a second client can acquire it and start working on the same resource at the same time.
What a Fencing Token Fixes and What It Doesn't
A fencing token attaches a monotonically increasing number to each lock acquisition, and the protected resource itself checks that the token it receives is newer than the last one it accepted. That stops the stale first client's write from landing after a second client has already taken over, because the resource rejects the old token.
It only works, though, if the resource you're protecting can actually validate a token. A call to a third-party payment API or an external webhook usually can't, so a fencing token protects your own database rows but not everything a locked operation touches.
When Redlock Is Overkill and When a Single Instance Isn't Enough
Redlock's multi-instance quorum approach protects against one Redis node going down mid-lock, at the cost of real operational complexity: running and maintaining several independent Redis instances just for locking. For most small and mid-sized teams, the failure you're actually protecting against is a slow or paused worker, not a Redis node failing over at the exact wrong millisecond, and a single well-configured Redis instance with a fencing token covers that case without the extra infrastructure.
Reach for a quorum-based approach when the cost of a duplicate execution is severe enough, real financial loss, not just an annoying retry, that you're willing to run the extra infrastructure to close that last gap.
A Safer Default for Most Teams
Set the lock's expiry a comfortable margin longer than your worst-case job duration under load, not your average case. Attach a fencing token so the resource itself can reject a stale holder. And design the underlying operation to be idempotent, so that if a duplicate execution ever does slip through, it's harmless rather than catastrophic. That last part is the fix most teams skip, and it's the one that actually saves you when the lock itself has a bad day.
Where This Actually Breaks in Production
- A lock expiry shorter than the job's actual runtime under load, not just under normal conditions.
- No fencing token, so the resource can't tell an old lock holder trying to finish from the current legitimate one.
- Treating the lock as the only safety net instead of also making the underlying operation idempotent.
- Forgetting to release the lock on the failure and cleanup path, so a crashed worker holds the resource hostage until the expiry finally runs out.
Testing a Lock the Way It Actually Fails
Most teams test a lock by confirming a second client can't acquire it while the first one holds it, then stop. That test never exercises the failure mode that actually causes incidents: a client that pauses mid-operation past the expiry and then resumes, unaware that its lock is already gone.
Add a test that deliberately holds the lock past its expiry, lets a second client acquire it and start work, then lets the first client attempt its write anyway. If the fencing token rejects that stale write, the lock is doing its job. If it doesn't, you've found the gap before a customer does, which is a far cheaper place to find it.
What Good Looks Like
Good distributed locking means a duplicate execution, if it ever happens, is harmless because the underlying operation is idempotent, not just prevented in the average case by the lock itself.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is Redlock overkill for a small team?
Often, yes. Redlock's multi-instance quorum protects against a single Redis node failing mid-lock, which is a real but rare failure mode. For most small teams, a single Redis instance with a fencing token and a realistic expiry covers the actual risk, a slow or paused worker, without the operational cost of running several Redis instances.
What is a fencing token and why does the lock alone not cover the same case?
A fencing token is a number that increases every time the lock is acquired, checked by the resource itself against the last token it accepted. The lock alone only controls who thinks they hold it; the token lets the resource reject a stale holder's write even after the lock has already changed hands.
How long should a distributed lock's expiry be?
Set it comfortably longer than your worst-case job duration under real load, not your typical case. Too short and a legitimate holder loses the lock mid-task; too long and a crashed worker blocks everyone else for that whole window. Pair whatever expiry you pick with an idempotent underlying operation as a backstop.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
The Real Cost of Vendor Lock-In (and When to Actually Migrate)
How to tell whether a vendor dependency is a real business risk or just an inconvenience, and what a realistic exit actually costs you.
When a Redis Lock Is Enough, and When It Isn't
Single-instance locks, Redlock, fencing tokens, and when to skip Redis entirely for a database advisory lock instead. A decision guide for CTOs.
Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)
A walkthrough of a real double-charge bug caused by a Redis lock's TTL expiring mid-job, and the fencing-token pattern that actually fixes it.
Redis Locks, Postgres Advisory Locks, or etcd: Picking a Locking Pattern
How to choose between a Redis lock, a Postgres advisory lock, and a dedicated coordination service like etcd when two processes must not run the same job twice.
Redis Lock, Postgres Advisory Lock, or Zookeeper: Picking One
A comparison of the three common ways to coordinate distributed locks: Redis-based locks, Postgres advisory locks, and a dedicated coordination service.
Picking a Distributed Lock That Won't Let Two Jobs Silently Run at Once
A comparison of distributed locking approaches for data pipelines, including where each one quietly fails under real conditions like network partitions.