Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Distributed Locks With Redis: What Actually Fails

Two workers both believe they hold the lock, and both process the same payout, because a single Redis key with an expiry isn't actually mutual exclusion once a client pauses or a network hiccups at the wrong moment.

Most teams reach for a distributed lock to solve exactly one problem: making sure only one worker acts on a given resource at a time. Getting there safely means understanding what a lock protects and what it doesn't.

The Naive SETNX Lock and Why It Isn't Enough

The common pattern is a single SET with NX and an expiry: whoever sets the key first holds the lock until it expires. That's fine as a first pass, but it has a real gap. If the client holding the lock pauses for garbage collection, a slow disk write, or a network delay longer than the lock's expiry, the lock releases while that client still believes it's holding it, and a second client can acquire it and start working on the same resource at the same time.

What a Fencing Token Fixes and What It Doesn't

A fencing token attaches a monotonically increasing number to each lock acquisition, and the protected resource itself checks that the token it receives is newer than the last one it accepted. That stops the stale first client's write from landing after a second client has already taken over, because the resource rejects the old token.

It only works, though, if the resource you're protecting can actually validate a token. A call to a third-party payment API or an external webhook usually can't, so a fencing token protects your own database rows but not everything a locked operation touches.

When Redlock Is Overkill and When a Single Instance Isn't Enough

Redlock's multi-instance quorum approach protects against one Redis node going down mid-lock, at the cost of real operational complexity: running and maintaining several independent Redis instances just for locking. For most small and mid-sized teams, the failure you're actually protecting against is a slow or paused worker, not a Redis node failing over at the exact wrong millisecond, and a single well-configured Redis instance with a fencing token covers that case without the extra infrastructure.

Reach for a quorum-based approach when the cost of a duplicate execution is severe enough, real financial loss, not just an annoying retry, that you're willing to run the extra infrastructure to close that last gap.

A Safer Default for Most Teams

Set the lock's expiry a comfortable margin longer than your worst-case job duration under load, not your average case. Attach a fencing token so the resource itself can reject a stale holder. And design the underlying operation to be idempotent, so that if a duplicate execution ever does slip through, it's harmless rather than catastrophic. That last part is the fix most teams skip, and it's the one that actually saves you when the lock itself has a bad day.

Where This Actually Breaks in Production

  • A lock expiry shorter than the job's actual runtime under load, not just under normal conditions.
  • No fencing token, so the resource can't tell an old lock holder trying to finish from the current legitimate one.
  • Treating the lock as the only safety net instead of also making the underlying operation idempotent.
  • Forgetting to release the lock on the failure and cleanup path, so a crashed worker holds the resource hostage until the expiry finally runs out.

Testing a Lock the Way It Actually Fails

Most teams test a lock by confirming a second client can't acquire it while the first one holds it, then stop. That test never exercises the failure mode that actually causes incidents: a client that pauses mid-operation past the expiry and then resumes, unaware that its lock is already gone.

Add a test that deliberately holds the lock past its expiry, lets a second client acquire it and start work, then lets the first client attempt its write anyway. If the fencing token rejects that stale write, the lock is doing its job. If it doesn't, you've found the gap before a customer does, which is a far cheaper place to find it.

Executive Capability Standard

What Good Looks Like

Good distributed locking means a duplicate execution, if it ever happens, is harmless because the underlying operation is idempotent, not just prevented in the average case by the lock itself.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map every job in your system that currently uses a lock and check whether its underlying operation would actually be safe if it ran twice.
2. Do Manually:Manually add a fencing token to your highest-risk locked resource, the one where a duplicate write would be expensive.
3. Delegate:Assign one engineer to own the lock library used across services so expiry and token conventions stay consistent.
4. Automate:Add automated tests that simulate a stale lock holder finishing late, to confirm the fencing token actually rejects it.
5. Buy:Bring in a backend specialist to review your highest-value locked workflows, like payouts or order fulfillment, before you scale them further.

How to Get Started

Frequently Asked Questions

Is Redlock overkill for a small team?

Often, yes. Redlock's multi-instance quorum protects against a single Redis node failing mid-lock, which is a real but rare failure mode. For most small teams, a single Redis instance with a fencing token and a realistic expiry covers the actual risk, a slow or paused worker, without the operational cost of running several Redis instances.

What is a fencing token and why does the lock alone not cover the same case?

A fencing token is a number that increases every time the lock is acquired, checked by the resource itself against the last token it accepted. The lock alone only controls who thinks they hold it; the token lets the resource reject a stale holder's write even after the lock has already changed hands.

How long should a distributed lock's expiry be?

Set it comfortably longer than your worst-case job duration under real load, not your typical case. Too short and a legitimate holder loses the lock mid-task; too long and a crashed worker blocks everyone else for that whole window. Pair whatever expiry you pick with an idempotent underlying operation as a backstop.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides