When a Redis Lock Is Enough, and When It Isn't
Distributed locks look simple, `SET key value NX EX ttl`, and the failure modes are not. A lock that expires while the job holding it is still running, or a network partition that makes two clients briefly believe they both hold the lock, can cause real damage: double-charged customers, corrupted batch jobs, duplicate emails sent.
Here's how to decide between a single-instance Redis lock, Redlock, and skipping Redis entirely in favor of a database-level lock, based on what actually happens if the lock fails.
The Single-Instance Lock and Its One Real Failure Mode
`SET NX EX` against a single Redis instance is simple and works for most cases: acquiring a lock, setting a TTL so it self-expires if the holder crashes, releasing it when done. The failure mode is the TTL itself. If the job holding the lock runs longer than the TTL, network hiccup, GC pause, slow downstream call, the lock expires while the job still thinks it holds it, and a second client can acquire it and start the same work concurrently. A watchdog thread that extends the TTL while the job is still alive closes most of this gap, but adds its own complexity.
Redlock: What It Buys You and What It Doesn't
Redlock acquires the lock across a majority of independent Redis instances instead of one, which is meant to survive a single instance failing or partitioning. Martin Kleppmann's well-known critique of the algorithm points out that it still doesn't protect against the TTL-expiry problem above, a process pause can make a client believe it holds the lock after it's actually expired regardless of how many instances agreed to grant it. Redlock buys you resilience against a single Redis node going down; it does not buy you a correctness guarantee against process pauses, and teams that adopt it expecting the latter are solving the wrong problem.
Fencing Tokens Are the Actual Fix for the TTL Problem
A fencing token is a monotonically increasing number returned with each lock acquisition, and every downstream system the lock protects, the database write, the API call, checks that the token it's seeing is the highest one it's seen before accepting the operation. This means even if two clients briefly both believe they hold the lock, only the one with the higher token's operations get accepted downstream. This requires the protected resource to support the check, which a Redis lock alone can't guarantee, so it's an added integration point, not a drop-in fix.
When to Skip Redis and Use a Database Advisory Lock Instead
If the work you're protecting already happens inside a transaction against your primary database, Postgres's `pg_advisory_lock` or a `SELECT ... FOR UPDATE` often does the job with fewer moving parts and no separate infrastructure to reason about. It's scoped to the database session, tied to the same transaction boundary as the work itself, and doesn't introduce a second system that can fail independently. Reach for Redis-based locking when the work spans multiple services or systems that don't share a transaction, not as a default for anything that needs mutual exclusion.
Designing for the Lock to Fail Safe
Whatever mechanism you pick, the job it protects should be idempotent, safe to run twice, wherever possible, so a lock failure degrades to "did the work an extra time" instead of "corrupted state." A distributed lock is a mitigation for a race condition, not a substitute for making the underlying operation safe to retry. Teams that treat the lock as the only line of defense are one clock skew or GC pause away from the exact failure they built the lock to prevent, which is why the idempotency work is worth doing even after the locking mechanism itself looks solid.
Match the mechanism to what a failed lock would cost you:
- Single-instance SET NX EX lock: simple and fine for most cases, but watch for a TTL that expires while the job is still running.
- Redlock: acquires the lock across a majority of independent Redis instances to survive one instance failing or partitioning.
- Fencing tokens: a monotonically increasing number that every downstream system checks, accepting only the highest token it has seen.
- Database advisory lock or SELECT FOR UPDATE: often enough when the protected work already runs inside a transaction on your primary database.
- Idempotent jobs: make the work safe to run twice, so a lock failure means duplicated effort instead of corrupted state.
Testing a Lock's Failure Modes Before Production Does It For You
Simulate the failure you're actually worried about in a lower environment: kill the process holding the lock mid-job and watch whether a second worker correctly picks up the work, or pause a process artificially to see whether its lock expires while it still believes it holds it. Most teams only discover their locking strategy has a gap during a real incident, when the failure mode is expensive to diagnose under pressure instead of cheap to reproduce deliberately ahead of time. A short, repeatable chaos test for your specific locking library is worth more than reading another article about Redlock's theoretical edge cases.
What Good Looks Like
A distributed lock should fail safe: if you're not certain the lock is still held, treat the operation as unprotected and stop, rather than assume it's fine.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need Redlock if we're already running Redis in a highly available cluster?
Usually not. A managed Redis cluster with automatic failover already handles single-node failure; Redlock's specific value is coordinating across independent Redis deployments, which most teams don't run. A single-instance lock with a watchdog for TTL extension covers most production cases.
What's the simplest way to add a fencing token without a big rewrite?
Store a monotonic version or timestamp alongside the resource the lock protects, and have the write path compare it before committing, similar to optimistic concurrency control. This is often a smaller change than it sounds, since many systems already have an updated_at or version column to build the check on.
How long should a lock's TTL be relative to expected job duration?
Long enough to cover the job's expected worst case with margin, then paired with a watchdog that extends it while the job is still running, rather than trying to guess a single fixed TTL that covers every scenario. A TTL set to the average case will expire under load, which is exactly when the lock matters most.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Reducing Vendor Lock-In Without Going Multi-Cloud
Vendor lock-in mitigation is mostly about contract terms and data portability, not a full abstraction layer. Here is where to actually spend the effort.
Distributed Locks With Redis: What Actually Fails
Why a simple Redis lock isn't mutual exclusion, what a fencing token fixes and doesn't, and a safer default for most small engineering teams.
Where Production Deployment Budgets Quietly Leak
The recurring places engineering teams overspend on production deployment architecture, and a practical order for fixing them without a full rebuild.
Choosing a Caching Strategy Without Overbuilding It
A decision guide for picking a caching approach that matches your actual read patterns, instead of defaulting to the most complex option available.
Redis Locks, Postgres Advisory Locks, or etcd: Picking a Locking Pattern
How to choose between a Redis lock, a Postgres advisory lock, and a dedicated coordination service like etcd when two processes must not run the same job twice.
Redis Lock, Postgres Advisory Lock, or Zookeeper: Picking One
A comparison of the three common ways to coordinate distributed locks: Redis-based locks, Postgres advisory locks, and a dedicated coordination service.