Distributed Locking With Redis: Where Redlock Actually Falls Short
A distributed lock is how you stop two processes from doing the same job at once when they're running on different machines: two workers picking up the same queued job, two schedulers both deciding it's time to run the same nightly close.
Redis is the default choice because it's already in most stacks and it's fast, but its locking guarantees are weaker than most teams assume, and knowing exactly where they're weak matters more than knowing the pattern exists.
The basic Redis lock, and why a plain set-if-not-exists isn't enough
A distributed lock with Redis starts with a conditional set that only succeeds if the key doesn't already exist, plus an expiry. The expiry matters because a worker that crashes while holding the lock must not hold it forever; it's what releases the lock automatically.
The gap in the basic version is that any process can delete any lock, since a plain delete doesn't check who set it. Fix that by setting a random value on acquire and only releasing with a script that checks the value matches before deleting, which stops a worker from releasing a lock it doesn't actually hold anymore.
Where Redlock's guarantees actually break down
Redlock extends the pattern across multiple independent Redis nodes so a single node failure doesn't silently drop your lock. Its known weak point is timing: if a process's clock pauses, a long garbage collection pause, a virtual machine migration, for longer than the lock's expiry, the lock can lapse while the process still believes it holds it, and a second process can acquire the same lock.
For most background job deduplication, that risk is acceptable because the cost of a rare double-run is low. For anything where a double-run causes real damage, like a financial settlement job, that risk usually isn't acceptable.
When to reach for a database lock instead
If you're already writing to Postgres as part of the job, a row lock or an advisory lock gives you locking with the same transactional guarantees as the rest of your write, without introducing Redis as a second source of truth for correctness. It's slower than Redis and ties the lock's lifetime to a database connection staying open, but for anything where correctness matters more than raw speed, that tradeoff is usually the right one.
Reach for Redis when the lock is protecting a side effect outside the database, like preventing two workers from both calling an external API for the same job.
A checklist before you ship a new distributed lock
- Does the lock have an expiry, and is it long enough to cover the slowest realistic run of the work it protects?
- Does release check ownership before deleting, so a stale worker can't release a lock it no longer holds?
- What happens if the lock is never released, either because a crash skipped cleanup or the expiry was set too short: does the job just not run, or does it silently double-run?
- Is the cost of a rare double-run genuinely low, or does this specific job need the stronger guarantees of a database-backed lock instead of Redis?
A worked example: two workers picking up the same job
Say a queue delivers the same job message twice, a normal occurrence under an at-least-once delivery guarantee, and two worker processes both pick it up within milliseconds of each other. Without a lock, both workers do the same expensive work, and if the job has a side effect like charging a card or sending an email, the customer sees it happen twice.
With a Redis lock keyed on the job's own ID, the second worker's acquire attempt fails immediately because the key already exists, and it can safely skip the job rather than run it. The lock's expiry has to be set with the job's realistic worst-case runtime in mind, since a job that runs longer than the lock's expiry can have a second worker pick it up midway through, right when the lock lapses.
Locks that outlive their usefulness
A lock acquired for a job that later gets refactored into something idempotent on its own no longer needs the lock at all, but teams rarely go back and remove it, so the locking overhead and a second point of failure stick around for no benefit. Periodically review which locks in the codebase are actually still protecting something that isn't already safe to run twice, and remove the ones that aren't, since every lock is one more thing that can silently fail closed and block a job that should have run.
What Good Looks Like
A good distributed locking setup uses an expiry long enough to survive the slowest realistic job run, checks ownership before releasing, and reserves Redis-based locks for work where a rare double-run is genuinely low-cost.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is Redis locking safe for financial or billing operations?
Generally not on its own. Redlock's timing guarantees can fail if a process's clock pauses longer than the lock's expiry, letting two processes briefly believe they both hold the lock. For anything where a double-run causes real damage, a database-backed lock with the same transactional guarantees as the rest of the write is usually the safer choice.
What's the difference between a plain Redis lock and Redlock?
A plain lock against a single Redis node is simple but has a single point of failure: if that node goes down, your locking breaks with it. Redlock spreads the same pattern across multiple independent nodes so a single node failure doesn't silently drop the lock, at the cost of more operational complexity.
How long should a distributed lock's expiry be?
Long enough to comfortably cover the slowest realistic run of the work it protects, with margin, since a lock that lapses mid-job lets a second process start the same work. Too long, and a crashed worker holds the lock far longer than necessary, delaying the next legitimate run.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Spotting Vendor Lock-In Before It Costs You an Exit
A practical checklist for spotting vendor lock-in in your identity, API, and infrastructure stack before switching costs become the deciding factor.
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
Where Caching Helps a Zero Trust API and Where It Creates Risk
Comparing where caching genuinely speeds up a zero trust API against where it creates a real revocation and permission risk.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.
Rolling Out OpenTelemetry Without Drowning in Spans
A practical rollout plan for OpenTelemetry distributed tracing, including sampling strategy, span naming, and the mistakes that make traces unusable.