API Security, Identity & Zero-TrustPlaybook3 min readUpdated September 2026

Distributed Locking With Redis: Where Redlock Actually Falls Short

A distributed lock is how you stop two processes from doing the same job at once when they're running on different machines: two workers picking up the same queued job, two schedulers both deciding it's time to run the same nightly close.

Redis is the default choice because it's already in most stacks and it's fast, but its locking guarantees are weaker than most teams assume, and knowing exactly where they're weak matters more than knowing the pattern exists.

The basic Redis lock, and why a plain set-if-not-exists isn't enough

A distributed lock with Redis starts with a conditional set that only succeeds if the key doesn't already exist, plus an expiry. The expiry matters because a worker that crashes while holding the lock must not hold it forever; it's what releases the lock automatically.

The gap in the basic version is that any process can delete any lock, since a plain delete doesn't check who set it. Fix that by setting a random value on acquire and only releasing with a script that checks the value matches before deleting, which stops a worker from releasing a lock it doesn't actually hold anymore.

Where Redlock's guarantees actually break down

Redlock extends the pattern across multiple independent Redis nodes so a single node failure doesn't silently drop your lock. Its known weak point is timing: if a process's clock pauses, a long garbage collection pause, a virtual machine migration, for longer than the lock's expiry, the lock can lapse while the process still believes it holds it, and a second process can acquire the same lock.

For most background job deduplication, that risk is acceptable because the cost of a rare double-run is low. For anything where a double-run causes real damage, like a financial settlement job, that risk usually isn't acceptable.

When to reach for a database lock instead

If you're already writing to Postgres as part of the job, a row lock or an advisory lock gives you locking with the same transactional guarantees as the rest of your write, without introducing Redis as a second source of truth for correctness. It's slower than Redis and ties the lock's lifetime to a database connection staying open, but for anything where correctness matters more than raw speed, that tradeoff is usually the right one.

Reach for Redis when the lock is protecting a side effect outside the database, like preventing two workers from both calling an external API for the same job.

A checklist before you ship a new distributed lock

  • Does the lock have an expiry, and is it long enough to cover the slowest realistic run of the work it protects?
  • Does release check ownership before deleting, so a stale worker can't release a lock it no longer holds?
  • What happens if the lock is never released, either because a crash skipped cleanup or the expiry was set too short: does the job just not run, or does it silently double-run?
  • Is the cost of a rare double-run genuinely low, or does this specific job need the stronger guarantees of a database-backed lock instead of Redis?

A worked example: two workers picking up the same job

Say a queue delivers the same job message twice, a normal occurrence under an at-least-once delivery guarantee, and two worker processes both pick it up within milliseconds of each other. Without a lock, both workers do the same expensive work, and if the job has a side effect like charging a card or sending an email, the customer sees it happen twice.

With a Redis lock keyed on the job's own ID, the second worker's acquire attempt fails immediately because the key already exists, and it can safely skip the job rather than run it. The lock's expiry has to be set with the job's realistic worst-case runtime in mind, since a job that runs longer than the lock's expiry can have a second worker pick it up midway through, right when the lock lapses.

Locks that outlive their usefulness

A lock acquired for a job that later gets refactored into something idempotent on its own no longer needs the lock at all, but teams rarely go back and remove it, so the locking overhead and a second point of failure stick around for no benefit. Periodically review which locks in the codebase are actually still protecting something that isn't already safe to run twice, and remove the ones that aren't, since every lock is one more thing that can silently fail closed and block a job that should have run.

Executive Capability Standard

What Good Looks Like

A good distributed locking setup uses an expiry long enough to survive the slowest realistic job run, checks ownership before releasing, and reserves Redis-based locks for work where a rare double-run is genuinely low-cost.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Audit every place in your codebase that currently uses a lock, distributed or otherwise, and check whether each one has an expiry and an ownership check on release.
2. Do Manually:Add a script-based release that checks the lock's value before deleting it to any lock currently using a plain delete.
3. Delegate:Assign a senior engineer to own the locking pattern used across services, so new locks don't get implemented inconsistently by whoever needs one next.
4. Automate:Standardize on a shared locking library or wrapper internally so every team gets the ownership check and expiry handling by default instead of reimplementing it.
5. Buy:A distributed systems specialist or fractional CTO is worth bringing in when a lock is protecting something with real financial or safety consequences and you need a second opinion on the design.

How to Get Started

Frequently Asked Questions

Is Redis locking safe for financial or billing operations?

Generally not on its own. Redlock's timing guarantees can fail if a process's clock pauses longer than the lock's expiry, letting two processes briefly believe they both hold the lock. For anything where a double-run causes real damage, a database-backed lock with the same transactional guarantees as the rest of the write is usually the safer choice.

What's the difference between a plain Redis lock and Redlock?

A plain lock against a single Redis node is simple but has a single point of failure: if that node goes down, your locking breaks with it. Redlock spreads the same pattern across multiple independent nodes so a single node failure doesn't silently drop the lock, at the cost of more operational complexity.

How long should a distributed lock's expiry be?

Long enough to comfortably cover the slowest realistic run of the work it protects, with margin, since a lock that lapses mid-job lets a second process start the same work. Too long, and a crashed worker holds the lock far longer than necessary, delaying the next legitimate run.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides