Redis Locks, Postgres Advisory Locks, or etcd: Picking a Locking Pattern
The problem is always the same shape: two instances of the same job, running on two different machines, and only one of them is supposed to actually do the work. Maybe it's a nightly billing run, a cache warm-up, or a cron job that scaled out along with your app servers. Without a lock, you get duplicate charges, duplicate emails, or a race condition that only shows up under load.
The tools for this are well established, but which one is right depends less on which is technically best and more on what you already run.
If You Already Run Postgres, Start With Advisory Locks
Postgres has a built-in advisory lock feature that costs you nothing extra to run: no new service, no new failure mode, just a function call against a database you're already depending on. You acquire a lock on an arbitrary integer key, do your work, and release it, or let it release automatically when the session ends. The tradeoff is that the lock is only as reliable as your database connection: if the process holding the lock loses its connection ungracefully, Postgres does clean up the lock, but you need to be careful that your application's connection pooling doesn't interfere with lock ownership.
If You Already Run Redis, a Simple Lock Beats Redlock for Most Cases
A basic Redis lock, using SET with NX and an expiry, is simple and fast, and for most job-deduplication use cases, that simplicity is exactly what you want. The expiry is what saves you if a process crashes mid-job: the lock releases itself instead of staying held forever. Redlock, the multi-node Redis locking algorithm, exists to guard against a narrower failure mode, a single Redis node failing over at exactly the wrong moment, and it adds real complexity to handle that edge case. Unless you've had an actual incident traceable to that specific failure, a single-node lock with a sane expiry and a bit of jitter to avoid thundering-herd retries covers the vast majority of real workloads.
When You Need etcd or Zookeeper Instead
A dedicated coordination service earns its keep when locking isn't a side feature of your architecture but a core one: leader election across a cluster, service discovery, configuration that must be consistent across every node at once. These systems are built around consensus protocols specifically to guarantee correctness under network partitions, which is more guarantee than most job-deduplication use cases actually need. If you're only trying to stop a cron job from double-running, running etcd for that alone is a lot of operational overhead for a problem Postgres or Redis already solves.
The Failure Mode That Breaks All Three
Every locking approach shares the same weak point: a lock that outlives the process holding it. If a worker acquires a lock, then gets stuck, killed, or partitioned from the network before it finishes, the lock can end up held far longer than the work should have taken. Always set an expiry on the lock, not just on the connection, and make sure the expiry is shorter than your alerting threshold for a stuck job. A lock with no expiry is a single point of failure waiting to happen the first time a process doesn't exit cleanly.
A Simple Decision Rule
If your job-deduplication need fits in a sentence, one job, one lock, a few seconds to a few minutes of expected runtime, start with whichever of Postgres or Redis you already operate. Reach for a dedicated coordination service only when locking is genuinely central to your architecture, not incidental to one background job. Most teams that reach straight for a heavier tool are solving a problem they don't actually have yet.
The short version, by what you already run:
- Postgres advisory locks fit when you already run Postgres and the need is one job, one lock, and a short runtime, with no new service to operate.
- A simple Redis lock using SET with NX and an expiry fits when you already run Redis, and the expiry frees the lock if a process crashes.
- etcd or Zookeeper fit when locking is central to the architecture, such as leader election, service discovery, or configuration that must be consistent everywhere.
- Whichever you pick, always set an expiry so a stuck or partitioned worker cannot hold the lock far longer than the work should take.
Testing a Lock Before You Trust It
A lock that's never been tested under a real crash is a lock you're only guessing works. Before relying on any locking approach in production, deliberately kill the process holding the lock mid-job, hard, not a graceful shutdown, and confirm two things: that the lock actually releases within its expiry window, and that a second process picking up the work doesn't duplicate whatever the first process already completed. That second part is easy to overlook: a lock stops two processes from running at the same time, but it does nothing to make a half-finished job safe to resume from the start. If a killed job can leave partial side effects, like a payment charged but not recorded, your job logic needs to be idempotent on top of being locked, not instead of it.
What Good Looks Like
A solid locking setup means every job that must not run twice has an explicit lock with an expiry shorter than your alert threshold, using whichever system, Postgres, Redis, or a coordination service, matches what the job actually needs.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What happens if two processes try to acquire the same Redis lock at the exact same millisecond?
Redis processes commands one at a time, so the SET NX operation is atomic. One process gets the lock, the other gets told the key already exists and moves on. There's no window where both processes believe they hold the lock simultaneously, as long as you're using the atomic set-if-not-exists form.
Do we need to worry about clock drift between servers when setting a lock expiry?
A little, but less than you'd think for most job-deduplication cases. Both Redis and Postgres track expiry on their own server clock, not the client's, so drift between your application servers doesn't affect it. Build in a reasonable buffer between the lock's expiry and your expected job duration to absorb normal variance.
Is a database-based lock too slow for high-frequency locking?
For a background job that runs once a minute or less, no, the overhead is negligible. If you're locking on every request in a hot path, hundreds or thousands of times a second, the overhead does start to matter, and that's a stronger signal to look at Redis or a purpose-built coordination service instead of a database lock.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
The Portability Audit: What It Costs to Leave a Vendor
A checklist for finding out what it would really take to leave a cloud vendor or platform, before you're forced to find out during a price increase.
Distributed Locks With Redis: What Actually Fails
Why a simple Redis lock isn't mutual exclusion, what a fencing token fixes and doesn't, and a safer default for most small engineering teams.
Redis Lock, Postgres Advisory Lock, or Zookeeper: Picking One
A comparison of the three common ways to coordinate distributed locks: Redis-based locks, Postgres advisory locks, and a dedicated coordination service.
Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)
A walkthrough of a real double-charge bug caused by a Redis lock's TTL expiring mid-job, and the fencing-token pattern that actually fixes it.
Choosing a Distributed Locking Pattern Without Overbuilding It
A decision guide for choosing a distributed locking approach, from a simple database row lock to a dedicated coordination service, based on what you need.
When a Redis Lock Is Enough, and When It Isn't
Single-instance locks, Redlock, fencing tokens, and when to skip Redis entirely for a database advisory lock instead. A decision guide for CTOs.