Redis, Postgres, or etcd: Choosing a Distributed Lock
The moment you run more than one instance of a service, some piece of work needs to happen exactly once even though several processes are racing to do it: a scheduled job, a migration, a webhook handler that must not double charge a customer. That's what a distributed lock is for, and there isn't one right answer for where to put it.
The right choice depends less on which lock is fastest and more on what you're already running, and how expensive it is if the lock is wrong exactly once.
What a distributed lock is actually protecting
Before picking a mechanism, be precise about the failure you're preventing. Some locks exist purely for efficiency, so two workers don't duplicate the same low-stakes work. Others exist because duplicate execution is a real incident, like charging a card twice or corrupting a shared file. The second category deserves a lock with real correctness guarantees, not just something that's usually right.
Write that distinction down for each lock in your system. It changes which of the options below is actually safe to use, and it's the first thing to check when a lock-related bug shows up months later.
For example, a lock around a nightly report email exists purely for efficiency: if two workers occasionally both send it, the cost is a duplicate email. A lock guarding a billing job is different, because a second execution can charge a customer twice. Label each lock in your code as efficiency or correctness, and give the correctness locks a mechanism with real guarantees, such as fencing tokens or a consensus store. Revisit the label whenever the lock's job changes, since a low-stakes lock can quietly pick up a higher-stakes responsibility over time.
Redis locks: fast, and dangerous if you skip fencing
A Redis lock, typically implemented with SETNX and an expiry, is simple and fast, which is why it's the default choice for low-stakes coordination. The danger is the expiry itself: if the process holding the lock stalls past the expiry, for example during a long garbage collection pause, another process can acquire the same lock while the first one is still running, unaware it lost ownership.
Fencing tokens, a monotonically increasing number checked by whatever the lock protects, close that gap by letting the protected resource reject a stale holder's write even after the lock itself has moved on. Skipping fencing is the single most common reason a Redis lock 'fails' in a way that's hard to reproduce.
Postgres advisory locks: free if you already trust Postgres
If your data already lives in Postgres, advisory locks give you locking with no new infrastructure and correctness guarantees you already trust for everything else. A session-level advisory lock is held until released or the connection drops, which ties its lifetime cleanly to something Postgres itself already tracks.
The tradeoff is throughput and portability. Advisory locks add load to a database that's usually your most constrained resource already, and they only work as long as everything that needs the lock can reach that same Postgres instance.
etcd and ZooKeeper: correctness first, latency second
For coordination that has to survive a network partition without silently granting the same lock to two holders, a purpose-built consensus store like etcd or ZooKeeper is the honest answer. They're built around the same algorithms that keep a distributed system agreeing on a single truth, at the cost of running and operating another stateful system.
For most small teams, that operational cost only pays off once the thing being protected is expensive enough, or dangerous enough, to justify it. Running etcd to protect a lock that gates a nightly report email is usually more infrastructure than the problem deserves.
Picking one for a small engineering team
If Postgres already holds your data, start with advisory locks and move only if the load it adds becomes a real problem. If you're already running Redis for caching, a lock with fencing tokens covers most everyday coordination cheaply. Reach for etcd or ZooKeeper only when the cost of a double-execution bug is high enough to justify running and maintaining a dedicated coordination service on top of everything else.
Whichever you pick, write a short note next to the lock in code explaining what it protects and what happens if it's ever acquired twice. That note is what saves the next engineer from assuming the lock is just a performance optimization when it's actually the only thing standing between the system and a duplicate charge.
Use this quick guide to match a lock to the stakes:
- If Postgres already holds your data, start with advisory locks and move only when the added database load becomes a real problem.
- If you already run Redis for caching, use a lock with fencing tokens for everyday coordination.
- Reach for etcd or ZooKeeper only when a double-execution bug is costly enough to justify operating another stateful system.
- Write a short note beside each lock in code saying what it protects and what happens if it is acquired twice.
What Good Looks Like
Good distributed locking means every lock in the system has a named failure mode you've thought through, not just a mechanism that's usually right.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is a Redis lock safe for something like a billing job?
Only with a fencing token that lets the protected resource reject a stale holder. Without one, a process that stalls past the lock's expiry can lose the lock without knowing it, and a second process can start the same job while the first is still finishing.
Do we need etcd or ZooKeeper if we're a small team?
Probably not yet. Both are built for correctness under network partitions, which is real value, but the operational cost of running one usually isn't worth it until the thing you're protecting is expensive enough to fail on.
What's the simplest safe option if we already use Postgres?
A Postgres advisory lock. It requires no new infrastructure, ties its lifetime to a connection Postgres already tracks, and is correct enough for most everyday coordination as long as the added load on your database is acceptable.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Keeping an Exit Ready When You Pick an Inference Vendor
How to pick an inference provider without losing your ability to leave: abstraction layers, data portability, and the exit costs worth checking up front.
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
Which Caching Strategy Actually Fits Your Inference Traffic
Comparing exact-match, semantic, and KV-cache reuse for AI model serving, and which one fits your actual traffic pattern.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
Rolling Out OpenTelemetry Without Drowning in Trace Data
How to instrument services with OpenTelemetry, choose a sampling strategy, and avoid the rollout mistake that leaves you with traces nobody reads.