Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

Choosing a Distributed Lock: Redis, Redlock, Postgres, or etcd

The right distributed lock is usually the lightest one that stops two instances of a job from running at once, often a single Redis lock or a Postgres advisory lock rather than a consensus system like etcd. Reach for etcd only when correctness truly demands it, since it adds an operational dependency most teams don't need.

This compares four approaches teams actually reach for, what each one protects against, and where each can still let two processes run at once despite holding what looks like a valid lock.

What You're Actually Protecting Against

Start by naming the failure you're avoiding, because it changes the right tool. Most locking in application code is about avoiding wasted work or duplicate side effects, like two workers sending the same email, not about a correctness guarantee that would cause data corruption if violated. For that class of problem, a lock that's almost always correct and cheap to run beats one that's provably correct and expensive to operate.

A smaller set of cases genuinely need a lock that can't fail open even under network partitions or clock drift, typically anything touching money movement or an irreversible external action. Be honest about which bucket your case falls into before reaching for the heavier options below; most locking problems belong in the first bucket.

A Single Redis Lock: Fast, and Fine for Most Cases

SET key value NX with an expiry is the simplest distributed lock: one Redis instance, one atomic command, a lock that self-expires if the holder crashes. It's fast and it's fine for the avoid-wasted-work bucket above. Its known failure mode is a process that holds the lock past its expiry, maybe from a slow garbage collection pause or a network stall, finishes its work, and releases a lock that a different process now holds, exposing someone else's work in progress to a second concurrent run.

Mitigate this with a lock value unique to the holder, checked on release so you only ever release your own lock, and a generous but bounded expiry with a renewal heartbeat for long-running work. For short jobs, this single-instance approach handles the overwhelming majority of real cases without needing anything more elaborate.

Redlock and the Debate Around It

Redlock extends the single-instance approach across several independent Redis nodes, requiring a majority to grant the lock, aiming to survive one node failing or being slow. It's a reasonable middle option, but it's worth knowing it's contested: the algorithm's guarantees under clock drift and process pauses have been publicly debated by distributed systems researchers, and the criticism is specifically that Redlock can still grant the same lock to two clients under conditions that aren't exotic in real production environments.

Use Redlock if you want meaningfully better failure tolerance than a single Redis instance without standing up a dedicated consensus system, and you've read enough of that debate to know it's not a silver bullet. Don't use it for the correctness-critical bucket from the first section; that's what the next two options are for.

Postgres Advisory Locks: No New Infrastructure

If you already run Postgres, advisory locks, pg_advisory_lock and its session or transaction-scoped variants, give you a distributed lock with no new infrastructure and no separate expiry mechanism to reason about, since the lock releases automatically when the holding session or transaction ends. That's genuinely simpler than tracking a Redis expiry.

The real limitation is connection pooling: if you're behind a pooler running in transaction mode, a session-scoped advisory lock won't behave the way you expect, since the session it was granted on doesn't reliably persist across the pooled connection. Use the transaction-scoped variant when you're behind a transaction-mode pooler, and test this specifically before relying on it, since the failure looks like the lock silently not working rather than an obvious error.

etcd or ZooKeeper for Real Consensus

Reach for a dedicated consensus system when the lock genuinely needs to survive network partitions with a correctness guarantee, not just usually work:

  • Leader election for a service that needs exactly one active instance at a time, where two active leaders would cause real damage, not just duplicate work.
  • Coordinating configuration changes across many nodes that must see a consistent view or not proceed at all.
  • Cases where the cost of a double-grant is significantly higher than the operational cost of running and monitoring a consensus cluster.
  • Cases where you already run etcd or ZooKeeper for another reason, since the marginal cost of using it for locking too is much lower than standing it up new.

For nearly everything else described earlier, this is more operational overhead than the problem calls for. Match the tool to the actual failure you're protecting against, not to which one sounds the most rigorous.

Executive Capability Standard

What Good Looks Like

Good distributed locking means the strategy matches the actual failure you're protecting against, wasted work versus a correctness-critical double-grant, rather than defaulting to whichever tool sounds most rigorous.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read through where your codebase currently uses locks and classify each one: is it preventing wasted work, or preventing genuine data corruption if it fails.
2. Do Manually:For a low-risk case, implement a single Redis lock with a unique holder value and a bounded expiry, and watch it in production before adding anything more elaborate.
3. Delegate:Have one engineer own the locking strategy for correctness-critical paths specifically, separate from the general-purpose locks used to avoid duplicate work.
4. Automate:Standardize a small internal locking library so every team reaches for the same pattern instead of each service inventing its own Redis lock logic.
5. Buy:Bring in outside distributed systems expertise before choosing a consensus system for a correctness-critical case if no one on the team has operated one before, since the operational failure modes are easy to get wrong.

How to Get Started

Frequently Asked Questions

Is a single Redis lock good enough for production?

For most cases, yes, as long as you're protecting against wasted work rather than a correctness-critical failure. Use a unique lock value so you only ever release your own lock, and a bounded expiry with a heartbeat for anything long-running. Reserve heavier options for cases where a double-grant would cause real harm.

Why is Redlock controversial?

Distributed systems researchers have publicly argued that Redlock's guarantees can break under clock drift and process pauses that aren't unusual in real production systems. It's still a reasonable middle option between a single Redis instance and a full consensus system, but treat it as improved reliability, not a provably correct lock.

Can we use Postgres advisory locks behind a connection pooler?

Only with the transaction-scoped variant if the pooler runs in transaction mode, since a session-scoped lock won't reliably persist across a pooled connection. Test this explicitly, because the failure mode is the lock silently not doing its job rather than a clear error.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides