Picking a Distributed Lock That Won't Let Two Jobs Silently Run at Once
Distributed locks look simple in a design doc and get complicated the moment a node pauses for garbage collection, a network partition splits your cluster, or a lock holder crashes mid task without releasing anything. The job of a distributed lock isn't just to coordinate the happy path; it's to fail safely when one of those things happens, and different locking patterns fail very differently.
For a pipeline where two workers processing the same partition at once means duplicated writes or a corrupted aggregate, the choice of locking pattern is worth getting right the first time rather than discovering the gap during an incident.
Is a single Redis lock enough for a distributed lock?
The simplest pattern sets a key in Redis with an expiration, and whichever worker sets it first holds the lock until it releases or the key expires. This works fine for low stakes coordination, but it has a real gap: if the lock holder pauses long enough, for a garbage collection stop or a slow network call, the key can expire while that worker still believes it holds the lock and keeps working.
That gap means a single Redis lock alone isn't safe for anything where a second worker acting concurrently causes real damage, like a financial calculation or a write that isn't idempotent. It's a good fit for lower stakes coordination, like preventing two identical cron triggers from double firing when the cost of an occasional double fire is low.
Redlock and its actual guarantees
Redlock extends the single instance pattern across multiple independent Redis nodes and requires a majority to agree before granting a lock, which is meant to tolerate a single node failing. It's a meaningful improvement over one instance, but it doesn't fully solve the pause problem described above; a worker that stalls after acquiring the lock can still have its lock expire while it keeps operating, since Redlock's safety guarantee is about consensus across nodes at acquisition time, not about what happens to a slow holder afterward.
If you use Redlock, pair it with a fencing token: a monotonically increasing number issued with the lock that downstream systems check before accepting a write, so a stale lock holder's write gets rejected even if it never learned its lock had expired.
Postgres advisory locks for pipelines already anchored to Postgres
If your pipeline already writes to Postgres, an advisory lock avoids adding a new coordination dependency entirely. Session level advisory locks release automatically if the holding connection dies, which sidesteps a common failure mode where a crashed worker leaves a lock held forever. The tradeoff is that advisory locks are scoped to a single Postgres instance, so they don't help if your workers are coordinating across separate databases.
They're a strong default for a moderate scale pipeline that's already Postgres centric, since the failure handling comes largely for free from how Postgres manages connection lifecycles, without the added operational surface of running a separate coordination service.
When should you use ZooKeeper or etcd for locking?
Systems built specifically for distributed consensus, like ZooKeeper or etcd, use ephemeral nodes tied to a session and a heartbeat, so a crashed or partitioned holder loses its lock automatically without relying on a fixed expiration guess. This handles the pause and partition problems more rigorously than Redis based approaches, at the cost of running and operating another stateful service.
This is the right tool when the cost of two workers processing the same partition is genuinely severe, like a payments ledger or an inventory system where a duplicate write corrupts a balance that's expensive to reconcile after the fact. For most data pipeline coordination, it's more infrastructure than the problem calls for.
Matching the pattern to what actually breaks in your pipeline
Before picking a pattern, write down what actually happens if two workers process the same partition at the same moment: does it produce a duplicate row your downstream deduplication already handles, or does it corrupt a running total that nobody notices until a customer complains. Those are very different risk profiles, and the honest answer usually points to a much cheaper solution than the one that feels safest on paper.
For a pipeline processing an event stream with idempotent writes already built in further downstream, a single Redis lock with a fencing token is often genuinely enough, and the extra operational cost of etcd buys safety the system doesn't actually need. Save the heavier tooling for the specific partitions where a duplicate write is expensive, rather than applying the strongest pattern everywhere by default.
Ask these questions before choosing a locking pattern:
- What happens if two workers process the same partition at once: a duplicate row that downstream deduplication absorbs, or a corrupted running total nobody notices?
- Can a paused worker, stalled by a garbage collection stop or a slow network call, keep writing after its lock has already expired?
- Does the pipeline already write to Postgres, so a session level advisory lock can release automatically when a crashed connection dies?
- Do you need coordination across multiple databases, or is a coordination failure severe enough to justify running ZooKeeper or etcd?
- Would a fencing token, checked by the system the lock protects, reject writes from a stale holder that never realized its lock expired?
What Good Looks Like
A distributed locking setup that holds under real conditions accounts for what happens when a lock holder pauses or crashes, not just how the lock is acquired on the happy path.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is a single Redis lock ever safe enough for production use?
Yes, for coordination where an occasional double execution is a minor inconvenience rather than a real problem, like deduplicating a cron trigger. It's not safe for anything where a stale lock holder writing data causes real damage, since a paused worker can keep operating after its key expires.
What's a fencing token and why does it matter for distributed locks?
It's a number that increases every time a lock is granted, passed along with the lock to whatever system the lock protects. That system rejects any write carrying an older token than the last one it accepted, which catches a stale lock holder even if the holder itself never realized its lock had expired.
Should we default to Postgres advisory locks if our pipeline already uses Postgres?
It's a reasonable default for moderate scale coordination within a single database, since the lock releases automatically if the holding connection dies. Move to a dedicated system like etcd only if you need to coordinate across multiple databases or the cost of a coordination failure is severe enough to justify the extra operational overhead.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Checklist for Keeping Your Event Pipeline Portable
A practical checklist for keeping a real-time event pipeline portable, so switching a managed provider stays a project instead of a rebuild.
Blue-Green, Canary, or Rolling: Deploying Stream Processors
A decision guide to rolling, blue-green, and canary deploys for stateful stream processors, plus the rollback plan most teams never actually test.
Cache-Aside, Write-Through, or Write-Behind for Streamed Data
A decision guide to cache-aside, write-through, and write-behind caching for data enriched by a stream, and how to invalidate a cache off real events.
Verifying Every Service That Talks to Your Pipeline
Which parts of zero-trust verification to build and which to buy, so every producer and consumer on a streaming pipeline proves its identity.
Distributed Locks With Redis: What Actually Fails
Why a simple Redis lock isn't mutual exclusion, what a fencing token fixes and doesn't, and a safer default for most small engineering teams.
Rolling Out OpenTelemetry Without Drowning Your Team in Spans
A practical rollout sequence for OpenTelemetry distributed tracing across a real time pipeline, including where to instrument first and how to control cost.