Data Engineering & Real-Time Event StreamsPlaybook3 min readUpdated September 2026

Picking a Distributed Lock That Won't Let Two Jobs Silently Run at Once

Distributed locks look simple in a design doc and get complicated the moment a node pauses for garbage collection, a network partition splits your cluster, or a lock holder crashes mid task without releasing anything. The job of a distributed lock isn't just to coordinate the happy path; it's to fail safely when one of those things happens, and different locking patterns fail very differently.

For a pipeline where two workers processing the same partition at once means duplicated writes or a corrupted aggregate, the choice of locking pattern is worth getting right the first time rather than discovering the gap during an incident.

Is a single Redis lock enough for a distributed lock?

The simplest pattern sets a key in Redis with an expiration, and whichever worker sets it first holds the lock until it releases or the key expires. This works fine for low stakes coordination, but it has a real gap: if the lock holder pauses long enough, for a garbage collection stop or a slow network call, the key can expire while that worker still believes it holds the lock and keeps working.

That gap means a single Redis lock alone isn't safe for anything where a second worker acting concurrently causes real damage, like a financial calculation or a write that isn't idempotent. It's a good fit for lower stakes coordination, like preventing two identical cron triggers from double firing when the cost of an occasional double fire is low.

Redlock and its actual guarantees

Redlock extends the single instance pattern across multiple independent Redis nodes and requires a majority to agree before granting a lock, which is meant to tolerate a single node failing. It's a meaningful improvement over one instance, but it doesn't fully solve the pause problem described above; a worker that stalls after acquiring the lock can still have its lock expire while it keeps operating, since Redlock's safety guarantee is about consensus across nodes at acquisition time, not about what happens to a slow holder afterward.

If you use Redlock, pair it with a fencing token: a monotonically increasing number issued with the lock that downstream systems check before accepting a write, so a stale lock holder's write gets rejected even if it never learned its lock had expired.

Postgres advisory locks for pipelines already anchored to Postgres

If your pipeline already writes to Postgres, an advisory lock avoids adding a new coordination dependency entirely. Session level advisory locks release automatically if the holding connection dies, which sidesteps a common failure mode where a crashed worker leaves a lock held forever. The tradeoff is that advisory locks are scoped to a single Postgres instance, so they don't help if your workers are coordinating across separate databases.

They're a strong default for a moderate scale pipeline that's already Postgres centric, since the failure handling comes largely for free from how Postgres manages connection lifecycles, without the added operational surface of running a separate coordination service.

When should you use ZooKeeper or etcd for locking?

Systems built specifically for distributed consensus, like ZooKeeper or etcd, use ephemeral nodes tied to a session and a heartbeat, so a crashed or partitioned holder loses its lock automatically without relying on a fixed expiration guess. This handles the pause and partition problems more rigorously than Redis based approaches, at the cost of running and operating another stateful service.

This is the right tool when the cost of two workers processing the same partition is genuinely severe, like a payments ledger or an inventory system where a duplicate write corrupts a balance that's expensive to reconcile after the fact. For most data pipeline coordination, it's more infrastructure than the problem calls for.

Matching the pattern to what actually breaks in your pipeline

Before picking a pattern, write down what actually happens if two workers process the same partition at the same moment: does it produce a duplicate row your downstream deduplication already handles, or does it corrupt a running total that nobody notices until a customer complains. Those are very different risk profiles, and the honest answer usually points to a much cheaper solution than the one that feels safest on paper.

For a pipeline processing an event stream with idempotent writes already built in further downstream, a single Redis lock with a fencing token is often genuinely enough, and the extra operational cost of etcd buys safety the system doesn't actually need. Save the heavier tooling for the specific partitions where a duplicate write is expensive, rather than applying the strongest pattern everywhere by default.

Ask these questions before choosing a locking pattern:

  • What happens if two workers process the same partition at once: a duplicate row that downstream deduplication absorbs, or a corrupted running total nobody notices?
  • Can a paused worker, stalled by a garbage collection stop or a slow network call, keep writing after its lock has already expired?
  • Does the pipeline already write to Postgres, so a session level advisory lock can release automatically when a crashed connection dies?
  • Do you need coordination across multiple databases, or is a coordination failure severe enough to justify running ZooKeeper or etcd?
  • Would a fencing token, checked by the system the lock protects, reject writes from a stale holder that never realized its lock expired?
Executive Capability Standard

What Good Looks Like

A distributed locking setup that holds under real conditions accounts for what happens when a lock holder pauses or crashes, not just how the lock is acquired on the happy path.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map every place in your pipeline where two workers processing the same data concurrently would cause a real problem, and what locking pattern each currently uses, if any.
2. Do Manually:Add a fencing token to your highest risk lock first, so a stale holder's write gets rejected downstream even if the lock itself already expired.
3. Delegate:Assign a senior engineer to own your coordination strategy across the pipeline and document which pattern is used where and why.
4. Automate:Build a shared locking library with fencing tokens built in by default, so new pipeline code doesn't reinvent unsafe locking from scratch.
5. Buy:Bring in a fractional CTO or distributed systems specialist if a duplicate processing incident has already caused a data integrity problem.

How to Get Started

Frequently Asked Questions

Is a single Redis lock ever safe enough for production use?

Yes, for coordination where an occasional double execution is a minor inconvenience rather than a real problem, like deduplicating a cron trigger. It's not safe for anything where a stale lock holder writing data causes real damage, since a paused worker can keep operating after its key expires.

What's a fencing token and why does it matter for distributed locks?

It's a number that increases every time a lock is granted, passed along with the lock to whatever system the lock protects. That system rejects any write carrying an older token than the last one it accepted, which catches a stale lock holder even if the holder itself never realized its lock had expired.

Should we default to Postgres advisory locks if our pipeline already uses Postgres?

It's a reasonable default for moderate scale coordination within a single database, since the lock releases automatically if the holding connection dies. Move to a dedicated system like etcd only if you need to coordinate across multiple databases or the cost of a coordination failure is severe enough to justify the extra operational overhead.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides