Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)

A billing job takes a lock on a customer ID before charging their card, sets a thirty-second TTL, and releases the lock when it's done. Most nights that works fine. One night, a slow response from the payment processor makes the job run for forty-five seconds. The lock expires at thirty. A second worker picks up the same customer, takes the now-available lock, and charges them again.

The fix isn't a longer TTL. It's understanding why any fixed TTL was the wrong tool for this job in the first place.

What actually happened

The lock's TTL was a guess at how long the critical section would take, and the critical section included a network call to a payment processor whose latency the job doesn't control. Most nights the guess held. The one night it didn't, the lock silently expired while the first worker was still mid-charge, and Redis had no way to know the job wasn't actually finished, because a TTL only tracks time, not whether the work inside it is done.

Why 'just extend the TTL' isn't the real fix

Raising the TTL to sixty seconds buys you some margin until the next slow response from the payment processor pushes past sixty too. The underlying problem is that the critical section's duration depends on an external system you don't control, so any fixed number is a bet, not a guarantee. A longer TTL also has its own cost: if a worker crashes mid-job instead of running long, that customer's lock now stays held for the full sixty seconds before anyone else can process them, which is a worse outcome for a common failure mode in exchange for protection against a rarer one.

The two fixes that actually address it

The first is to get the slow external call out of the locked section entirely, using a reserve-then-confirm pattern: reserve the charge (write an intent record) inside the lock, release the lock, make the external call, then confirm or roll back the intent based on the result. The lock only needs to be held for the fast, local part of the work. The second is a lock with automatic lease extension: a background renewal (a watchdog) keeps extending the TTL as long as the holding worker is still alive and working, so the lock only expires early if the worker itself has actually died, not just because a call ran long.

Either approach should also carry a fencing token: a monotonically increasing number issued with the lock, checked by the downstream system (here, the charge itself) before it acts. If a worker's lock expired and a second worker took over, the first worker's fencing token is now stale, and the charge system rejects it even if that first worker wakes up (from a pause, a slow GC cycle, or a network partition) still believing it holds the lock.

Where a Postgres advisory lock is the simpler tool

For lower-stakes coordination that's already happening inside a database transaction, a Postgres advisory lock avoids the whole class of problem: it's tied to the transaction or session lifetime rather than a separately managed TTL, and there's no extra network hop to a lock service that can itself become a point of failure. The tradeoff is scope: it only coordinates work happening against that one Postgres instance, so it doesn't help when the thing you're locking spans multiple services or data stores, which is exactly the case the billing job above was in.

The mistake underneath the mistake: no fencing token

Even a well-tuned TTL and a careful watchdog don't fully close this gap without a fencing token, because a paused worker (a long garbage collection pause, a brief network partition) can wake up still believing it holds a lock that has since been reassigned. The fencing token is what lets the downstream system say no to a stale writer, regardless of what the worker itself believes. Skipping it is the most common reason a team's second attempt at distributed locking still has the same bug as the first, just harder to reproduce, because a shorter TTL or a watchdog narrows the window without closing it.

Testing this without waiting for a slow night in production

You can reproduce the original bug on purpose: add an artificial delay to the payment call in a test environment so it reliably exceeds the lock's TTL, then run two workers against the same customer ID and confirm only one charge goes through. That same test, run again after adding lease renewal and a fencing token, is the actual proof the fix works, rather than an absence of errors in production being read as success. A fix that isn't exercised this way tends to get 'confirmed' by a quiet week, which is exactly what the original bug also looked like right up until the payment processor had a slow night.

To reproduce the double-charge bug on purpose, follow these steps:

  1. In a test environment, add an artificial delay to the payment call so it reliably runs longer than the lock's TTL.
  2. Start two workers against the same customer ID, so the second can take the lock after it expires.
  3. Confirm whether one charge or two goes through, and treat a second charge as a failed test.
  4. Run the same test again after adding lease renewal and a fencing token, and confirm only one charge succeeds.
Executive Capability Standard

What Good Looks Like

A production distributed lock keeps any slow external call outside the locked section, uses lease renewal instead of a single fixed TTL where that's not possible, and pairs with a fencing token checked by whatever the lock is protecting.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map every place in the codebase where a lock's critical section includes a call to an external service you don't control the latency of.
2. Do Manually:Rework the highest-risk of those locks to use a reserve-then-confirm pattern that moves the external call outside the locked section.
3. Delegate:Have a senior engineer add fencing tokens to the remaining locked operations and get them checked by the systems they protect.
4. Automate:Add lease renewal (a watchdog that extends the TTL while the worker is alive) for any lock that still can't move its slow call out.
5. Buy:Bring in backend or distributed-systems advisory if a duplicate-execution bug like this has already reached production and the fix needs to happen fast.

How to Get Started

Frequently Asked Questions

Is Redlock the right pattern for this kind of lock?

Redlock addresses coordinating a lock across multiple Redis nodes, which is a different problem from the TTL-versus-critical-section mismatch here. You can use Redlock and still hit this exact bug if the critical section's duration isn't bounded and there's no fencing token checked by the system being protected.

How long should a lock's TTL be?

Long enough to cover the fast, local part of the work with margin, after you've moved any slow external call outside the locked section. If you can't move the external call out, use lease renewal instead of a single fixed TTL, since no fixed number reliably covers a duration you don't control.

Do we need a fencing token if we already have lease renewal?

Yes. Lease renewal reduces how often a lock expires early, but it doesn't eliminate a paused worker waking up after its lock was reassigned. The fencing token is what protects the downstream system in that specific case, and it's cheap to add compared with debugging a duplicate charge in production.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides