Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)
A billing job takes a lock on a customer ID before charging their card, sets a thirty-second TTL, and releases the lock when it's done. Most nights that works fine. One night, a slow response from the payment processor makes the job run for forty-five seconds. The lock expires at thirty. A second worker picks up the same customer, takes the now-available lock, and charges them again.
The fix isn't a longer TTL. It's understanding why any fixed TTL was the wrong tool for this job in the first place.
What actually happened
The lock's TTL was a guess at how long the critical section would take, and the critical section included a network call to a payment processor whose latency the job doesn't control. Most nights the guess held. The one night it didn't, the lock silently expired while the first worker was still mid-charge, and Redis had no way to know the job wasn't actually finished, because a TTL only tracks time, not whether the work inside it is done.
Why 'just extend the TTL' isn't the real fix
Raising the TTL to sixty seconds buys you some margin until the next slow response from the payment processor pushes past sixty too. The underlying problem is that the critical section's duration depends on an external system you don't control, so any fixed number is a bet, not a guarantee. A longer TTL also has its own cost: if a worker crashes mid-job instead of running long, that customer's lock now stays held for the full sixty seconds before anyone else can process them, which is a worse outcome for a common failure mode in exchange for protection against a rarer one.
The two fixes that actually address it
The first is to get the slow external call out of the locked section entirely, using a reserve-then-confirm pattern: reserve the charge (write an intent record) inside the lock, release the lock, make the external call, then confirm or roll back the intent based on the result. The lock only needs to be held for the fast, local part of the work. The second is a lock with automatic lease extension: a background renewal (a watchdog) keeps extending the TTL as long as the holding worker is still alive and working, so the lock only expires early if the worker itself has actually died, not just because a call ran long.
Either approach should also carry a fencing token: a monotonically increasing number issued with the lock, checked by the downstream system (here, the charge itself) before it acts. If a worker's lock expired and a second worker took over, the first worker's fencing token is now stale, and the charge system rejects it even if that first worker wakes up (from a pause, a slow GC cycle, or a network partition) still believing it holds the lock.
Where a Postgres advisory lock is the simpler tool
For lower-stakes coordination that's already happening inside a database transaction, a Postgres advisory lock avoids the whole class of problem: it's tied to the transaction or session lifetime rather than a separately managed TTL, and there's no extra network hop to a lock service that can itself become a point of failure. The tradeoff is scope: it only coordinates work happening against that one Postgres instance, so it doesn't help when the thing you're locking spans multiple services or data stores, which is exactly the case the billing job above was in.
The mistake underneath the mistake: no fencing token
Even a well-tuned TTL and a careful watchdog don't fully close this gap without a fencing token, because a paused worker (a long garbage collection pause, a brief network partition) can wake up still believing it holds a lock that has since been reassigned. The fencing token is what lets the downstream system say no to a stale writer, regardless of what the worker itself believes. Skipping it is the most common reason a team's second attempt at distributed locking still has the same bug as the first, just harder to reproduce, because a shorter TTL or a watchdog narrows the window without closing it.
Testing this without waiting for a slow night in production
You can reproduce the original bug on purpose: add an artificial delay to the payment call in a test environment so it reliably exceeds the lock's TTL, then run two workers against the same customer ID and confirm only one charge goes through. That same test, run again after adding lease renewal and a fencing token, is the actual proof the fix works, rather than an absence of errors in production being read as success. A fix that isn't exercised this way tends to get 'confirmed' by a quiet week, which is exactly what the original bug also looked like right up until the payment processor had a slow night.
To reproduce the double-charge bug on purpose, follow these steps:
- In a test environment, add an artificial delay to the payment call so it reliably runs longer than the lock's TTL.
- Start two workers against the same customer ID, so the second can take the lock after it expires.
- Confirm whether one charge or two goes through, and treat a second charge as a failed test.
- Run the same test again after adding lease renewal and a fencing token, and confirm only one charge succeeds.
What Good Looks Like
A production distributed lock keeps any slow external call outside the locked section, uses lease renewal instead of a single fixed TTL where that's not possible, and pairs with a fencing token checked by whatever the lock is protecting.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is Redlock the right pattern for this kind of lock?
Redlock addresses coordinating a lock across multiple Redis nodes, which is a different problem from the TTL-versus-critical-section mismatch here. You can use Redlock and still hit this exact bug if the critical section's duration isn't bounded and there's no fencing token checked by the system being protected.
How long should a lock's TTL be?
Long enough to cover the fast, local part of the work with margin, after you've moved any slow external call outside the locked section. If you can't move the external call out, use lease renewal instead of a single fixed TTL, since no fixed number reliably covers a duration you don't control.
Do we need a fencing token if we already have lease renewal?
Yes. Lease renewal reduces how often a lock expires early, but it doesn't eliminate a paused worker waking up after its lock was reassigned. The fencing token is what protects the downstream system in that specific case, and it's cheap to add compared with debugging a duplicate charge in production.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Reducing Vendor Lock-In Without Slowing Your Team Down
How to tell real vendor lock-in from ordinary switching costs, where it actually bites, and why a multi-cloud abstraction often costs more than it saves.
Cache Invalidation Is Still the Hard Part
A practical guide to choosing a caching layer and, more importantly, keeping it from serving stale or wrong data across a distributed system.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
A Worksheet for Deciding What to Instrument With OpenTelemetry First
A simple worksheet for prioritizing which services get OpenTelemetry instrumentation first, based on incident history and blast radius, not ease of setup.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.