Choosing a Distributed Locking Pattern Without Overbuilding It
Distributed locking gets reached for whenever two processes might touch the same resource at once, but the right implementation depends entirely on how much correctness you actually need and how bad a false lock failure would be. Reaching for a heavyweight coordination service to prevent two cron jobs from double-firing is usually overkill. Reaching for a flimsy one to protect a financial transaction is a real risk.
Use the criteria below to match the lock to the job instead of defaulting to whatever pattern the last team you worked with happened to use.
Criterion one: what happens if the lock briefly fails
If two workers occasionally both run the same idempotent cleanup task, the cost of a lock failing is wasted compute, not corrupted data. If two workers both decrement the same inventory count, the cost is a real data integrity bug. Start by writing down the actual consequence of the lock not holding, since that answer determines whether you need a lock that's merely helpful or one that's provably correct under failure.
A common mistake is choosing a lock by what the team already runs rather than by what a failure costs. A cache is already deployed, so a TTL lock gets used for a job that moves money, and nobody writes down what happens if two workers both believe they hold it. A better habit is to record the failure consequence next to the lock, in the code or a short design note, so a later reviewer can see whether the chosen pattern matches the stakes and challenge it if the job has become more important since it was written.
When is a database row lock enough for distributed locking?
If your resource already lives in a relational database, a row-level lock acquired with a SELECT FOR UPDATE inside a transaction is often the simplest correct option, since it uses infrastructure you already operate and already understand. It fits well for coordinating access to a specific record, like preventing two requests from both processing the same order. It doesn't fit well when the resource isn't naturally row-shaped, or when lock hold times would be long enough to create contention on a hot table.
When is a TTL-based cache lock good enough?
A lock implemented as a key with a time-to-live in a cache like Redis is fast and works well for short-lived, best-effort coordination, such as making sure only one instance of a scheduled job runs at a time. Its weakness is the TTL itself: if the process holding the lock is slow or crashes, another process can acquire the lock before the slow one finishes, and correctness depends entirely on getting the TTL and renewal logic right. Treat this pattern as good for reducing duplicate work, not as a guarantee against it.
Option: a dedicated coordination service
Tools built specifically for distributed consensus give you strong guarantees, including fencing tokens that prevent a lock holder that's fallen behind from taking action after it's lost the lock. This is the right tool when the cost of two processes both believing they hold the lock is genuinely unacceptable, such as coordinating a leader election for a system that must never have two active leaders. It's also the heaviest option operationally, since it's another piece of infrastructure to run, monitor, and understand during an incident.
A decision shortcut
If a false failure just means duplicate work that's safe to redo, use the simplest lock your existing infrastructure supports. If a false failure means data corruption or a genuinely unsafe dual-write, invest in a pattern with fencing tokens and treat the added operational cost as the price of that correctness. Most production incidents involving locks trace back to a team picking the lightweight option for a problem that needed the strong guarantee, not the other way around.
Match the lock to the job with these rules of thumb:
- If a false failure only means duplicate work that is safe to redo, use the simplest lock your existing infrastructure supports.
- If the resource is already a database record, a row lock inside a transaction is often the simplest correct option.
- For short-lived, best-effort coordination such as one scheduled job at a time, a TTL lock in a cache layer usually fits.
- If a false failure means data corruption or an unsafe dual write, use a coordination service with fencing tokens and accept the operational cost.
A worked example: the double-fired payout job
Say a scheduled job pays out vendor invoices once a day, protected by a TTL lock in a cache layer with no fencing token. If that process pauses for a garbage collection cycle long enough for the TTL to expire, a second instance can acquire the lock and start paying out invoices while the first instance is still mid-run, unaware it's lost its lock. The result is duplicate payments, not a crash, which makes it far harder to notice until finance reconciles the ledger. A fencing token would have let the payout system reject the first instance's writes the moment its lock expired.
What Good Looks Like
Good distributed locking matches the strength of the lock to the actual cost of it failing, using the simplest option that provides that guarantee rather than defaulting to the heaviest available tool.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is a TTL-based lock ever safe for critical operations?
It can be, but only when paired with a fencing token: a monotonically increasing value the lock holder must present, so a downstream system can reject an action from a holder whose lock has already expired. Without that check, a TTL lock alone isn't a strong enough guarantee for anything where a duplicate action causes real harm.
Do we need distributed locking if we're running a single instance?
Usually not for that instance alone, but check whether anything else, like a deployment rolling out a second instance briefly, could still create two processes running at once. A lot of single-instance assumptions quietly break the first time autoscaling or a rolling deploy overlaps two copies of the same service.
How do we test that our locking actually prevents the race condition?
Write a test that deliberately starts two workers at nearly the same time against the same resource and asserts only one of them completes the protected action. A lock that looks correct in code review can still have a race condition that only shows up under genuine concurrency, so the test needs to create that concurrency directly.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
The Vendor Lock-In Checklist for Your Vector Search Stack
A practical checklist for keeping your RAG and vector search stack portable, from embedding format to index rebuild cost, before you're stuck with one vendor.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Three Places to Cache in a RAG Pipeline, and What Each Buys You
Embedding caches, chunk-set caches, and shared versus per-instance caching each solve a different RAG cost or latency problem. Here's how to pick.
Distributed Locks With Redis: What Actually Fails
Why a simple Redis lock isn't mutual exclusion, what a fencing token fixes and doesn't, and a safer default for most small engineering teams.
When a Circuit Breaker Helps, and When It Just Hides a Bug
A decision guide for using circuit breakers and bulkheads to stop one failing dependency from cascading, and where a circuit breaker can mask a real problem.
Setting Up Distributed Tracing Without Drowning in Spans
A practical guide to rolling out OpenTelemetry distributed tracing: what to instrument first, and how to keep trace data useful instead of overwhelming.