Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Choosing a Distributed Locking Pattern Without Overbuilding It

Distributed locking gets reached for whenever two processes might touch the same resource at once, but the right implementation depends entirely on how much correctness you actually need and how bad a false lock failure would be. Reaching for a heavyweight coordination service to prevent two cron jobs from double-firing is usually overkill. Reaching for a flimsy one to protect a financial transaction is a real risk.

Use the criteria below to match the lock to the job instead of defaulting to whatever pattern the last team you worked with happened to use.

Criterion one: what happens if the lock briefly fails

If two workers occasionally both run the same idempotent cleanup task, the cost of a lock failing is wasted compute, not corrupted data. If two workers both decrement the same inventory count, the cost is a real data integrity bug. Start by writing down the actual consequence of the lock not holding, since that answer determines whether you need a lock that's merely helpful or one that's provably correct under failure.

A common mistake is choosing a lock by what the team already runs rather than by what a failure costs. A cache is already deployed, so a TTL lock gets used for a job that moves money, and nobody writes down what happens if two workers both believe they hold it. A better habit is to record the failure consequence next to the lock, in the code or a short design note, so a later reviewer can see whether the chosen pattern matches the stakes and challenge it if the job has become more important since it was written.

When is a database row lock enough for distributed locking?

If your resource already lives in a relational database, a row-level lock acquired with a SELECT FOR UPDATE inside a transaction is often the simplest correct option, since it uses infrastructure you already operate and already understand. It fits well for coordinating access to a specific record, like preventing two requests from both processing the same order. It doesn't fit well when the resource isn't naturally row-shaped, or when lock hold times would be long enough to create contention on a hot table.

When is a TTL-based cache lock good enough?

A lock implemented as a key with a time-to-live in a cache like Redis is fast and works well for short-lived, best-effort coordination, such as making sure only one instance of a scheduled job runs at a time. Its weakness is the TTL itself: if the process holding the lock is slow or crashes, another process can acquire the lock before the slow one finishes, and correctness depends entirely on getting the TTL and renewal logic right. Treat this pattern as good for reducing duplicate work, not as a guarantee against it.

Option: a dedicated coordination service

Tools built specifically for distributed consensus give you strong guarantees, including fencing tokens that prevent a lock holder that's fallen behind from taking action after it's lost the lock. This is the right tool when the cost of two processes both believing they hold the lock is genuinely unacceptable, such as coordinating a leader election for a system that must never have two active leaders. It's also the heaviest option operationally, since it's another piece of infrastructure to run, monitor, and understand during an incident.

A decision shortcut

If a false failure just means duplicate work that's safe to redo, use the simplest lock your existing infrastructure supports. If a false failure means data corruption or a genuinely unsafe dual-write, invest in a pattern with fencing tokens and treat the added operational cost as the price of that correctness. Most production incidents involving locks trace back to a team picking the lightweight option for a problem that needed the strong guarantee, not the other way around.

Match the lock to the job with these rules of thumb:

  • If a false failure only means duplicate work that is safe to redo, use the simplest lock your existing infrastructure supports.
  • If the resource is already a database record, a row lock inside a transaction is often the simplest correct option.
  • For short-lived, best-effort coordination such as one scheduled job at a time, a TTL lock in a cache layer usually fits.
  • If a false failure means data corruption or an unsafe dual write, use a coordination service with fencing tokens and accept the operational cost.

A worked example: the double-fired payout job

Say a scheduled job pays out vendor invoices once a day, protected by a TTL lock in a cache layer with no fencing token. If that process pauses for a garbage collection cycle long enough for the TTL to expire, a second instance can acquire the lock and start paying out invoices while the first instance is still mid-run, unaware it's lost its lock. The result is duplicate payments, not a crash, which makes it far harder to notice until finance reconciles the ledger. A fencing token would have let the payout system reject the first instance's writes the moment its lock expired.

Executive Capability Standard

What Good Looks Like

Good distributed locking matches the strength of the lock to the actual cost of it failing, using the simplest option that provides that guarantee rather than defaulting to the heaviest available tool.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand the failure modes of whichever locking pattern your team currently uses, especially what happens if the lock holder crashes mid-operation.
2. Do Manually:Walk through your highest-risk shared resource and manually trace what happens if two processes both believe they hold the lock at once.
3. Delegate:Give one engineer ownership of the shared locking utility so every service uses the same tested implementation instead of reinventing it.
4. Automate:Add monitoring that alerts when lock acquisition failures or contention spike, since that's often an early signal of a scaling problem.
5. Buy:Adopt a managed coordination service for your highest-stakes locks rather than running your own consensus infrastructure.

How to Get Started

Frequently Asked Questions

Is a TTL-based lock ever safe for critical operations?

It can be, but only when paired with a fencing token: a monotonically increasing value the lock holder must present, so a downstream system can reject an action from a holder whose lock has already expired. Without that check, a TTL lock alone isn't a strong enough guarantee for anything where a duplicate action causes real harm.

Do we need distributed locking if we're running a single instance?

Usually not for that instance alone, but check whether anything else, like a deployment rolling out a second instance briefly, could still create two processes running at once. A lot of single-instance assumptions quietly break the first time autoscaling or a rolling deploy overlaps two copies of the same service.

How do we test that our locking actually prevents the race condition?

Write a test that deliberately starts two workers at nearly the same time against the same resource and asserts only one of them completes the protected action. A lock that looks correct in code review can still have a race condition that only shows up under genuine concurrency, so the test needs to create that concurrency directly.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides