Cache Invalidation Is Still the Hard Part
Adding a cache is the easy part of distributed caching. Every hard problem in this space is really about invalidation: knowing when cached data is wrong, and clearing it before a user sees something stale or a downstream service acts on outdated information.
This is a practical way to pick a caching strategy and, more importantly, a way to keep it honest.
Match the caching pattern to how the data actually changes
Cache-aside, where the application checks the cache first and populates it on a miss, works well for read-heavy data that changes infrequently. Write-through, updating the cache at the same time as the source of truth, fits data where staleness is unacceptable even briefly.
Picking the wrong pattern for the data's actual change frequency is the root of most caching bugs. A write-heavy resource cached with a naive cache-aside pattern and no active invalidation will serve stale data constantly, no matter how well the cache itself performs.
Set a TTL that reflects real staleness tolerance, not a round number
A default five-minute TTL applied everywhere because it seemed reasonable is rarely actually reasonable for every piece of data in your system. A user's session data might tolerate a few seconds of staleness; a product catalog might tolerate an hour; an account balance might tolerate none at all without active invalidation.
Set TTLs per data type based on how long stale data is actually acceptable for that specific use case, not a single value copied across every cached resource because it's easier to configure once.
Invalidate actively for anything staleness genuinely can't tolerate
TTL-based expiry alone means data can be wrong for up to the full TTL window after it changes, which is fine for some data and unacceptable for other data in the same system. For anything in the second category, invalidate the cache entry explicitly, as part of the same operation that changes the underlying data, rather than waiting for it to expire naturally.
This needs to be reliable: a write path that sometimes forgets to invalidate the cache, because someone added a new way to update the data without updating the invalidation logic, is a recurring source of stale-data bugs that are hard to reproduce after the fact.
A worked example: a price shown wrong after a change
Say a product's price gets updated in the source database, and the cache entry for that product's page, with a fifteen-minute TTL, keeps serving the old price for up to fifteen minutes after the change. A customer checks out during that window at the stale price, and now there's a pricing discrepancy to resolve after the fact.
Adding explicit cache invalidation to the price-update code path, clearing the specific cache key the moment the price changes, removes the staleness window entirely for this specific case, while leaving the TTL as a safety net for cases the explicit invalidation might miss.
Where caching strategies break down
- One TTL value applied uniformly across data with very different staleness tolerance
- Write paths that update the source of truth without invalidating the corresponding cache entry
- No monitoring on cache hit rate, so a cache that's stopped being effective goes unnoticed
- Caching layered on top of a service without anyone owning invalidation correctness for new features
Monitor hit rate and staleness incidents, not just latency
A cache that's working shows up in latency graphs as an improvement, but that alone doesn't tell you whether it's serving correct data. Track hit rate to catch a cache that's stopped being effective, and track stale-data incidents specifically, tagged as such, to catch invalidation bugs before they become a pattern.
A cache with a great hit rate and no visibility into staleness incidents can still be quietly serving wrong data most of the time. The two metrics answer different questions, and you need both to trust the cache is actually helping.
Be deliberate about caching negative results too
It's easy to only think about caching successful lookups and forget that a repeated failed lookup, a not-found or a permission check that returns no, hits your backend just as hard as a successful one would. Caching negative results, with a shorter TTL than positive ones, protects against a specific pattern: a client repeatedly checking for something that doesn't exist yet.
This matters most for lookups triggered by client-side polling or retries, where a missing resource being checked for repeatedly can generate meaningful load if every check bypasses the cache. A short negative-result TTL absorbs that pattern without risking long-lived incorrect data if the resource does show up shortly after.
What Good Looks Like
A well-run caching layer matches its pattern and TTL to each data type's actual staleness tolerance, invalidates explicitly wherever staleness genuinely matters, and is monitored for both hit rate and stale-data incidents.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should every read path in a distributed system have a cache in front of it?
No, only where read volume and data staleness tolerance both justify the added complexity. A low-traffic endpoint with strict correctness requirements often isn't worth caching at all, since the invalidation risk outweighs the performance benefit.
What's a safe default TTL if we're not sure how stale data can be?
Start short, seconds to a couple of minutes, and lengthen it deliberately once you've confirmed how stale the data can actually get before it causes a real problem. It's far easier to extend a TTL that's too conservative than to discover one was too long after a stale-data incident.
How do we catch a write path that forgot to invalidate the cache?
Add a test that writes data through every known update path and asserts the corresponding cache entry is cleared. This catches the gap in CI rather than in production, where it usually surfaces as a confusing, hard-to-reproduce bug report.
Is a distributed cache like Redis always better than an in-process cache?
Not always. An in-process cache is faster and simpler for data that's fine to be slightly inconsistent across instances, while a shared distributed cache is worth the added latency and complexity when every instance needs to see the same value immediately after a change.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.
Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)
A walkthrough of a real double-charge bug caused by a Redis lock's TTL expiring mid-job, and the fencing-token pattern that actually fixes it.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
A Worksheet for Deciding What to Instrument With OpenTelemetry First
A simple worksheet for prioritizing which services get OpenTelemetry instrumentation first, based on incident history and blast radius, not ease of setup.
How to Run a Real Security Audit on a Distributed System
A working method for auditing service boundaries, credentials, and patch timelines across a distributed system instead of filling out a compliance checklist.