Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Cache Invalidation Is Still the Hard Part

Adding a cache is the easy part of distributed caching. Every hard problem in this space is really about invalidation: knowing when cached data is wrong, and clearing it before a user sees something stale or a downstream service acts on outdated information.

This is a practical way to pick a caching strategy and, more importantly, a way to keep it honest.

Match the caching pattern to how the data actually changes

Cache-aside, where the application checks the cache first and populates it on a miss, works well for read-heavy data that changes infrequently. Write-through, updating the cache at the same time as the source of truth, fits data where staleness is unacceptable even briefly.

Picking the wrong pattern for the data's actual change frequency is the root of most caching bugs. A write-heavy resource cached with a naive cache-aside pattern and no active invalidation will serve stale data constantly, no matter how well the cache itself performs.

Set a TTL that reflects real staleness tolerance, not a round number

A default five-minute TTL applied everywhere because it seemed reasonable is rarely actually reasonable for every piece of data in your system. A user's session data might tolerate a few seconds of staleness; a product catalog might tolerate an hour; an account balance might tolerate none at all without active invalidation.

Set TTLs per data type based on how long stale data is actually acceptable for that specific use case, not a single value copied across every cached resource because it's easier to configure once.

Invalidate actively for anything staleness genuinely can't tolerate

TTL-based expiry alone means data can be wrong for up to the full TTL window after it changes, which is fine for some data and unacceptable for other data in the same system. For anything in the second category, invalidate the cache entry explicitly, as part of the same operation that changes the underlying data, rather than waiting for it to expire naturally.

This needs to be reliable: a write path that sometimes forgets to invalidate the cache, because someone added a new way to update the data without updating the invalidation logic, is a recurring source of stale-data bugs that are hard to reproduce after the fact.

A worked example: a price shown wrong after a change

Say a product's price gets updated in the source database, and the cache entry for that product's page, with a fifteen-minute TTL, keeps serving the old price for up to fifteen minutes after the change. A customer checks out during that window at the stale price, and now there's a pricing discrepancy to resolve after the fact.

Adding explicit cache invalidation to the price-update code path, clearing the specific cache key the moment the price changes, removes the staleness window entirely for this specific case, while leaving the TTL as a safety net for cases the explicit invalidation might miss.

Where caching strategies break down

  • One TTL value applied uniformly across data with very different staleness tolerance
  • Write paths that update the source of truth without invalidating the corresponding cache entry
  • No monitoring on cache hit rate, so a cache that's stopped being effective goes unnoticed
  • Caching layered on top of a service without anyone owning invalidation correctness for new features

Monitor hit rate and staleness incidents, not just latency

A cache that's working shows up in latency graphs as an improvement, but that alone doesn't tell you whether it's serving correct data. Track hit rate to catch a cache that's stopped being effective, and track stale-data incidents specifically, tagged as such, to catch invalidation bugs before they become a pattern.

A cache with a great hit rate and no visibility into staleness incidents can still be quietly serving wrong data most of the time. The two metrics answer different questions, and you need both to trust the cache is actually helping.

Be deliberate about caching negative results too

It's easy to only think about caching successful lookups and forget that a repeated failed lookup, a not-found or a permission check that returns no, hits your backend just as hard as a successful one would. Caching negative results, with a shorter TTL than positive ones, protects against a specific pattern: a client repeatedly checking for something that doesn't exist yet.

This matters most for lookups triggered by client-side polling or retries, where a missing resource being checked for repeatedly can generate meaningful load if every check bypasses the cache. A short negative-result TTL absorbs that pattern without risking long-lived incorrect data if the resource does show up shortly after.

Executive Capability Standard

What Good Looks Like

A well-run caching layer matches its pattern and TTL to each data type's actual staleness tolerance, invalidates explicitly wherever staleness genuinely matters, and is monitored for both hit rate and stale-data incidents.

Building The Capability (5-Stage Skill Ladder)

1. Learn:List your cached data types and, for each one, write down honestly how stale it can be before it causes a real problem.
2. Do Manually:Adjust TTLs by hand to match that staleness tolerance instead of leaving a single default value applied everywhere.
3. Delegate:Give the team that owns each data type responsibility for its cache invalidation logic, rather than treating caching as separate infrastructure.
4. Automate:Add tests that verify every write path invalidates its corresponding cache entry, catching gaps in CI instead of production.
5. Buy:Bring in a performance engineering specialist once your caching layer spans enough services that invalidation correctness is hard to reason about by hand.

How to Get Started

Frequently Asked Questions

Should every read path in a distributed system have a cache in front of it?

No, only where read volume and data staleness tolerance both justify the added complexity. A low-traffic endpoint with strict correctness requirements often isn't worth caching at all, since the invalidation risk outweighs the performance benefit.

What's a safe default TTL if we're not sure how stale data can be?

Start short, seconds to a couple of minutes, and lengthen it deliberately once you've confirmed how stale the data can actually get before it causes a real problem. It's far easier to extend a TTL that's too conservative than to discover one was too long after a stale-data incident.

How do we catch a write path that forgot to invalidate the cache?

Add a test that writes data through every known update path and asserts the corresponding cache entry is cleared. This catches the gap in CI rather than in production, where it usually surfaces as a confusing, hard-to-reproduce bug report.

Is a distributed cache like Redis always better than an in-process cache?

Not always. An in-process cache is faster and simpler for data that's fine to be slightly inconsistent across instances, while a shared distributed cache is worth the added latency and complexity when every instance needs to see the same value immediately after a change.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides