Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Choosing a Caching Strategy Without Creating a Consistency Nightmare

Caching is one of the few performance fixes that can make a system both faster and more confusing at the same time, because the moment cached data can diverge from the source of truth, every bug report starts with the question of whether you're looking at stale data or a real problem.

The right caching strategy depends less on which library you pick and more on how tolerant your data is of being briefly wrong, which is a product decision as much as a technical one.

When is cache-aside the right caching default?

The application checks the cache first, and on a miss, reads from the source of truth and populates the cache for next time. It's simple to reason about and works well for read-heavy data that changes infrequently. The gap is invalidation: when the underlying data changes, something has to actively remove or update the stale cache entry, and forgetting that step is the single most common source of confusing stale-data bugs.

Cache-aside is a reasonable default for most read-heavy endpoints, as long as invalidation is treated as a required part of every write path that touches cached data, not an optional cleanup step.

Write-through: consistency at the cost of write latency

Every write goes to the cache and the source of truth together, synchronously, so the cache is never stale after a successful write. This trades a bit of write latency, since the write isn't complete until both are updated, for much stronger consistency guarantees, which is the right trade for data where staleness would cause real problems: pricing, inventory counts, permissions.

Use write-through selectively, on the specific data where a stale read would actually cause harm, rather than as a default everywhere, since the latency cost adds up across a system with many write paths.

Write-behind: fast writes, with a real risk to weigh

Writes go to the cache immediately and are persisted to the source of truth asynchronously afterward. This is the fastest option for the write path, and it carries real risk: if the cache fails before the asynchronous write completes, that data is genuinely lost, not just temporarily stale. Reserve this pattern for data where that loss is tolerable, high-volume analytics events, non-critical activity logs, not for anything you'd need to reconstruct if it disappeared.

Most application teams reach for write-behind for the speed and underestimate the loss scenario until it actually happens once, which tends to be a memorable lesson.

Which layer needs a cache: CDN, application or database?

A CDN cache in front of largely static or slowly changing content solves a completely different problem than an application-level cache for computed results, which is different again from a database query cache. Identify which layer is actually the bottleneck before choosing a pattern, since an application-level cache added on top of an already-fast database query doesn't help and just adds another place for staleness to hide.

Measure first: if the database isn't actually slow, caching in front of it trades a small, well-understood latency cost for a new, harder-to-debug consistency problem, in exchange for a speed improvement nobody would have noticed.

A common mistake: caching before you've set a staleness budget

"How stale is acceptable" is a question worth answering explicitly, in seconds or minutes, before choosing a caching pattern, not an implicit assumption that gets discovered the first time someone notices outdated data and asks why nobody flagged it sooner. A five second staleness budget on a price and a five minute one on a rarely-changing profile field call for genuinely different invalidation strategies, and treating them the same either over-engineers the profile field or under-serves the price.

Write the staleness budget down per piece of cached data, the same way you'd write down a latency budget, so the caching decision has an actual target to be measured against instead of an implicit, unexamined one that only gets discovered by accident.

Revisit the budget whenever the product usage of that data changes. A field that was harmless to leave stale for five minutes when it only appeared on an internal dashboard can become a real problem once a new feature starts making pricing or availability decisions based on it, and nobody thinks to re-check the caching assumption when the new feature ships on top of the old cache.

Match the pattern to the data like this:

  • Use cache-aside for read heavy data that changes infrequently, with an explicit step to invalidate entries when the source changes.
  • Use write-through when staleness would cause real problems, such as pricing, inventory counts or permissions, and accept slightly slower writes.
  • Use write-behind only for disposable or reproducible data such as high volume analytics events, since a cache failure can lose unpersisted writes.
  • Set an explicit staleness budget for each type of data before choosing any pattern.
Executive Capability Standard

What Good Looks Like

Caching is working when every cached piece of data has an explicit staleness budget and an invalidation path that's actually exercised on every relevant write, not an assumption nobody's tested.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Measure actual latency at each layer, CDN, application, database, before assuming you know where a cache would help most.
2. Do Manually:Add cache-aside caching manually to your highest-traffic, infrequently-changing read paths, with explicit invalidation on the corresponding writes.
3. Delegate:Assign an engineer to own your caching strategy and document staleness budgets per data type as new caches get added.
4. Automate:Build event-driven invalidation so cached data updates automatically when the source of truth changes, instead of relying on manual cache-clearing code.
5. Buy:Bring in a CDN or caching platform specialist once your caching needs span multiple layers and manual invalidation logic is becoming hard to reason about.

How to Get Started

Frequently Asked Questions

Should we cache database queries or just cache at the application layer?

Start by measuring where the actual latency is coming from. If the database itself is fast and the slowness is in application-level computation, an application cache of the computed result helps more than a database query cache, and vice versa; caching the wrong layer adds complexity without fixing the real bottleneck.

How do we handle cache invalidation across multiple services?

Publish an event when the underlying data changes, and have every service that caches it subscribe and invalidate its own copy. That is more reliable than each service polling or guessing when to refresh, which tends to drift out of sync over time.

Is write-behind caching ever appropriate for critical data?

Generally no, because the risk of losing data that was only written to the cache before the source of truth update completes is usually not worth the latency savings for anything you'd need to reconstruct. Reserve it for genuinely disposable or reproducible data.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides