Choosing a Caching Strategy Without Creating a Bigger Problem
Caching is easy to add and easy to get subtly wrong in a way that doesn't show up until a customer sees stale or incorrect data and asks why. The lookup-speed part of caching is the easy part. The part that actually causes problems is knowing when to throw the cached value away.
Here's how to choose a strategy without creating a harder problem than the one you started with.
Cache-aside versus write-through: what each one actually trades off
Cache-aside checks the cache first, falls back to the real data source on a miss, and stores the result for next time. It's simple and works well when reads vastly outnumber writes, but leaves a window where the cache can be stale after the underlying data changes, until something invalidates or expires it.
Write-through updates the cache at the same time as the underlying data, keeping them in sync, at the cost of added latency and complexity on every write. It's a better fit when staleness is genuinely unacceptable for that particular data, and a worse fit for data that changes often but isn't read often enough to justify the extra write overhead.
The real hard part: invalidation, not lookup speed
Almost any caching approach makes reads fast. The genuinely hard problem is knowing exactly when a cached value is no longer accurate and needs to be thrown away, especially when the same underlying data can be updated from more than one place in your system. Missing one of those update paths is how a cache quietly starts serving wrong data.
Before choosing a caching strategy, map every place the underlying data can change. If you can't name them all confidently, that's a sign the invalidation logic isn't ready yet, regardless of which caching pattern you pick.
Where a local cache is enough and where you need a shared one
A local, in-process cache is fast and simple, and it's a reasonable choice for data that's expensive to compute but tolerant of being slightly different across instances for a short window. A shared cache, like a dedicated caching layer all your instances read from, is the better choice when consistency across instances actually matters, or when the data is large enough that duplicating it in every instance's memory is wasteful.
Default to local for anything where a brief, minor inconsistency between instances genuinely doesn't matter, since it's the simpler system to reason about and operate.
TTLs are a guess, treat them like one
A time-to-live setting is really a bet about how long a piece of data stays accurate enough to serve, and teams often set it once, based on a rough feeling, and never revisit it. A TTL that's too long serves stale data longer than it should. One that's too short defeats much of the point of caching by forcing frequent, unnecessary refetches.
Set TTLs based on how often the underlying data actually changes in practice, and revisit them if the data's real update frequency shifts, rather than treating the original guess as permanent.
A worked example: a cache that served stale prices for hours
Say pricing data gets updated from two different internal tools, a pricing dashboard and a bulk import job. If only the dashboard's update path invalidates the cache, prices changed through the bulk import keep serving from cache until the TTL eventually expires, potentially for hours, with no error or warning anywhere in the system. Mapping every update path in advance, as described above, is exactly what would have caught this gap before it reached a real customer.
Deciding whether you have a caching problem at all
Not every slow read needs a cache. If the underlying query or call is simply inefficient, fixing that directly is usually a better first step than caching around it, since a cache adds real complexity (invalidation, staleness, another thing that can be wrong) that a genuinely fast, well-built data path doesn't need at all. Reach for caching once the underlying operation is already reasonably efficient and still too slow or too expensive to run on every request.
Before you add a cache, work through these checks:
- Confirm the slow read is not simply an inefficient query or call, since fixing that directly is usually a better first step than caching it.
- Map every path that can change the underlying data, including internal tools and bulk imports, before you ship the cache.
- Write a test that updates the data through each of those paths and confirms the cache reflects the change.
- Set the TTL from how often that specific data really changes, treating it as a bet and not a single default for everything.
- Decide whether a local in-process cache is enough, or whether the data needs a shared cache that every instance sees consistently.
What Good Looks Like
Good here means every cached value has a clear, tested invalidation path covering every place the underlying data can change, not just the one path someone happened to think of first.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How do we decide what TTL to use for a specific piece of data?
Base it on how often that specific data actually changes in practice, not a single default applied everywhere. Data that updates rarely can tolerate a longer TTL; anything that changes frequently, or where staleness has real consequences, needs a much shorter one or an explicit invalidation path instead.
Is it better to cache at the database level or the application level?
It depends on what you're optimizing for. Application-level caching is more flexible and easier to reason about for specific, expensive computations. Database-level caching can help broadly across many queries without custom logic, but gives you less control over exactly when a specific value gets invalidated.
How do we catch a caching bug before it reaches customers?
Map every path that can change the underlying data before you ship the cache, and write a test that updates the data through each of those paths and confirms the cache reflects the change. Most caching bugs come from an update path nobody accounted for, and this is the step that surfaces it early.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Choosing a Caching Strategy Without Creating a Consistency Nightmare
A comparison of cache-aside, write-through, and write-behind caching, and how to decide which layer, CDN, application, or database, actually needs one.
How to Ship a Risky Change Without a 2am Rollback
A concrete walkthrough of how to plan a risky production deployment: how to split it, what to watch, and when to decide the rollback trigger.
Picking a Caching Approach Without Creating a Consistency Mess
A comparison of common distributed caching approaches, with the consistency and invalidation tradeoffs each one actually carries in production.
Diagnosing Slow Requests Before You Blame the Database
A step-by-step way to find out whether a slowdown is the network, the app, or the database, before you add caching or upgrade infrastructure to fix it.
Build or Buy for Verifying Every Device That Connects?
How to split device identity from device posture checking, what building either one in house actually costs, and where a platform earns its keep instead.
Choosing a Caching Strategy Without Overbuilding It
A decision guide for picking a caching approach that matches your actual read patterns, instead of defaulting to the most complex option available.