Three Places to Cache in a RAG Pipeline, and What Each Buys You
A RAG pipeline has three distinct places worth caching: the query embedding, the set of retrieved chunks, and the final generated answer, and each solves a different problem. Adding a cache isn't a specific enough plan, and picking the wrong layer wastes effort on something that wasn't your bottleneck.
Why cache the embedding, not just the final answer?
The same query text, or near-identical text, embedded repeatedly wastes an API call and adds latency for no benefit. A simple exact-match cache keyed on the query string, storing the resulting embedding vector, is low-risk and cheap to build, since an identical input always deserves an identical embedding output with no risk of serving something wrong. This is usually the first and safest caching layer worth adding.
Cache retrieved chunk sets, separate from the final generated answer
Caching the set of chunks a query retrieved, rather than the final generated answer, lets you skip the search step for a repeated query while still regenerating a fresh response each time. This matters when the underlying model or prompt template changes often, since a cached final answer would need invalidating on every prompt change, while a cached chunk set stays valid as long as the underlying documents haven't changed, which is a much less frequent event.
Should you use a shared cache or a per-instance cache?
A cache kept in each application instance's own memory is fast, with no network hop, but starts cold after every deploy or scale-out event, and different instances can end up with different cached content until they've each independently warmed up. A shared distributed cache, backed by something like Redis, stays warm across deploys and is consistent across instances, at the cost of a network round trip on every lookup. For a pipeline that deploys often, the shared option usually wins despite the added latency, since a cold per-instance cache after every deploy defeats much of the point.
Invalidate on the write path, not on a timer alone
A cache that only expires on a fixed time-to-live serves stale results for however long that window is, even when a source document changed the moment the cache entry was written. When your ingestion pipeline updates a document, actively invalidate any cache entries derived from it as part of that same write, rather than waiting for a TTL to catch up. A TTL is still worth keeping as a backstop for cases the write-path invalidation misses, but it shouldn't be the only mechanism.
For example, a support knowledge base updates its refund policy document at nine in the morning. With a one-hour time-to-live as the only invalidation, customers keep receiving chunks from the old policy until the entries expire, even though the new text was ingested immediately. If the ingestion job instead deletes every cache entry derived from that document as part of the same write, the next query retrieves the new policy right away. The time-to-live still stays in place as a backstop for anything the write-path invalidation misses.
Watch cache hit rate as its own metric, not an assumption
A caching layer that looked like a good idea on paper sometimes barely gets used in practice, if query traffic turns out to be more unique than expected. Track hit rate explicitly for each cache layer, and be willing to remove one that isn't earning its complexity. A cache with a low hit rate adds code, adds a failure mode of its own, an invalidation bug serving stale data, and delivers little of the cost or latency benefit it was built for.
Set an explicit size limit and eviction policy
An unbounded cache eventually consumes whatever memory is available to it, crowding out other processes on a shared instance or driving up the bill on a managed cache service. Set a deliberate size limit and an eviction policy, typically least-recently-used, so the cache holds what's actually being reused rather than accumulating every query ever seen. Size the limit against real hit rate data, since a cache too small to hold your actual working set delivers little benefit no matter how well-designed the rest of it is.
Keep cache failures from becoming pipeline failures
A caching layer should make the pipeline faster and cheaper, not less reliable. If the cache service itself becomes unavailable, the pipeline should fall through to computing the result directly rather than failing the whole request. Treat every cache read and write as something that can fail independently of the underlying retrieval and generation logic, and write the calling code so a cache outage degrades performance rather than breaking functionality outright.
Before adding a cache layer, check that it meets these criteria:
- It has a clear job, such as skipping an embedding call, skipping a search, or reusing an answer, and that job is your actual bottleneck.
- Its keys include tenant or permission context wherever results depend on who is asking.
- Entries are invalidated on the write path when a source document changes, with a time-to-live kept only as a backstop.
- It has an explicit size limit and an eviction policy, typically least-recently-used.
- Hit rate is tracked as its own metric, and a layer that isn't earning its complexity gets removed.
- A cache outage falls through to computing the result directly instead of failing the request.
What Good Looks Like
The caching standard is a shared, distributed cache for embeddings and chunk sets with write-path invalidation, distinct from the final generated answer, and hit rate tracked per layer to confirm each one is earning its complexity.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Which cache layer should we build first if we're starting from nothing?
The embedding cache, since it's the lowest-risk and cheapest to build: an exact text match always deserves the same embedding, so there's no risk of serving a wrong result. It also directly reduces embedding API cost, which is often the most immediately visible line item on the bill.
How do we know if our cache hit rate is good enough to justify the complexity?
There's no universal threshold, but compare the infrastructure and maintenance cost of the cache against the API cost and latency it's actually saving, measured from real hit rate data rather than an assumption made when it was first built. A cache saving a small fraction of calls at meaningful maintenance cost may not be worth keeping.
Is it safe to cache results that include tenant-specific or permission-filtered data?
Only if the cache key includes the tenant or permission context, not just the query text, since two users with different access asking the same question should never share a cached result derived from documents one of them can't see. Treat this as a hard requirement, not an edge case to handle later.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Where RAG Latency Actually Goes, and How to Budget It
Break a RAG request into its four latency stages, find out which one is actually slow, and set a budget for each before you start tuning blindly.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.
Choosing a Distributed Locking Pattern Without Overbuilding It
A decision guide for choosing a distributed locking approach, from a simple database row lock to a dedicated coordination service, based on what you need.
Setting Up Distributed Tracing Without Drowning in Spans
A practical guide to rolling out OpenTelemetry distributed tracing: what to instrument first, and how to keep trace data useful instead of overwhelming.
How Vector Search Throughput Degrades as Your Index Grows
Throughput doesn't fall off gradually as a vector index grows. Here's why it degrades in steps, and how sharding, replicas, and quantization each help.