Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Where RAG Latency Actually Goes, and How to Budget It

"The RAG pipeline is slow" isn't a diagnosis, it's a symptom with four possible causes. Before you tune anything, find out which stage of the request is actually eating the time.

Most teams guess it's the vector search. Often it isn't.

How do you break a RAG request into its four stages?

A typical RAG request has four legs: embedding the query, searching the vector index, an optional reranking pass, and the model generating a response from the retrieved context. Instrument each one separately with timing spans, not just an overall request timer. On many stacks the embedding call and the generation call, both round trips to a hosted model, dwarf the vector search itself, which can return in single-digit milliseconds even against millions of vectors.

Once you have per-stage numbers, you'll know whether you're solving a search problem or a network problem.

Distributed tracing makes this much easier than manually timestamping code. A single trace that shows all four spans nested under one request lets you see not just how long each stage took, but whether they ran sequentially when they could have run in parallel; embedding the query and checking a cache, for instance, don't need to block each other.

Tune the index instead of guessing

If the search stage genuinely is slow, the fix is usually a parameter, not a rewrite. Approximate nearest-neighbor indexes like HNSW trade recall for speed through a search-time parameter (often called ef_search or similar), and IVF-style indexes do the same through the number of clusters probed. Turning either one down speeds up search and lowers recall; turning it up does the opposite. Measure recall against a fixed query set before and after any change, since a faster index that returns worse matches just moves the problem downstream to the model.

Test this on a copy of your production corpus, not a small synthetic one, since recall behavior at ten thousand vectors doesn't reliably predict behavior at ten million. A parameter setting that looks safe in a small test set can quietly drop recall once the index reaches production scale, where near-duplicate and borderline-similar content becomes more common.

Decide whether you need a reranker at all

A cross-encoder reranking pass over the top candidates from vector search improves result quality, but it adds a full model call to the request path, often more latency than the vector search step it's refining. It's worth it when your first-stage retrieval returns a lot of near-miss results and quality visibly suffers without it. It's not worth it if your embedding model and chunking already produce clean top results; you'd be paying latency for a quality gain nobody notices.

A middle option worth testing before committing either way: rerank only when the score gap between the first and fifth vector search result is small, since a large gap usually means the top result is already clearly best and reranking won't change the answer, just the latency.

Cache what actually repeats

Look at your query logs before building a cache. If the same or near-duplicate questions recur, common on a support or documentation site, a semantic cache keyed on query similarity can skip the search and generation stages entirely for a meaningful share of traffic. If queries are mostly unique, as they tend to be on an internal analytics tool, a query cache won't help much and you're better off spending the effort on the embedding and generation legs instead.

A semantic cache needs its own tuning too: too loose a similarity threshold and it serves a wrong answer for a genuinely different question; too tight and it barely fires. Start conservative and loosen the threshold only after checking a sample of what it's matching.

How should you set a latency budget for a RAG pipeline?

A p50 latency number hides the requests that actually frustrate users. Set a budget per stage, embedding, search, rerank, generation, and alert on p95 or p99 for each independently. A search stage that's usually 8ms but spikes to 400ms under load points at a resourcing problem on the vector database, while a generation stage with a long, flat tail usually points at a model provider issue you can't fix locally, only route around.

Turn the per-stage numbers into budgets with these rules:

  • Set a separate latency budget for embedding, search, reranking, and generation instead of relying on one overall request timer.
  • Alert on p95 or p99 for each stage independently, because a p50 number hides the requests that actually frustrate users.
  • Read the pattern: a search stage that spikes under load points to database resourcing, while a long flat generation tail points to the model provider.
  • Track cold starts separately from steady state, since a scaled-to-zero embedding or reranking function can leave p50 normal and p99 terrible.

Watch for the cold-start tax separately from steady-state latency

A serverless embedding or reranking function that scales to zero looks fast in a load test, which tends to keep it warm, and then adds a full cold-start penalty to the first request after any idle period in production. This shows up as a confusing pattern in your metrics: a normal p50 and a terrible p99, with no obvious cause in the application code. Check whether any stage in your pipeline runs on infrastructure that can scale to zero, and either keep a minimum number of instances warm or account for the cold-start tax explicitly in your latency budget instead of chasing it as a mystery.

Executive Capability Standard

What Good Looks Like

The latency standard is per-stage timing with independent budgets for embedding, search, reranking, and generation, tuned against a measured recall target rather than a single end-to-end number.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Add timing spans around each of the four request stages so you can see where time actually goes before changing anything.
2. Do Manually:Run a fixed query set through the pipeline by hand at a few different index parameter settings and record the recall and latency tradeoff.
3. Delegate:Have a backend engineer own the index tuning parameters and revisit them whenever the corpus size changes materially.
4. Automate:Build a recall-and-latency benchmark into CI so a reindex or a parameter change that regresses either one gets caught before deploy.
5. Buy:Bring in a search or infrastructure specialist once you're operating at a scale where index tuning has real cost and latency stakes.

How to Get Started

Frequently Asked Questions

Is vector search usually the slow part of a RAG request?

Less often than people assume. Approximate nearest-neighbor search is typically the fastest leg of the request. The embedding call and the final generation call, both round trips to a model, are the more common bottlenecks. Measure before you optimize the search index.

Will increasing recall always hurt latency?

Usually, yes, since higher recall means examining more candidates before returning results. The relationship isn't always linear, though: a poorly tuned index can be both slower and less accurate than a well-tuned one at the same recall target, so retuning sometimes improves both at once.

Should we cache the final generated answer or the retrieved chunks?

Cache the retrieved chunks first. They're stable across near-duplicate queries and reusing them still lets you regenerate a fresh answer. Caching the full generated answer is faster but riskier, since it can serve a stale response for a question whose underlying data has since changed.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides