Where RAG Latency Actually Goes, and How to Budget It
"The RAG pipeline is slow" isn't a diagnosis, it's a symptom with four possible causes. Before you tune anything, find out which stage of the request is actually eating the time.
Most teams guess it's the vector search. Often it isn't.
How do you break a RAG request into its four stages?
A typical RAG request has four legs: embedding the query, searching the vector index, an optional reranking pass, and the model generating a response from the retrieved context. Instrument each one separately with timing spans, not just an overall request timer. On many stacks the embedding call and the generation call, both round trips to a hosted model, dwarf the vector search itself, which can return in single-digit milliseconds even against millions of vectors.
Once you have per-stage numbers, you'll know whether you're solving a search problem or a network problem.
Distributed tracing makes this much easier than manually timestamping code. A single trace that shows all four spans nested under one request lets you see not just how long each stage took, but whether they ran sequentially when they could have run in parallel; embedding the query and checking a cache, for instance, don't need to block each other.
Tune the index instead of guessing
If the search stage genuinely is slow, the fix is usually a parameter, not a rewrite. Approximate nearest-neighbor indexes like HNSW trade recall for speed through a search-time parameter (often called ef_search or similar), and IVF-style indexes do the same through the number of clusters probed. Turning either one down speeds up search and lowers recall; turning it up does the opposite. Measure recall against a fixed query set before and after any change, since a faster index that returns worse matches just moves the problem downstream to the model.
Test this on a copy of your production corpus, not a small synthetic one, since recall behavior at ten thousand vectors doesn't reliably predict behavior at ten million. A parameter setting that looks safe in a small test set can quietly drop recall once the index reaches production scale, where near-duplicate and borderline-similar content becomes more common.
Decide whether you need a reranker at all
A cross-encoder reranking pass over the top candidates from vector search improves result quality, but it adds a full model call to the request path, often more latency than the vector search step it's refining. It's worth it when your first-stage retrieval returns a lot of near-miss results and quality visibly suffers without it. It's not worth it if your embedding model and chunking already produce clean top results; you'd be paying latency for a quality gain nobody notices.
A middle option worth testing before committing either way: rerank only when the score gap between the first and fifth vector search result is small, since a large gap usually means the top result is already clearly best and reranking won't change the answer, just the latency.
Cache what actually repeats
Look at your query logs before building a cache. If the same or near-duplicate questions recur, common on a support or documentation site, a semantic cache keyed on query similarity can skip the search and generation stages entirely for a meaningful share of traffic. If queries are mostly unique, as they tend to be on an internal analytics tool, a query cache won't help much and you're better off spending the effort on the embedding and generation legs instead.
A semantic cache needs its own tuning too: too loose a similarity threshold and it serves a wrong answer for a genuinely different question; too tight and it barely fires. Start conservative and loosen the threshold only after checking a sample of what it's matching.
How should you set a latency budget for a RAG pipeline?
A p50 latency number hides the requests that actually frustrate users. Set a budget per stage, embedding, search, rerank, generation, and alert on p95 or p99 for each independently. A search stage that's usually 8ms but spikes to 400ms under load points at a resourcing problem on the vector database, while a generation stage with a long, flat tail usually points at a model provider issue you can't fix locally, only route around.
Turn the per-stage numbers into budgets with these rules:
- Set a separate latency budget for embedding, search, reranking, and generation instead of relying on one overall request timer.
- Alert on p95 or p99 for each stage independently, because a p50 number hides the requests that actually frustrate users.
- Read the pattern: a search stage that spikes under load points to database resourcing, while a long flat generation tail points to the model provider.
- Track cold starts separately from steady state, since a scaled-to-zero embedding or reranking function can leave p50 normal and p99 terrible.
Watch for the cold-start tax separately from steady-state latency
A serverless embedding or reranking function that scales to zero looks fast in a load test, which tends to keep it warm, and then adds a full cold-start penalty to the first request after any idle period in production. This shows up as a confusing pattern in your metrics: a normal p50 and a terrible p99, with no obvious cause in the application code. Check whether any stage in your pipeline runs on infrastructure that can scale to zero, and either keep a minimum number of instances warm or account for the cold-start tax explicitly in your latency budget instead of chasing it as a mystery.
What Good Looks Like
The latency standard is per-stage timing with independent budgets for embedding, search, reranking, and generation, tuned against a measured recall target rather than a single end-to-end number.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is vector search usually the slow part of a RAG request?
Less often than people assume. Approximate nearest-neighbor search is typically the fastest leg of the request. The embedding call and the final generation call, both round trips to a model, are the more common bottlenecks. Measure before you optimize the search index.
Will increasing recall always hurt latency?
Usually, yes, since higher recall means examining more candidates before returning results. The relationship isn't always linear, though: a poorly tuned index can be both slower and less accurate than a well-tuned one at the same recall target, so retuning sometimes improves both at once.
Should we cache the final generated answer or the retrieved chunks?
Cache the retrieved chunks first. They're stable across near-duplicate queries and reusing them still lets you regenerate a fresh answer. Caching the full generated answer is faster but riskier, since it can serve a stale response for a question whose underlying data has since changed.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
How to Benchmark Your RAG API Gateway Without Fooling Yourself
A methodology for benchmarking API gateway latency in front of a RAG pipeline, and the common mistakes that make a benchmark misleading.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.
How Vector Search Throughput Degrades as Your Index Grows
Throughput doesn't fall off gradually as a vector index grows. Here's why it degrades in steps, and how sharding, replicas, and quantization each help.
Building a Golden Set to Catch RAG Regressions Before Users Do
A step-by-step approach to building a RAG evaluation set from real queries, scoring retrieval and generation separately, and gating on regressions.
What to Check First in a RAG Pipeline Security Audit
A practical order of operations for auditing a production RAG pipeline: data exposure, prompt injection, access control, logging, and vendor risk.