Designing a Load Test That Finds Where RAG Actually Breaks
A useful RAG load test replays a realistic mix of real queries, ramps load gradually to find where latency climbs disproportionately, and exercises ingestion alongside queries. A test that hammers one repeated query at increasing volume mostly measures your cache, not your system.
How should you model your real concurrent query mix?
Pull a sample of real queries from your logs and replay a realistic distribution, not a single query repeated thousands of times. Include the popularity skew that real traffic has, some queries recur constantly, most are closer to unique, since that skew affects cache hit rates and index access patterns in ways a uniform test won't reveal. A load test built on unrealistic traffic gives you a confident, precise number that doesn't actually predict what will happen in production.
Include a mix of cheap and expensive queries too, short filters against a narrow collection alongside broad, high-top-k requests that trigger reranking, since a test built only from the easy end of that spectrum understates how much load the expensive end actually places on the system.
Why ramp gradually and watch for the knee, not just the ceiling?
Increase load steadily rather than jumping straight to a target queries-per-second number, and watch for the point where latency starts climbing disproportionately faster than load, commonly called the knee in the curve. That point, not your test's maximum throughput number, is the more useful planning figure, since it tells you where the system starts to strain well before it actually falls over, giving you an early warning threshold rather than just a breaking point.
Test cold and warm separately
A freshly started system with an unwarmed index or a cold cache behaves very differently under load than one that's been serving steady traffic for hours. Since production experiences a cold state after every deploy and every scale-out event, not just at initial launch, test both explicitly and know which condition any given result describes. A load test result that only covers the warm, steady-state case tells you nothing about the minutes right after your next deploy, which is often exactly when load testing would have mattered most.
Load test ingestion concurrently with queries
Production doesn't pause ingestion while serving query traffic, so a load test that exercises only one at a time misses the real contention between them, ingestion write load competing with query read load on the same underlying infrastructure. Run a realistic ingestion workload alongside your query load test to see how they actually interact, since a system that handles each one fine in isolation can still struggle when both are happening at once, which is the normal state in production.
Watch the failure mode, not just the pass or fail line
When the system does start failing under load, the important question is how: responses getting slower but still correct is a far safer failure mode than requests erroring out or the service crashing outright. Design toward graceful degradation deliberately, and use your load test specifically to confirm which failure mode you actually get under real pressure, rather than only checking whether the system holds up at your target load and stopping there.
Turn findings into an explicit capacity plan
A load test that produces a number nobody acts on isn't worth much more than not running one. Translate the knee-point findings into a concrete plan: at what load do you scale up, what's the lead time to do it, and how does that compare against your availability target, since a downtime budget shrinks fast at higher availability tiers1 and a capacity shortfall that isn't caught ahead of time eats directly into that budget.
A load test that predicts production runs in this order:
- Replay a sample of real queries with realistic popularity skew instead of one query repeated thousands of times.
- Ramp load steadily and note the knee, the point where latency climbs faster than load.
- Run cold and warm states separately, since every deploy and scale-out starts cold.
- Run ingestion concurrently with queries to expose contention on shared infrastructure.
- Record how the system fails, since slower but correct is safer than errors or crashes.
- Convert the knee point into a capacity trigger with a defined scale-up lead time.
Share results in a format the rest of the team will actually read
A load test report full of raw percentile tables tends to get skimmed once and forgotten. Summarize the knee point, the failure mode observed, and the resulting capacity trigger in plain language at the top, with the detailed numbers available for whoever wants to dig in. A result that only the engineer who ran the test understands doesn't actually change how the team plans for growth, and a plan nobody outside that one engineer can act on isn't really a plan yet.
What Good Looks Like
The load testing standard is a gradual ramp against a realistic query mix, tested both cold and warm, with ingestion running concurrently, translated into a specific capacity plan tied to your availability target.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should we re-run our load test?
After any meaningful change to chunking, embedding dimensionality, or index configuration, since each of those shifts the resource profile the original test measured. Outside of specific triggers like that, a quarterly re-run catches drift from gradual growth in corpus size and traffic that no single change would have prompted on its own.
Should load testing happen against production or a separate environment?
A separate, production-scale environment is safer and lets you push well past normal limits without risking real user traffic. If a fully separate environment isn't practical, a carefully bounded test against production, with a defined blast radius and a fast abort plan, is better than not load testing at all, but treat it with real caution.
What's the most common mistake teams make when load testing a RAG pipeline?
Testing the query path in isolation and never testing it alongside a realistic ingestion workload. Since production runs both simultaneously, a load test that only exercises one gives you a number that looks solid in a report but doesn't hold up against how the system actually behaves once it's live.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Catching Retrieval API Schema Drift Before It Breaks Things
Consumer-driven contract tests catch a retrieval API's silent schema drift, a changed field type or a dropped value, before it breaks a caller in production.
Building a Golden Set to Catch RAG Regressions Before Users Do
A step-by-step approach to building a RAG evaluation set from real queries, scoring retrieval and generation separately, and gating on regressions.
What Synthetic Monitoring Catches That Your Alerts Don't
How synthetic transaction probes catch outages that server metrics and error-rate alerts miss, and how to set them up without drowning in false alarms.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
A Checklist for Spinning Up Test Environments on Demand
A checklist for building ephemeral, per-branch test environments, covering the pitfalls that turn a promising idea into a slow, flaky, expensive one.
Running a Load Test That Actually Tells You Something Useful
A step-by-step approach to load testing that finds your real breaking point, not just a green checkmark that traffic below some threshold works fine.