How Vector Search Throughput Degrades as Your Index Grows
Vector search throughput degrades in steps as an index grows, not gradually: it stays roughly flat while the index fits in memory, then drops sharply once the engine starts paging to disk. Fine results at a hundred thousand vectors and a sharp drop past a few million is a predictable pattern, not a fault.
Why does throughput degrade in steps instead of gradually?
As long as an index fits comfortably in memory, query throughput stays roughly flat as it grows. Once it outgrows available memory and the engine starts paging to disk, throughput drops sharply rather than gracefully, since a disk read is orders of magnitude slower than a memory access. Benchmark at several sizes, not just your current one and a rough guess at double, since the interesting behavior is finding where that step actually happens for your specific hardware and index configuration, not assuming it scales smoothly until it suddenly doesn't.
Sharding trades single-query latency for aggregate throughput
Splitting an index across shards lets you parallelize both ingestion and query load across more machines, raising aggregate throughput. The tradeoff is that a single query often needs to fan out to every shard and merge the results, which can add tail latency even as overall system throughput improves. This is worth benchmarking explicitly: aggregate throughput under load is a different number from p99 latency for an individual request, and a sharding decision that helps one can hurt the other.
Read replicas scale query throughput, not write throughput
Adding read replicas helps a read-heavy workload, more concurrent search queries, but does nothing for an ingestion-heavy one, since writes still have to go through the primary and propagate. If your bottleneck is actually ingestion throughput, a bulk backfill or a high-volume real-time ingestion pipeline, adding replicas is the wrong lever entirely, and the added infrastructure cost buys you nothing for that specific problem. Diagnose which side of the workload is actually constrained before reaching for either fix.
Does quantization change throughput, not just storage?
Reducing vector precision, from full floating-point to int8 or binary representations, shrinks the memory footprint, which means more of the index fits in memory or fast cache at a given hardware size. This often raises effective throughput meaningfully, not just cuts storage cost, since the bottleneck for many workloads is memory bandwidth rather than raw compute. Test the recall tradeoff carefully before committing, since a faster index that returns worse matches has just moved the cost somewhere else.
Benchmark with your real query mix, not a uniform synthetic one
A synthetic benchmark that queries uniformly across the whole index behaves differently from real production traffic, which is usually skewed: some documents or topics get queried constantly, others rarely at all. That skew affects cache hit rates and effective throughput in ways a uniform random benchmark won't reveal. Pull a representative sample of real query patterns, including the popularity skew, before trusting a benchmark's numbers as a predictor of production behavior.
For example, a team benchmarks with queries drawn uniformly across the whole index and reports comfortable headroom. In production, a small set of popular topics gets queried constantly, so caches stay warm and typical throughput looks better than the benchmark, while the long tail of rarely seen documents still hits cold storage and sets the worst-case latency. Replaying a sample of real queries, popularity skew included, shows both effects and gives the team a benchmark that predicts what production will actually do, instead of a number that merely looks reassuring.
Ingestion throughput has its own separate ceiling
It's easy to benchmark query throughput carefully and never separately measure how fast the pipeline can ingest and index new documents, until a large backfill or a burst of new content takes far longer than expected. Embedding API rate limits, index build time, and write contention with concurrent queries all cap ingestion throughput independently of anything affecting search. Measure this on its own, with a realistic bulk-ingestion scenario, rather than assuming query performance numbers tell you anything meaningful about it.
Re-benchmark after any change to chunking or embedding dimensionality
A change to chunk size affects how many vectors exist per document, and a change to embedding dimensionality affects how much memory each vector consumes, both of which shift exactly the thresholds this article is about. A throughput benchmark run before either change stops being a reliable predictor of behavior after it. Treat a chunking or embedding change as a trigger to re-run your scaling benchmark, not just your recall benchmark, since the two numbers can move in opposite directions from the same change.
Match each scaling lever to the bottleneck it actually addresses:
- Sharding raises aggregate throughput and parallelizes ingestion, but fan-out and result merging can add tail latency for individual queries.
- Read replicas add query capacity for read-heavy workloads and do nothing for ingestion-heavy ones, since writes still go through the primary.
- Quantization shrinks memory so more of the index stays in fast memory, at a recall cost you should test before committing.
- Re-run the scaling benchmark after any chunking or embedding dimensionality change, since both move the memory threshold.
What Good Looks Like
The scaling standard is throughput and latency benchmarked at multiple index sizes against your real query distribution, with a known memory threshold and a deliberate choice between sharding, replicas, and quantization based on which resource is actually constrained.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How do we find the memory threshold where our index starts paging to disk?
Run a load test at increasing index sizes, watching both query latency and the host's memory and disk I/O metrics together. The threshold shows up as a sudden jump in latency that correlates with disk read activity starting to appear, rather than a smooth curve, and it's specific to your hardware and index configuration, not a fixed number you can look up.
Is sharding worth the added complexity for a mid-sized index?
Usually not until a single well-provisioned machine genuinely can't hold the index in memory or can't keep up with query volume on its own. Sharding adds real operational complexity, result merging, uneven shard load, that's worth avoiding until the simpler single-node approach has actually been pushed to its limit.
Does quantization hurt recall enough to matter for most use cases?
It depends on how much recall margin your current setup has above what your use case actually needs. Many production systems have more headroom than they realize, since the difference between very good and slightly-less-good retrieval often isn't noticeable to end users, but this is worth measuring on your own data rather than assuming.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Where RAG Latency Actually Goes, and How to Budget It
Break a RAG request into its four latency stages, find out which one is actually slow, and set a budget for each before you start tuning blindly.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.
Building a Golden Set to Catch RAG Regressions Before Users Do
A step-by-step approach to building a RAG evaluation set from real queries, scoring retrieval and generation separately, and gating on regressions.
What to Check First in a RAG Pipeline Security Audit
A practical order of operations for auditing a production RAG pipeline: data exposure, prompt injection, access control, logging, and vendor risk.
Catching Retrieval API Schema Drift Before It Breaks Things
Consumer-driven contract tests catch a retrieval API's silent schema drift, a changed field type or a dropped value, before it breaks a caller in production.