Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

How Vector Search Throughput Degrades as Your Index Grows

Vector search throughput degrades in steps as an index grows, not gradually: it stays roughly flat while the index fits in memory, then drops sharply once the engine starts paging to disk. Fine results at a hundred thousand vectors and a sharp drop past a few million is a predictable pattern, not a fault.

Why does throughput degrade in steps instead of gradually?

As long as an index fits comfortably in memory, query throughput stays roughly flat as it grows. Once it outgrows available memory and the engine starts paging to disk, throughput drops sharply rather than gracefully, since a disk read is orders of magnitude slower than a memory access. Benchmark at several sizes, not just your current one and a rough guess at double, since the interesting behavior is finding where that step actually happens for your specific hardware and index configuration, not assuming it scales smoothly until it suddenly doesn't.

Sharding trades single-query latency for aggregate throughput

Splitting an index across shards lets you parallelize both ingestion and query load across more machines, raising aggregate throughput. The tradeoff is that a single query often needs to fan out to every shard and merge the results, which can add tail latency even as overall system throughput improves. This is worth benchmarking explicitly: aggregate throughput under load is a different number from p99 latency for an individual request, and a sharding decision that helps one can hurt the other.

Read replicas scale query throughput, not write throughput

Adding read replicas helps a read-heavy workload, more concurrent search queries, but does nothing for an ingestion-heavy one, since writes still have to go through the primary and propagate. If your bottleneck is actually ingestion throughput, a bulk backfill or a high-volume real-time ingestion pipeline, adding replicas is the wrong lever entirely, and the added infrastructure cost buys you nothing for that specific problem. Diagnose which side of the workload is actually constrained before reaching for either fix.

Does quantization change throughput, not just storage?

Reducing vector precision, from full floating-point to int8 or binary representations, shrinks the memory footprint, which means more of the index fits in memory or fast cache at a given hardware size. This often raises effective throughput meaningfully, not just cuts storage cost, since the bottleneck for many workloads is memory bandwidth rather than raw compute. Test the recall tradeoff carefully before committing, since a faster index that returns worse matches has just moved the cost somewhere else.

Benchmark with your real query mix, not a uniform synthetic one

A synthetic benchmark that queries uniformly across the whole index behaves differently from real production traffic, which is usually skewed: some documents or topics get queried constantly, others rarely at all. That skew affects cache hit rates and effective throughput in ways a uniform random benchmark won't reveal. Pull a representative sample of real query patterns, including the popularity skew, before trusting a benchmark's numbers as a predictor of production behavior.

For example, a team benchmarks with queries drawn uniformly across the whole index and reports comfortable headroom. In production, a small set of popular topics gets queried constantly, so caches stay warm and typical throughput looks better than the benchmark, while the long tail of rarely seen documents still hits cold storage and sets the worst-case latency. Replaying a sample of real queries, popularity skew included, shows both effects and gives the team a benchmark that predicts what production will actually do, instead of a number that merely looks reassuring.

Ingestion throughput has its own separate ceiling

It's easy to benchmark query throughput carefully and never separately measure how fast the pipeline can ingest and index new documents, until a large backfill or a burst of new content takes far longer than expected. Embedding API rate limits, index build time, and write contention with concurrent queries all cap ingestion throughput independently of anything affecting search. Measure this on its own, with a realistic bulk-ingestion scenario, rather than assuming query performance numbers tell you anything meaningful about it.

Re-benchmark after any change to chunking or embedding dimensionality

A change to chunk size affects how many vectors exist per document, and a change to embedding dimensionality affects how much memory each vector consumes, both of which shift exactly the thresholds this article is about. A throughput benchmark run before either change stops being a reliable predictor of behavior after it. Treat a chunking or embedding change as a trigger to re-run your scaling benchmark, not just your recall benchmark, since the two numbers can move in opposite directions from the same change.

Match each scaling lever to the bottleneck it actually addresses:

  • Sharding raises aggregate throughput and parallelizes ingestion, but fan-out and result merging can add tail latency for individual queries.
  • Read replicas add query capacity for read-heavy workloads and do nothing for ingestion-heavy ones, since writes still go through the primary.
  • Quantization shrinks memory so more of the index stays in fast memory, at a recall cost you should test before committing.
  • Re-run the scaling benchmark after any chunking or embedding dimensionality change, since both move the memory threshold.
Executive Capability Standard

What Good Looks Like

The scaling standard is throughput and latency benchmarked at multiple index sizes against your real query distribution, with a known memory threshold and a deliberate choice between sharding, replicas, and quantization based on which resource is actually constrained.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull your current index size and query volume, and compare them against your last load test, if one exists, to see how stale that data is.
2. Do Manually:Run a load test at your current size and a projected future size by hand, watching memory and disk I/O alongside latency.
3. Delegate:Assign an engineer to own a recurring benchmarking cadence as the index and traffic grow, rather than testing only once at launch.
4. Automate:Build the recall-and-throughput benchmark into a scheduled job that runs against production-scale data regularly, not just before a big migration.
5. Buy:Bring in a search infrastructure specialist once index size and query volume are large enough that scaling decisions carry real cost and latency stakes.

How to Get Started

Frequently Asked Questions

How do we find the memory threshold where our index starts paging to disk?

Run a load test at increasing index sizes, watching both query latency and the host's memory and disk I/O metrics together. The threshold shows up as a sudden jump in latency that correlates with disk read activity starting to appear, rather than a smooth curve, and it's specific to your hardware and index configuration, not a fixed number you can look up.

Is sharding worth the added complexity for a mid-sized index?

Usually not until a single well-provisioned machine genuinely can't hold the index in memory or can't keep up with query volume on its own. Sharding adds real operational complexity, result merging, uneven shard load, that's worth avoiding until the simpler single-node approach has actually been pushed to its limit.

Does quantization hurt recall enough to matter for most use cases?

It depends on how much recall margin your current setup has above what your use case actually needs. Many production systems have more headroom than they realize, since the difference between very good and slightly-less-good retrieval often isn't noticeable to end users, but this is worth measuring on your own data rather than assuming.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides