Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Sizing Your Vector Database Before It Falls Over in Production

Most teams size their first vector database off whatever ran during the proof of concept: a few thousand documents, one query at a time, no rerank step. That number has almost nothing to do with what the index needs once real users start asking questions against a real corpus.

Capacity planning for a production retrieval-augmented generation (RAG) system means sizing three things separately: how much memory your index needs, how many queries it can serve at once, and how much of both you keep in reserve before you're forced into an emergency migration.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you size an index from vectors and dimension, not documents?

Document count is the wrong unit. What actually drives memory is the number of vector embeddings you store, the dimension of each one, and the index type you pick. An HNSW index, the graph-based approach most vector databases default to, keeps a neighbor list per vector on top of the raw floats, so its memory footprint runs well above the raw vector data alone. Switching from a 384-dimension embedding model to a 1,536-dimension one doesn't just scale memory with dimension: it also changes how many vectors fit per shard and how long each similarity search takes.

Work this out before you provision anything: multiply your expected vector count by the dimension, add the index overhead your vector database's documentation states for its default index type, then add replica copies. If you chunk documents into passages rather than storing one vector per document, your vector count is a multiple of document count, not equal to it, and that multiple is usually the single biggest sizing error teams make.

Separate query throughput from corpus size

Corpus size tells you how big the index is. It says nothing about how many searches that index has to answer per second, and those two numbers scale independently. A support-ticket RAG system with 50,000 documents and 200 concurrent users has a completely different throughput profile than a code-search tool with two million documents and five internal users.

Break your latency budget into hops: the embedding call that turns the user's question into a vector, the approximate nearest-neighbor search itself, any reranking pass, and the generation call. Each hop has its own concurrency limit, and the slowest one sets your real ceiling no matter how fast the others run. If your embedding model runs on a shared inference endpoint, that endpoint's queue, not your vector database, is often the actual bottleneck, so profile the whole path before you conclude the vector store needs more nodes.

How much headroom should you build in before the corpus doubles?

Say your index currently holds 500,000 vectors and comfortably serves your traffic. If your ingestion pipeline adds new source documents every week, work out how many months it takes to double that count, and provision for the doubled number now rather than when you hit the current ceiling. Vector database migrations under load are harder than most schema migrations: you're often re-embedding content with a newer model at the same time you're trying to move it, and the two changes compound.

A simple rule that holds up: keep enough spare capacity that you could absorb a sudden doubling in either corpus size or query volume without a re-architecture, and revisit that buffer every quarter as both numbers move. Treat the buffer as a number you track, not a one-time provisioning decision you make and forget.

Watch the ingestion and re-embedding cost, not just storage

The overlooked cost in vector search capacity planning is compute, not storage. Every time you change embedding models, whether for quality or to fix a vendor deprecation, you have to re-embed your entire corpus and rebuild the index, and that job scales with vector count and embedding model latency, not with how much disk the vectors take up. Plan for this as a recurring operational cost with its own capacity needs, inference throughput for the re-embedding job, not a one-off migration.

If your ingestion pipeline batches new documents nightly, check whether that batch job's runtime is trending upward as the corpus grows. A batch that finishes in twenty minutes today can quietly grow past your ingestion window in a few months if nobody's watching the trend line, and you won't notice until documents start showing up stale in search results.

Common capacity-planning mistakes

A few patterns show up again and again in production RAG deployments:

  • Sizing off the proof-of-concept's document count instead of the chunked vector count
  • Assuming replicas fix throughput problems that are actually caused by the embedding or rerank hop
  • Treating re-embedding as a one-time migration instead of a recurring capacity need
  • Provisioning for today's traffic with no plan for what doubles it
  • Never revisiting the headroom number after the first sizing exercise

Any one of these is survivable. Two or three together are usually what turns a routine growth quarter into an emergency migration.

Executive Capability Standard

What Good Looks Like

Good capacity planning for a production RAG system means you can state, at any time, how much headroom you have in both index memory and query throughput, and you revisit both numbers on a schedule rather than when something breaks.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your vector database's documentation on index memory overhead for the specific index type you use, HNSW, IVF, or otherwise, since the multiplier over raw vector size varies a lot between them.
2. Do Manually:Track vector count, dimension, and query volume in a spreadsheet you update monthly, and calculate headroom by hand before each planning cycle.
3. Delegate:Have an infrastructure engineer own capacity headroom as a named responsibility with a recurring check-in, instead of leaving it as something everyone assumes someone else is watching.
4. Automate:Set up alerting on index memory usage and p95 query latency so you get a warning while you still have runway, not an incident when you don't.
5. Buy:Use your cloud or vector database provider's built-in capacity dashboards and autoscaling where they're available, rather than building your own monitoring layer from scratch.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

How much headroom should I build into a new vector search deployment?

Enough to absorb a doubling in either corpus size or query volume without a re-architecture is a reasonable starting point. Track both numbers monthly and revisit the buffer every quarter rather than setting it once at launch. Teams that only check capacity when something slows down are always planning reactively instead of ahead of the growth curve.

Does adding replicas fix capacity problems in a RAG system?

Only if the bottleneck is actually search throughput. Replicas add read capacity to the vector index itself, but if your latency problem is really in the embedding call or the reranking step, adding replicas won't touch it. Profile each hop in the retrieval path before you spend on more nodes.

When should I move off a single-node vector store?

When your index no longer fits comfortably in memory on that node, or a single node can't serve your peak concurrent query volume without queueing. Both are measurable: watch memory headroom and p95 query latency under real load, and move before either one is consistently tight, not after users notice slow answers.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides