Sizing Your Vector Database Before It Falls Over in Production
Most teams size their first vector database off whatever ran during the proof of concept: a few thousand documents, one query at a time, no rerank step. That number has almost nothing to do with what the index needs once real users start asking questions against a real corpus.
Capacity planning for a production retrieval-augmented generation (RAG) system means sizing three things separately: how much memory your index needs, how many queries it can serve at once, and how much of both you keep in reserve before you're forced into an emergency migration.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you size an index from vectors and dimension, not documents?
Document count is the wrong unit. What actually drives memory is the number of vector embeddings you store, the dimension of each one, and the index type you pick. An HNSW index, the graph-based approach most vector databases default to, keeps a neighbor list per vector on top of the raw floats, so its memory footprint runs well above the raw vector data alone. Switching from a 384-dimension embedding model to a 1,536-dimension one doesn't just scale memory with dimension: it also changes how many vectors fit per shard and how long each similarity search takes.
Work this out before you provision anything: multiply your expected vector count by the dimension, add the index overhead your vector database's documentation states for its default index type, then add replica copies. If you chunk documents into passages rather than storing one vector per document, your vector count is a multiple of document count, not equal to it, and that multiple is usually the single biggest sizing error teams make.
Separate query throughput from corpus size
Corpus size tells you how big the index is. It says nothing about how many searches that index has to answer per second, and those two numbers scale independently. A support-ticket RAG system with 50,000 documents and 200 concurrent users has a completely different throughput profile than a code-search tool with two million documents and five internal users.
Break your latency budget into hops: the embedding call that turns the user's question into a vector, the approximate nearest-neighbor search itself, any reranking pass, and the generation call. Each hop has its own concurrency limit, and the slowest one sets your real ceiling no matter how fast the others run. If your embedding model runs on a shared inference endpoint, that endpoint's queue, not your vector database, is often the actual bottleneck, so profile the whole path before you conclude the vector store needs more nodes.
How much headroom should you build in before the corpus doubles?
Say your index currently holds 500,000 vectors and comfortably serves your traffic. If your ingestion pipeline adds new source documents every week, work out how many months it takes to double that count, and provision for the doubled number now rather than when you hit the current ceiling. Vector database migrations under load are harder than most schema migrations: you're often re-embedding content with a newer model at the same time you're trying to move it, and the two changes compound.
A simple rule that holds up: keep enough spare capacity that you could absorb a sudden doubling in either corpus size or query volume without a re-architecture, and revisit that buffer every quarter as both numbers move. Treat the buffer as a number you track, not a one-time provisioning decision you make and forget.
Watch the ingestion and re-embedding cost, not just storage
The overlooked cost in vector search capacity planning is compute, not storage. Every time you change embedding models, whether for quality or to fix a vendor deprecation, you have to re-embed your entire corpus and rebuild the index, and that job scales with vector count and embedding model latency, not with how much disk the vectors take up. Plan for this as a recurring operational cost with its own capacity needs, inference throughput for the re-embedding job, not a one-off migration.
If your ingestion pipeline batches new documents nightly, check whether that batch job's runtime is trending upward as the corpus grows. A batch that finishes in twenty minutes today can quietly grow past your ingestion window in a few months if nobody's watching the trend line, and you won't notice until documents start showing up stale in search results.
Common capacity-planning mistakes
A few patterns show up again and again in production RAG deployments:
- Sizing off the proof-of-concept's document count instead of the chunked vector count
- Assuming replicas fix throughput problems that are actually caused by the embedding or rerank hop
- Treating re-embedding as a one-time migration instead of a recurring capacity need
- Provisioning for today's traffic with no plan for what doubles it
- Never revisiting the headroom number after the first sizing exercise
Any one of these is survivable. Two or three together are usually what turns a routine growth quarter into an emergency migration.
What Good Looks Like
Good capacity planning for a production RAG system means you can state, at any time, how much headroom you have in both index memory and query throughput, and you revisit both numbers on a schedule rather than when something breaks.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
If you're working toward SOC 2, Drata can pull evidence that you monitor capacity and availability on a schedule, instead of you manually screenshotting a dashboard before every audit.
Vanta works the same way for continuous capacity and availability monitoring evidence, which is worth setting up before an auditor asks for it rather than after.
Frequently Asked Questions
How much headroom should I build into a new vector search deployment?
Enough to absorb a doubling in either corpus size or query volume without a re-architecture is a reasonable starting point. Track both numbers monthly and revisit the buffer every quarter rather than setting it once at launch. Teams that only check capacity when something slows down are always planning reactively instead of ahead of the growth curve.
Does adding replicas fix capacity problems in a RAG system?
Only if the bottleneck is actually search throughput. Replicas add read capacity to the vector index itself, but if your latency problem is really in the embedding call or the reranking step, adding replicas won't touch it. Profile each hop in the retrieval path before you spend on more nodes.
When should I move off a single-node vector store?
When your index no longer fits comfortably in memory on that node, or a single node can't serve your peak concurrent query volume without queueing. Both are measurable: watch memory headroom and p95 query latency under real load, and move before either one is consistently tight, not after users notice slow answers.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Sizing Platform Capacity Around How Often Your Team Ships
A way to size infrastructure headroom against your traffic pattern, deploy cadence, and uptime target, instead of picking a round percentage and hoping.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
How to Build an Infrastructure Headroom Worksheet Before You Need One
A worksheet-based way for CTOs to track infrastructure headroom by service, so capacity decisions happen before an outage forces them.
How Much Infrastructure Headroom Is Actually Enough?
Capacity planning usually means reacting to a page instead of a forecast. Here is how to pick a headroom target and spot your next constraint before it hits.
How Much Cloud Headroom Should You Actually Keep?
A practical way to size compute headroom against real traffic spikes, so engineering isn't paying for capacity it never uses or scrambling when demand jumps.
A Capacity Planning Runbook for Teams Tired of Fire Drills
A concrete way to set headroom targets, watch the right leading indicators, and decide what to pre-provision before the next launch catches you flat.