The Real Cost Drivers in a RAG Pipeline
Say your team's monthly bill for a RAG feature comes in higher than expected. Before cutting anything, it helps to know that the cost lives in four different places, and the fix for one can make another worse.
Here's how to find where your money is actually going.
Why do embedding costs scale with ingestion, not just queries?
It's easy to budget for query-time embedding costs and forget that ingestion embeds every chunk of every document, often repeatedly if your pipeline re-embeds on every update rather than only on change. If you're re-indexing a full corpus nightly instead of incrementally updating changed documents, you're paying to re-embed unchanged content every single day. Check your ingestion logs for how much of your embedding spend is redundant re-embedding versus genuinely new content.
A content hash comparison before re-embedding is a cheap fix: store a hash of each chunk's text alongside its vector, and only re-embed when the hash changes. This alone eliminates the most common source of wasted embedding spend without touching anything else in the pipeline.
Storage and compute scale with index size and replica count
Vector database cost usually scales with the number of vectors stored, their dimensionality, and how many replicas you run for availability. Reducing embedding dimensionality, some models offer a smaller output size at a modest quality cost, or quantizing vectors to int8 or binary representations can cut storage and memory footprint substantially, since the index has to fit in memory for fast search on most engines. This is worth testing against your recall benchmark before committing to it broadly.
Replica count is worth revisiting on its own schedule too. A team that added a third replica during an early scaling scare sometimes never revisits whether it's still needed once traffic patterns settle down, and it's a fixed cost that compounds every month it goes unreviewed.
Reranking and generation cost more per token than the search itself
Say your pipeline retrieves the top 20 chunks and reranks them before passing the top 5 to the model. Each of those steps has a per-call or per-token cost that's usually far higher than the vector search itself, which is typically a cheap, self-hosted or low-cost managed operation. If you're retrieving more chunks than you actually use in the final prompt, you're paying reranking cost for candidates that get discarded. Tightening your top-k before reranking is often the cheapest lever available.
It's worth pricing out the reranker separately from the generation model, since teams often bundle both into a single "the AI part is expensive" line item and never find out which one actually drives the bill. In many pipelines, generation, not reranking, is the larger cost, simply because the final prompt carries far more tokens than a reranking pass does.
How does padded context inflate generation cost?
If your chunking strategy produces overlapping or oversized chunks to be safe, you're paying generation-time token cost for redundant context on every single query, not just occasionally. Say a document is chunked with 500-token overlaps between adjacent chunks; retrieving five of those chunks means paying for roughly 2,000 tokens of duplicated content. Tightening chunk boundaries and overlap, and testing whether retrieval quality actually degrades without the extra padding, is worth doing before assuming you need a bigger context window.
Know when switching to a cheaper embedding model is worth it
Not all embedding models cost the same, and the price difference between a larger model and a smaller, cheaper one can be substantial at ingestion volume. The tradeoff is retrieval quality: a cheaper model with lower-dimensional embeddings can lose enough recall that your reranker, or your users, notice. Before switching, re-embed a representative sample of your corpus with the cheaper model and measure recall against your current baseline, not just against a generic published benchmark, since the quality gap varies a lot by domain and document type.
Batch what doesn't need to happen in real time
Say your ingestion pipeline embeds documents the moment they're uploaded, one at a time, through a real-time API call. Most embedding providers offer a batch endpoint at a meaningfully lower per-token price for workloads that can tolerate a delay of minutes to hours. If your source documents aren't time-sensitive, routing bulk or backfill ingestion through the batch path, while keeping real-time embedding only for genuinely urgent updates, can cut a meaningful share of ingestion cost without touching query-time behavior at all.
Work through these cost levers, starting with ingestion:
- Check ingestion logs for redundant re-embedding, and store a content hash per chunk so you only re-embed text that changed.
- Reduce vector dimensionality or quantize vectors to lower storage and memory cost, testing recall before you commit.
- Retrieve only as many chunks as the final prompt uses, since extra reranked chunks cost more per token than the search itself.
- Tighten chunk boundaries and overlaps so you do not pay for duplicated context on every generation call.
- Route bulk or backfill ingestion through a provider's batch endpoint when the documents are not time-sensitive.
What Good Looks Like
The cost standard is knowing your spend broken out by ingestion, storage, reranking, and generation separately, with incremental ingestion in place so unchanged documents aren't re-embedded on every run.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should we cut costs by reducing how many chunks we retrieve?
Test it before assuming it hurts quality. Many pipelines retrieve more chunks than the model actually needs, as a hedge against poor retrieval. If your embedding and chunking are solid, dropping from ten retrieved chunks to five often costs little in answer quality while cutting reranking and generation spend proportionally.
Is quantizing vectors worth the accuracy tradeoff?
It depends on your recall margin. If your current setup has recall well above what your use case needs, quantization to int8 or binary vectors can cut memory and storage cost significantly with a small, often acceptable accuracy loss. Measure recall before and after on your actual query set rather than trusting a vendor's general benchmark.
Where should a small team look first when a RAG bill spikes unexpectedly?
Check ingestion first, not queries. A spike is more often caused by a pipeline re-embedding an entire corpus unnecessarily, or a bug causing duplicate ingestion, than by a real increase in user query volume. Query costs tend to grow gradually with usage; ingestion cost spikes are usually a configuration problem.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
A FinOps Checklist for Teams Before Their First Big Cloud Bill
The cost-optimization checklist to run before your cloud bill becomes a board topic, plus the five mistakes that quietly undo every fix on the list.
Three Ways to Cut Cloud Spend, and When Each One Works
Rightsizing, committed-use discounts, and architecture changes all cut cloud spend differently. Here's how to pick the right one for your situation.
Build vs. Buy for Your Security Tooling Stack
A decision framework for when to build DevSecOps tooling in-house versus buying a platform, based on team size, maintenance burden and audit needs.
A CTO's Framework for Cutting Infrastructure Costs
A decision framework for engineering leaders trying to cut cloud and tooling spend without slowing the team down or cutting into future capacity.
Where RAG Latency Actually Goes, and How to Budget It
Break a RAG request into its four latency stages, find out which one is actually slow, and set a budget for each before you start tuning blindly.