Setting Spend Caps on a RAG Pipeline Without Breaking It
To cap spend on a RAG pipeline without breaking it, set limits per tenant instead of globally, measure cost in tokens across every stage, and degrade gracefully at the limit instead of rejecting requests. A global cap built in a hurry punishes every user for one tenant's traffic, and a hard rejection looks like an outage.
What unit should you cap: requests or tokens?
A request-based cap is simple for users to understand but a poor proxy for cost, since one request with a long document and a large top-k can cost far more than ten simple ones. A token-based cap tracks cost accurately but is harder for a user to reason about in advance. Many teams end up with both: a token-based cap for actual cost control, translated into a simpler request-based indicator shown to users so the limit doesn't feel arbitrary.
Whichever unit you choose, measure it consistently across every stage of the pipeline. A cap that counts tokens at the generation step but ignores tokens spent on embedding and reranking will systematically undercount real cost, and undercounting compounds quietly until the actual bill stops matching what the cap implied it should be.
Cap per tenant, not just globally
A single global cap means one tenant running an unusually heavy workload can consume the entire budget and degrade the experience for everyone else sharing it. Per-tenant caps, sized to each tenant's actual plan or usage tier, with a shared overflow pool for genuine bursts, keep one noisy tenant from starving the rest while still letting you use spare capacity efficiently rather than leaving it idle.
For example, suppose two tenants share one budget and one of them starts a bulk workflow that embeds long documents with a large top-k. Under a single global cap, the second tenant's ordinary queries start failing even though nothing about their usage changed. With a per-tenant cap and a shared overflow pool, the heavy tenant reaches its own ceiling and draws on spare capacity, while the second tenant is unaffected. The team can also see exactly which tenant drove the spend, which makes a plan upgrade conversation concrete instead of speculative.
What should happen at the limit: queue, degrade, or reject?
Queueing a request past the limit adds latency but preserves the full experience once it processes, which fits a background or batch workflow where a short delay is acceptable. Degrading gracefully, skipping the reranking pass, retrieving fewer chunks, still returns something useful and fits an interactive feature where users expect a response now. A hard rejection is the simplest to implement and the worst experience, and it's worth reserving for genuine abuse rather than routine peak usage.
The three responses at the limit suit different workloads:
- Queue the request when a short delay is acceptable, such as background or batch work, so it still gets the full experience once it runs.
- Degrade gracefully by skipping the reranking pass or retrieving fewer chunks when users expect an answer right now.
- Reject outright only for genuine abuse, since it's the simplest to build and the worst experience for routine peak usage.
Rate-limit ingestion separately from queries
If ingestion and query-time embedding share the same provider quota, a large bulk upload can silently starve query traffic that has nothing to do with it, showing up as a mysterious latency spike with no obvious cause in the query path itself. Give ingestion its own quota, or explicitly prioritize query traffic over ingestion when both compete for the same upstream capacity, so a backfill job never degrades the live user experience.
Show the user something before they hit the wall
A usage indicator, or a warning once a tenant crosses most of its allotment, turns a hard cutoff from a surprise into an expected event. This is a small addition on top of the underlying rate limiting logic, but it's the difference between a limit that feels like a broken feature and one that feels like a plan boundary the user already understood.
Account for the reranker and generation model separately
A cap sized only around embedding and search cost will miss the fact that reranking and generation usually cost more per request than the retrieval step itself. If your rate limiting logic only tracks the initial search call, a tenant can still drive up spend through expensive reranking or long generation outputs without ever tripping the limit. Track cost across the whole request path, not just the stage that happens to be easiest to instrument first.
Revisit caps as usage patterns change, not once and never again
A cap sized around a tenant's usage at signup tends to become stale within a few months, either too tight for a tenant that's grown into heavier legitimate use or too loose for one whose usage pattern shifted toward more expensive query types. Review caps on a regular schedule against actual usage data, rather than only adjusting them reactively when a tenant complains or a cost spike forces the question.
What Good Looks Like
The rate limiting standard is a token-based, per-tenant cap with a defined behavior at the limit, queue, degrade, or reject, chosen deliberately per use case, and ingestion quota kept separate from query quota.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should free-tier and paid-tier users share the same rate limiting logic?
The underlying mechanism can be the same, but the thresholds and what happens at the limit should differ. A reasonable pattern is generous queueing or graceful degradation for paid tiers and firmer rejection for free tiers, using the same infrastructure with different configuration per tier rather than building separate systems.
How do we set the initial cap for a new tenant with no usage history?
Start conservative, based on their plan tier rather than an assumption of typical usage, and monitor closely during their first weeks. It's easier to raise a cap for a tenant who's clearly hitting it with legitimate usage than to walk back a cap that was set too generously from the start.
What's the risk of degrading rather than rejecting at the limit?
Users may not immediately notice the quality drop from skipped reranking or fewer retrieved chunks, which can mask a capacity problem that would otherwise prompt an upgrade conversation or a capacity fix. Track degraded responses as their own metric so the team can see how often it's happening, not just that the pipeline never technically returned an error.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Setting Rate Limits Without Breaking Your Best Customers
A decision guide for setting per-tier rate limits and spend caps that protect your infrastructure without throttling the customers you most want to keep.
Setting Rate Limits and Spend Caps That Don't Break Real Usage
How to set rate limits and spend caps that stop abuse and runaway costs without throttling your actual customers, with a worked example.
The Rate-Limit Gaps a 30-Minute Audit Usually Finds
A short, practical checklist for finding the rate-limiting and spend-cap gaps that let one bad actor or one buggy client burn through your budget.
Setting Rate Limits That Protect Budget, Not Just Uptime
A practical checklist for designing rate limits and spend caps that stop runaway costs and abuse without breaking legitimate customer usage.
Setting Spend Caps Before Your Agents Set Them for You
A decision guide for setting rate limits and spend caps on agentic workloads, so a stuck loop or a bad actor can't turn into an open-ended bill.
Setting Rate Limits and Spend Caps Without Breaking Real Users
How to set request-rate limits and spend caps for AI model serving separately, with a worksheet for setting your first cap.