Setting Rate Limits and Spend Caps Without Breaking Real Users
Rate limiting for a model-serving endpoint has two different jobs that get lumped into one setting way too often: protecting your infrastructure from being overwhelmed, and protecting your budget from a runaway integration or an abusive caller. Those two jobs need different limits, checked at different points.
A single global rate limit tends to under-protect one of the two goals no matter how you tune it.
Two different caps, not one
- A request-rate cap, protecting infrastructure, set close to what your GPUs can actually handle before latency degrades for everyone.
- A spend cap, protecting budget, set per caller or per feature based on expected usage, independent of whether the infrastructure could technically handle more.
A caller can be well under the request-rate cap and still blow through a spend cap if their requests are unusually large, long prompts, high output token counts, so track both separately rather than assuming one implies the other.
Where to check the limit: gateway versus model server
Checking the limit at your API gateway, before a request reaches the model server, is cheaper and catches abuse earlier, but it needs to know enough about the request, expected token count, which model it's headed for, to make a meaningful decision.
Checking it at the model server catches everything the gateway might miss, but by then you've already paid for whatever work happened before the check. The practical answer is both: a coarse check at the gateway to reject obvious abuse fast, and a finer per-model, per-token check closer to where the actual cost is incurred.
A worksheet for setting your first spend cap
Work through four numbers for each caller type: typical requests per day, typical tokens per request, your cost per thousand tokens, and the multiple above typical usage you're comfortable allowing before you want a human to look at it, often somewhere between three and five times normal.
Multiply the first three together for a typical daily cost, then apply your comfort multiple for the actual cap. Review actual usage against this cap monthly and adjust it; a cap set once from a guess and never revisited either blocks legitimate growth or fails to catch a real problem.
What happens when a caller hits the limit
A hard rejection is the simplest response and the easiest to get wrong for a legitimate customer having an unusually busy day. Consider a soft response instead for the spend cap specifically: throttle rather than reject outright, and notify someone on your team so a human can decide whether to raise the cap.
Reserve hard rejection for the request-rate cap protecting infrastructure, where the risk of letting traffic through unchecked is a real outage rather than an unexpected bill you can review after the fact and, if needed, follow up on with the customer directly.
A useful decision rule is to ask what the worst outcome is if the request goes through. If the worst outcome is an outage for every other customer, reject it. If the worst outcome is a larger bill for one caller, throttle and notify. For example, a customer who starts a large batch import can be slowed to a lower rate with a message explaining why, while a runaway loop sending the same request over and over should be blocked at the gateway. Writing that rule down means the on-call engineer is not inventing policy in the middle of an incident.
Communicating limits to customers before they hit them
A caller who hits an undocumented limit assumes something is broken, not that they've been capped. Publish your rate limits and, where practical, your spend-cap defaults, so integrators can design around them instead of discovering them through failed requests.
For spend caps specifically, consider a usage dashboard or a proactive notification well before the cap, at a comfortable majority of it, so a customer scaling up gets a heads-up instead of a surprise rejection at the worst possible moment for whatever they're building on top of you.
Rate-limiting mistakes that hit real customers
- One global limit applied to every caller regardless of their actual usage pattern or contract.
- No monitoring on how close callers run to their limit, so you find out about a problem only after they've already been blocked.
- Treating a spend cap breach the same as a security incident, when it's usually just a customer using the product more than expected.
The goal is protecting the business, not punishing growth. A customer who outgrows a cap because their own usage is genuinely growing is a sales conversation, not a security event.
What Good Looks Like
A working rate-limiting setup tracks request-rate and spend caps as separate numbers, checks them at both the gateway and closer to the model, and routes a cap breach to a person rather than only to an automatic hard rejection.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should request-rate limits and spend caps be the same setting?
No. A request-rate cap protects your infrastructure from being overwhelmed and should be based on what your GPUs can actually handle. A spend cap protects your budget and depends on request size, not just count. A caller can stay under one and still exceed the other, so track them separately.
Where should we check rate limits, at the gateway or the model server?
Both, for different reasons. A coarse check at the gateway rejects obvious abuse before it costs you anything. A finer check closer to the model server catches what the gateway can't see, like actual token counts, but by then you've already paid for some of the work.
What should happen when a customer hits their spend cap?
For a spend cap specifically, consider throttling rather than a hard rejection, with a notification to your team so a person can decide whether to raise it. Reserve hard rejection for the request-rate cap protecting your infrastructure, where letting traffic through unchecked risks a real outage.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Setting Rate Limits Without Breaking Your Best Customers
A decision guide for setting per-tier rate limits and spend caps that protect your infrastructure without throttling the customers you most want to keep.
Setting Spend Caps Before Your Agents Set Them for You
A decision guide for setting rate limits and spend caps on agentic workloads, so a stuck loop or a bad actor can't turn into an open-ended bill.
Setting Rate Limits and Spend Caps That Don't Break Real Usage
How to set rate limits and spend caps that stop abuse and runaway costs without throttling your actual customers, with a worked example.
The Rate-Limit Gaps a 30-Minute Audit Usually Finds
A short, practical checklist for finding the rate-limiting and spend-cap gaps that let one bad actor or one buggy client burn through your budget.
Setting Rate Limits That Protect Budget, Not Just Uptime
A practical checklist for designing rate limits and spend caps that stop runaway costs and abuse without breaking legitimate customer usage.
What to Do When an Upstream API Starts Rate Limiting You
A checklist for surviving upstream rate limits: reading the response headers, backing off correctly, and knowing when to buy more quota instead.