AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

Setting Rate Limits and Spend Caps Without Breaking Real Users

Rate limiting for a model-serving endpoint has two different jobs that get lumped into one setting way too often: protecting your infrastructure from being overwhelmed, and protecting your budget from a runaway integration or an abusive caller. Those two jobs need different limits, checked at different points.

A single global rate limit tends to under-protect one of the two goals no matter how you tune it.

Two different caps, not one

  • A request-rate cap, protecting infrastructure, set close to what your GPUs can actually handle before latency degrades for everyone.
  • A spend cap, protecting budget, set per caller or per feature based on expected usage, independent of whether the infrastructure could technically handle more.

A caller can be well under the request-rate cap and still blow through a spend cap if their requests are unusually large, long prompts, high output token counts, so track both separately rather than assuming one implies the other.

Where to check the limit: gateway versus model server

Checking the limit at your API gateway, before a request reaches the model server, is cheaper and catches abuse earlier, but it needs to know enough about the request, expected token count, which model it's headed for, to make a meaningful decision.

Checking it at the model server catches everything the gateway might miss, but by then you've already paid for whatever work happened before the check. The practical answer is both: a coarse check at the gateway to reject obvious abuse fast, and a finer per-model, per-token check closer to where the actual cost is incurred.

A worksheet for setting your first spend cap

Work through four numbers for each caller type: typical requests per day, typical tokens per request, your cost per thousand tokens, and the multiple above typical usage you're comfortable allowing before you want a human to look at it, often somewhere between three and five times normal.

Multiply the first three together for a typical daily cost, then apply your comfort multiple for the actual cap. Review actual usage against this cap monthly and adjust it; a cap set once from a guess and never revisited either blocks legitimate growth or fails to catch a real problem.

What happens when a caller hits the limit

A hard rejection is the simplest response and the easiest to get wrong for a legitimate customer having an unusually busy day. Consider a soft response instead for the spend cap specifically: throttle rather than reject outright, and notify someone on your team so a human can decide whether to raise the cap.

Reserve hard rejection for the request-rate cap protecting infrastructure, where the risk of letting traffic through unchecked is a real outage rather than an unexpected bill you can review after the fact and, if needed, follow up on with the customer directly.

A useful decision rule is to ask what the worst outcome is if the request goes through. If the worst outcome is an outage for every other customer, reject it. If the worst outcome is a larger bill for one caller, throttle and notify. For example, a customer who starts a large batch import can be slowed to a lower rate with a message explaining why, while a runaway loop sending the same request over and over should be blocked at the gateway. Writing that rule down means the on-call engineer is not inventing policy in the middle of an incident.

Communicating limits to customers before they hit them

A caller who hits an undocumented limit assumes something is broken, not that they've been capped. Publish your rate limits and, where practical, your spend-cap defaults, so integrators can design around them instead of discovering them through failed requests.

For spend caps specifically, consider a usage dashboard or a proactive notification well before the cap, at a comfortable majority of it, so a customer scaling up gets a heads-up instead of a surprise rejection at the worst possible moment for whatever they're building on top of you.

Rate-limiting mistakes that hit real customers

  • One global limit applied to every caller regardless of their actual usage pattern or contract.
  • No monitoring on how close callers run to their limit, so you find out about a problem only after they've already been blocked.
  • Treating a spend cap breach the same as a security incident, when it's usually just a customer using the product more than expected.

The goal is protecting the business, not punishing growth. A customer who outgrows a cap because their own usage is genuinely growing is a sales conversation, not a security event.

Executive Capability Standard

What Good Looks Like

A working rate-limiting setup tracks request-rate and spend caps as separate numbers, checks them at both the gateway and closer to the model, and routes a cap breach to a person rather than only to an automatic hard rejection.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Work through the spend-cap worksheet, typical requests, typical tokens, cost per thousand tokens, comfort multiple, for your highest-volume caller type.
2. Do Manually:Review actual usage against your current limits by hand for a month and note anyone running close to a cap.
3. Delegate:Assign someone to own rate-limit and spend-cap settings and review them monthly against real usage, not just at launch.
4. Automate:Build monitoring that flags callers approaching their cap before they hit it, and route spend-cap breaches to a person instead of a silent rejection.
5. Buy:Bring in infrastructure help to design gateway-level rate limiting if you're fielding enough traffic that manual tuning can't keep up.

How to Get Started

Frequently Asked Questions

Should request-rate limits and spend caps be the same setting?

No. A request-rate cap protects your infrastructure from being overwhelmed and should be based on what your GPUs can actually handle. A spend cap protects your budget and depends on request size, not just count. A caller can stay under one and still exceed the other, so track them separately.

Where should we check rate limits, at the gateway or the model server?

Both, for different reasons. A coarse check at the gateway rejects obvious abuse before it costs you anything. A finer check closer to the model server catches what the gateway can't see, like actual token counts, but by then you've already paid for some of the work.

What should happen when a customer hits their spend cap?

For a spend cap specifically, consider throttling rather than a hard rejection, with a notification to your team so a person can decide whether to raise it. Reserve hard rejection for the request-rate cap protecting your infrastructure, where letting traffic through unchecked risks a real outage.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides