API Security, Identity & Zero-TrustPlaybook3 min readUpdated September 2026

Sizing API Rate Limits So They Actually Protect You

A rate limit set by guessing is either too loose to stop abuse or too tight to let real usage through, and most teams only find out which once something goes wrong. Sizing it properly takes a few concrete inputs, not a gut feeling.

Walk through this with your own numbers rather than copying a default from a framework's documentation.

How do you size a rate limit from your heaviest legitimate user?

Say your typical customer makes 200 API calls a day, but your single heaviest customer, running a legitimate batch sync job overnight, makes 8,000 in a ten-minute window. A limit set from the average would throttle that customer constantly. Pull your actual top-percentile usage from logs, not from what you assume "normal" looks like, and set the limit above that, with headroom for growth.

This single step catches the most common rate-limiting mistake: building the limit around a mental model of typical usage instead of the real distribution, which is almost always far more skewed than people expect.

Pull at least 30 days of logs before setting anything, since a shorter window can hide a legitimate monthly or quarterly batch job that only runs once. A limit tuned on two weeks of data will throttle that job the first time it actually runs, and the resulting support ticket is a worse way to discover the gap than checking the data upfront.

Should rate limits apply per identity or per IP address?

An IP-based limit punishes every user behind a shared corporate NAT or VPN for one bad actor's behavior, and it does nothing to stop an attacker who simply rotates IPs. Tie limits to the authenticated caller's identity instead, whether that's an API key, an OAuth client, or a user session, since that's the identity zero trust already requires you to establish before the request is processed anyway.

Keep a lightweight IP-based limit as a backstop against unauthenticated abuse, but treat identity-based limiting as the primary mechanism, not the fallback.

Give different endpoints different limits based on what they cost

A read of a cached, lightweight resource and a call that triggers an expensive downstream AI model or a heavy database query shouldn't share the same quota. Set limits per endpoint category based on actual cost to serve, not a single blanket number applied everywhere, or you'll end up either overprotecting cheap endpoints or underprotecting expensive ones.

This matters most for anything that calls out to a metered third-party service, where a burst of legitimate-looking traffic can generate a real, unexpected bill before anyone notices a spike.

Build a spend cap as a second, independent layer above rate limits

A rate limit caps request volume; a spend cap caps dollars, and they're not the same protection. A caller making requests within their rate limit can still generate a large bill if each request is expensive enough, especially for endpoints backed by usage-priced infrastructure. Set an explicit dollar ceiling, checked independently of the request-count limit, so a cost spike gets caught even when the volume looks unremarkable.

Alert well before the cap is hit, not only when it's reached, so someone can investigate whether the spend is legitimate growth or something worth stopping. For example, if an endpoint proxies calls to a metered third-party model and a caller's usage jumps from $40 a day to $400 a day, that's worth a look before the monthly bill arrives, not after.

Decide what happens at the limit before you hit it in production

A caller that exceeds its limit should get a clear, documented response, a specific status code and a retry-after header, not a generic error that looks like an outage. Decide this behavior and test it deliberately, including for your own internal callers, so a burst of legitimate traffic degrades predictably instead of producing a confusing failure that looks like a bug in the API itself.

Review your limits again after any customer's usage roughly doubles, the same way you'd revisit any other capacity assumption, rather than waiting for a complaint to prompt the review.

Size and roll out a limit in this order:

  1. Pull a long enough window of real usage logs to include monthly or quarterly batch jobs, and find your heaviest legitimate callers.
  2. Set the limit above that top-percentile usage with headroom for growth, instead of basing it on an average customer.
  3. Tie the limit to the authenticated identity, such as an API key, OAuth client or user session, rather than an IP address.
  4. Give expensive endpoints their own lower quotas based on the cost to serve them.
  5. Add a separate dollar spend cap, then decide and test what callers see at the limit, including a clear status code and retry-after header.
Executive Capability Standard

What Good Looks Like

Good practice means rate limits are set from actual top-percentile usage data per identity and per endpoint, and a spend cap exists as an independent, second layer of protection above the request-count limit.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull your actual usage distribution from logs, including your heaviest legitimate users, before setting or revising any limit.
2. Do Manually:Set limits by hand for your highest-traffic endpoints based on that data, then monitor for a week to see how often real customers approach them.
3. Delegate:Assign an engineer to own rate limit and spend cap configuration, with a standing task to review it whenever a customer's usage changes significantly.
4. Automate:Build automated alerting on approach-to-limit and approach-to-spend-cap, so you find out before a customer complains or a bill surprises you.
5. Buy:Use your API gateway's built-in rate limiting and cost-tracking features rather than building custom throttling logic, since this is a well-solved problem most gateways already handle.

How to Get Started

Frequently Asked Questions

Should rate limits be the same across every customer?

Not necessarily. A tiered approach, where limits scale with plan level or a documented higher-usage agreement, is common and reasonable, as long as the underlying method for setting each tier's limit, based on real usage data, stays consistent.

What's the difference between a rate limit and a spend cap?

A rate limit caps how many requests a caller can make in a window; a spend cap caps how much those requests can cost in dollars. They need to be enforced independently, because a caller within their rate limit can still generate an unexpectedly large bill if individual requests are expensive.

How do we know if our current rate limits are too tight?

Look for support tickets about unexpected throttling from legitimate customers, and compare your limit against actual top-percentile usage in your logs rather than against the average. A limit that's frequently hit by real customers doing normal things is a sign it was set too close to average usage instead of above it.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides