The Rate-Limit Gaps a 30-Minute Audit Usually Finds
A short rate-limit audit usually finds gaps beyond the public API edge, such as uncapped internal calls to paid APIs, open webhook endpoints, and background jobs with no concurrency ceiling. Rate limiting is typically built to protect a public API from abuse, and these other gaps rarely surface until a spend spike or an outage.
Here's a short audit that surfaces the most common ones.
Which outbound calls need a limit? Anything you pay for by usage
Start with anything metered: a third-party API billed per call, a cloud service billed per request or per compute-second, an AI provider billed per token. For each one, confirm there's a hard ceiling somewhere in the call path, not just a soft expectation of "normal" traffic. A retry loop with no cap, triggered by a transient error, can turn a brief outage on the other end into a large and completely avoidable bill on your end.
Give each metered dependency an owner responsible for knowing its pricing model and its current usage trend, not just the engineer who happened to integrate it first. Pricing models change, a provider moves from flat-rate to usage-based, or a free tier shrinks, and a dependency with no clear owner is the one most likely to still be running under last year's outdated assumptions about cost.
Check whether internal services trust each other's rate limits
Rate limiting at your public API edge protects you from external abuse, but it does nothing for a bug in an internal service that calls another internal service in a tight loop. Internal-to-internal calls are exactly where rate limiting gets skipped, since it feels unnecessary between systems you control. A retry-without-backoff bug in one service can still take down another internal service or burn through a shared quota just as effectively as an external attacker, and it's often harder to spot because nobody's watching internal traffic the way they watch the public edge, so it can run unnoticed for hours before anyone connects the symptom to the cause.
For example, an internal reporting service starts retrying a failed call to a paid enrichment API without backoff. The public edge shows normal traffic, so nothing alerts, yet the metered bill climbs for hours. A per-service ceiling on that outbound call, plus an alert when usage departs from its usual pattern, would have stopped it early. A useful decision rule: if a service can spend money or exhaust a shared quota, it needs its own limit, even when it only talks to systems you control.
Where should rate limits be enforced? Close to where the cost happens
A rate limit enforced at your load balancer catches abusive traffic before it reaches your application, but if the expensive resource is a downstream API call or a database query several layers deeper, a limiter that only sits at the edge won't catch a legitimate-looking request that triggers an expensive operation internally. Put a limit as close as practical to whatever actually costs money or capacity: the specific database query, the specific external API call, not just the front door.
Set per-tenant limits, not just global ones
A single global rate limit protects the system as a whole but doesn't protect one customer's experience from another's. If your product has multiple tenants or customers, a limit scoped per tenant (or per API key, per user) prevents one noisy or misbehaving customer from degrading service for everyone else. This matters most for shared resources like a database connection pool or a rate-limited third-party API where usage isn't naturally isolated by tenant already.
- Set a global ceiling to protect the system as a whole
- Set per-tenant or per-key ceilings so one customer can't exhaust shared capacity
- Alert when any tenant approaches its own ceiling, not just when the global one is hit
- Review actual usage quarterly and adjust ceilings that were guessed rather than measured
Per-tenant limits also protect you internally, not just customers from each other. A single large customer's usage pattern, a bulk import running overnight, say, can otherwise look identical to an attack from the system's point of view, and a global-only limiter has no way to tell the difference between "our biggest customer is having a busy day" and "something is wrong," while a per-tenant limit lets both be true without one disrupting the other.
What Good Looks Like
Every metered external call has a hard ceiling enforced close to where the cost happens, internal service-to-service calls aren't exempt from limits, and per-tenant ceilings exist wherever a shared resource could otherwise be exhausted by one customer.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the fastest way to find our biggest rate-limiting gap?
Pull your last three months of cloud and third-party API bills and look for spikes with no corresponding traffic increase. That mismatch almost always points at a retry loop, a missing cap, or a single tenant's usage that went unnoticed.
Should rate limits ever be soft, just logging instead of blocking?
For anything genuinely metered or costly, no; a soft limit that only logs doesn't stop the spend or the outage, it just gives you evidence after the fact. Reserve soft, warning-only limits for lower-stakes internal calls where you want visibility before deciding whether a hard cap is warranted.
How do we set a sensible limit without historical usage data to base it on?
Start conservative, based on what a reasonable single user or tenant would realistically need, and monitor closely in the first weeks after launch. It's much easier to raise a limit that's turning out to be too tight than to explain an outage or bill spike caused by one that was too loose.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Setting Rate Limits Without Breaking Your Best Customers
A decision guide for setting per-tier rate limits and spend caps that protect your infrastructure without throttling the customers you most want to keep.
Setting Rate Limits and Spend Caps That Don't Break Real Usage
How to set rate limits and spend caps that stop abuse and runaway costs without throttling your actual customers, with a worked example.
Setting Rate Limits That Protect Budget, Not Just Uptime
A practical checklist for designing rate limits and spend caps that stop runaway costs and abuse without breaking legitimate customer usage.
Setting Spend Caps on a RAG Pipeline Without Breaking It
Per-tenant quotas, graceful degradation instead of hard rejection, and separating ingestion from query traffic: a practical guide to RAG rate limits.
Sizing API Rate Limits So They Actually Protect You
A worked example for setting rate limits and spend caps on your APIs so they catch real abuse without throttling your legitimate customers.
Setting Rate Limits and Spend Caps Without Breaking Real Users
How to set request-rate limits and spend caps for AI model serving separately, with a worksheet for setting your first cap.