Setting Rate Limits Before a Bad Actor, or Your Own Cron Job, Sets Them For You
Most teams add rate limiting only after something has already gone wrong: a scraper hammering an endpoint, a customer's retry loop looping too fast, or a misconfigured internal job quietly consuming ten times its normal quota. By then it's a firefight instead of a design decision.
This is how to set limits that hold before that happens, and stay out of the way of legitimate traffic when it does.
Pick the algorithm that matches your actual traffic shape
A fixed window counter is simple and cheap but lets traffic spike right at the window boundary, twice the intended rate for a brief moment. A sliding window or token bucket smooths that out at a small added cost in complexity and state.
Most APIs don't need the most sophisticated option available. Token bucket handles bursty legitimate traffic, a user clicking rapidly, better than a strict fixed window, and it's usually the right default unless you have a specific reason to need something else.
Rate limit per tenant, not just per API key globally
A global limit protects your infrastructure but not your customers from each other. One noisy tenant, a runaway script or an aggressive integration, can eat the shared budget and degrade the experience for everyone else calling the same endpoint.
Per-tenant quotas, enforced independently, keep one customer's mistake from becoming every customer's incident. This costs more to implement than a single global counter, but it's the difference between one account having a bad day and your whole API looking unreliable.
Return enough information for a client to actually back off correctly
A bare 429 with no other information leaves the caller guessing how long to wait, which usually means they guess wrong and retry too soon. Standard rate-limit headers, remaining quota, reset time, tell a well-behaved client exactly when it's safe to try again.
This matters more than it sounds like it should: a client that retries immediately after a 429 makes the overload worse, not better, and good headers are what let a reasonable retry implementation actually do the right thing instead of guessing.
A worked example: the internal job that looked like an attack
Say a nightly reconciliation job starts calling an internal API in a tight loop after a bug removes its pagination delay, and the on-call engineer's first read of the traffic graph looks exactly like a credential-stuffing attack: thousands of requests per minute from one source.
Tracing the source back to an internal service account, not an external IP, changes the whole response, no incident response for an attack, just a bug fix and a quota on that service account so the same mistake can't repeat at that scale again. Internal traffic needs the same quotas external traffic gets, since a bug in your own code can look identical to abuse from the outside.
Where rate limiting setups have gaps
- Limits set once at launch and never revisited as traffic patterns change
- No distinction between a legitimate burst and sustained abuse, so real users get blocked during normal spikes
- Internal service-to-service calls exempted from limits entirely, until one of them misbehaves
- No dashboard showing who's actually near their limit, so problems surface as support tickets instead of a graph
Treat your limits as a living configuration, not a launch-day setting
Traffic patterns change as your product grows, and a limit that made sense for your first hundred customers can be either too tight or dangerously loose for your next thousand. Review actual usage against configured limits on a regular cadence, quarterly is reasonable for most APIs, and adjust before a limit becomes either a support burden or a real vulnerability.
This is a small recurring task, closer to a cost review than a security audit, but skipping it for a year is exactly how limits end up disconnected from the traffic they're supposed to govern.
Give large, legitimate customers a documented path to more headroom
A hard limit with no escalation path pushes your biggest, most valuable customers to either work around it awkwardly or quietly look at a competitor with more generous defaults. A documented process, request a higher tier, get reviewed, get approved, keeps that growth from turning into churn.
This doesn't mean raising limits on request with no scrutiny. It means having an actual process, rather than an ad hoc Slack message to an engineer who happens to know how to change the config, so the exception is tracked and reviewable instead of invisible.
What Good Looks Like
Rate limiting that holds under real traffic uses an algorithm matched to your traffic shape, enforces quotas per tenant rather than globally, and gives callers enough information in the response to back off correctly.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's a reasonable default rate limit for a new API endpoint?
There's no universal number, it depends on what the endpoint does and what a legitimate caller's normal usage pattern looks like. Start by measuring actual usage from your heaviest real customers, then set the limit at a comfortable multiple above that, not below it.
Should rate limits differ by pricing tier?
Usually yes, and it's a natural way to make the limit legible to customers: a higher tier gets a higher quota, tied to what they're paying for rather than an arbitrary technical number they have no context for.
How do we rate limit without breaking legitimate burst traffic, like a bulk import?
Token bucket algorithms handle this well by design, allowing a burst up to the bucket's capacity before throttling kicks in. For genuinely large bulk operations, a separate batch endpoint with its own higher limit is often cleaner than trying to make one endpoint serve both patterns.
Should rate limits be documented publicly in the API docs?
Yes, specific and current. A vague 'reasonable use' policy leaves integrators guessing and building brittle retry logic around assumptions that may be wrong, while a documented number lets them design correctly against it from the start.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Setting Rate Limits Without Breaking Your Best Customers
A decision guide for setting per-tier rate limits and spend caps that protect your infrastructure without throttling the customers you most want to keep.
The Rate-Limit Gaps a 30-Minute Audit Usually Finds
A short, practical checklist for finding the rate-limiting and spend-cap gaps that let one bad actor or one buggy client burn through your budget.
Setting Rate Limits and Spend Caps That Don't Break Real Usage
How to set rate limits and spend caps that stop abuse and runaway costs without throttling your actual customers, with a worked example.
Setting Rate Limits That Protect Budget, Not Just Uptime
A practical checklist for designing rate limits and spend caps that stop runaway costs and abuse without breaking legitimate customer usage.
Setting Spend Caps Before Your Agents Set Them for You
A decision guide for setting rate limits and spend caps on agentic workloads, so a stuck loop or a bad actor can't turn into an open-ended bill.
Setting Spend Caps on a RAG Pipeline Without Breaking It
Per-tenant quotas, graceful degradation instead of hard rejection, and separating ingestion from query traffic: a practical guide to RAG rate limits.