Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Setting Rate Limits Before a Bad Actor, or Your Own Cron Job, Sets Them For You

Most teams add rate limiting only after something has already gone wrong: a scraper hammering an endpoint, a customer's retry loop looping too fast, or a misconfigured internal job quietly consuming ten times its normal quota. By then it's a firefight instead of a design decision.

This is how to set limits that hold before that happens, and stay out of the way of legitimate traffic when it does.

Pick the algorithm that matches your actual traffic shape

A fixed window counter is simple and cheap but lets traffic spike right at the window boundary, twice the intended rate for a brief moment. A sliding window or token bucket smooths that out at a small added cost in complexity and state.

Most APIs don't need the most sophisticated option available. Token bucket handles bursty legitimate traffic, a user clicking rapidly, better than a strict fixed window, and it's usually the right default unless you have a specific reason to need something else.

Rate limit per tenant, not just per API key globally

A global limit protects your infrastructure but not your customers from each other. One noisy tenant, a runaway script or an aggressive integration, can eat the shared budget and degrade the experience for everyone else calling the same endpoint.

Per-tenant quotas, enforced independently, keep one customer's mistake from becoming every customer's incident. This costs more to implement than a single global counter, but it's the difference between one account having a bad day and your whole API looking unreliable.

Return enough information for a client to actually back off correctly

A bare 429 with no other information leaves the caller guessing how long to wait, which usually means they guess wrong and retry too soon. Standard rate-limit headers, remaining quota, reset time, tell a well-behaved client exactly when it's safe to try again.

This matters more than it sounds like it should: a client that retries immediately after a 429 makes the overload worse, not better, and good headers are what let a reasonable retry implementation actually do the right thing instead of guessing.

A worked example: the internal job that looked like an attack

Say a nightly reconciliation job starts calling an internal API in a tight loop after a bug removes its pagination delay, and the on-call engineer's first read of the traffic graph looks exactly like a credential-stuffing attack: thousands of requests per minute from one source.

Tracing the source back to an internal service account, not an external IP, changes the whole response, no incident response for an attack, just a bug fix and a quota on that service account so the same mistake can't repeat at that scale again. Internal traffic needs the same quotas external traffic gets, since a bug in your own code can look identical to abuse from the outside.

Where rate limiting setups have gaps

  • Limits set once at launch and never revisited as traffic patterns change
  • No distinction between a legitimate burst and sustained abuse, so real users get blocked during normal spikes
  • Internal service-to-service calls exempted from limits entirely, until one of them misbehaves
  • No dashboard showing who's actually near their limit, so problems surface as support tickets instead of a graph

Treat your limits as a living configuration, not a launch-day setting

Traffic patterns change as your product grows, and a limit that made sense for your first hundred customers can be either too tight or dangerously loose for your next thousand. Review actual usage against configured limits on a regular cadence, quarterly is reasonable for most APIs, and adjust before a limit becomes either a support burden or a real vulnerability.

This is a small recurring task, closer to a cost review than a security audit, but skipping it for a year is exactly how limits end up disconnected from the traffic they're supposed to govern.

Give large, legitimate customers a documented path to more headroom

A hard limit with no escalation path pushes your biggest, most valuable customers to either work around it awkwardly or quietly look at a competitor with more generous defaults. A documented process, request a higher tier, get reviewed, get approved, keeps that growth from turning into churn.

This doesn't mean raising limits on request with no scrutiny. It means having an actual process, rather than an ad hoc Slack message to an engineer who happens to know how to change the config, so the exception is tracked and reviewable instead of invisible.

Executive Capability Standard

What Good Looks Like

Rate limiting that holds under real traffic uses an algorithm matched to your traffic shape, enforces quotas per tenant rather than globally, and gives callers enough information in the response to back off correctly.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull your actual traffic distribution for your busiest endpoints and find your heaviest legitimate callers' normal usage pattern.
2. Do Manually:Set an initial limit by hand based on that data and watch it against real traffic for a few weeks before treating it as final.
3. Delegate:Give one team ownership of rate-limit configuration across the API, so limits aren't set ad hoc by whoever built each endpoint.
4. Automate:Move from a fixed window to a token bucket or sliding window implementation with proper rate-limit response headers.
5. Buy:Bring in an API gateway or platform engineering specialist once you're managing per-tenant quotas across dozens of endpoints by hand.

How to Get Started

Frequently Asked Questions

What's a reasonable default rate limit for a new API endpoint?

There's no universal number, it depends on what the endpoint does and what a legitimate caller's normal usage pattern looks like. Start by measuring actual usage from your heaviest real customers, then set the limit at a comfortable multiple above that, not below it.

Should rate limits differ by pricing tier?

Usually yes, and it's a natural way to make the limit legible to customers: a higher tier gets a higher quota, tied to what they're paying for rather than an arbitrary technical number they have no context for.

How do we rate limit without breaking legitimate burst traffic, like a bulk import?

Token bucket algorithms handle this well by design, allowing a burst up to the bucket's capacity before throttling kicks in. For genuinely large bulk operations, a separate batch endpoint with its own higher limit is often cleaner than trying to make one endpoint serve both patterns.

Should rate limits be documented publicly in the API docs?

Yes, specific and current. A vague 'reasonable use' policy leaves integrators guessing and building brittle retry logic around assumptions that may be wrong, while a documented number lets them design correctly against it from the start.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides