API managementPlaybook3 min readUpdated September 2026

Rate Limiting an API: Limits, Headers and 429 Errors

The core rules for rate limiting an API are to limit by authenticated identity rather than IP alone, use a token bucket or sliding window, and return HTTP 429 with a Retry-After header. Publish the limits so clients can plan around them.

Rate limiting protects three things: your availability, your bill, and your other customers. It isn't mainly about stopping attackers. Most limit breaches come from a well-meaning client with a retry loop, which is why the design of the error response matters as much as the number.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Which rate limiting algorithm should you pick?

Four approaches cover almost every case. The right one depends on how bursty your legitimate clients are:

  • Fixed window: count requests per minute or hour, reset on the boundary. Simple to build, but a client can send a full quota at the end of one window and again at the start of the next.
  • Sliding window: counts requests over the trailing period, which removes the boundary spike. Slightly more storage per key.
  • Token bucket: tokens refill at a steady rate and each request spends one. It allows short bursts up to the bucket size, which fits most real client behavior.
  • Concurrency limit: caps in-flight requests instead of rate. Best for expensive endpoints such as report generation or search.

If you're unsure, start with a token bucket per API key. It's forgiving to bursty clients and easy to explain in documentation.

What should you rate limit by?

IP address alone is a poor key. Many customers sit behind one corporate NAT, and one attacker can rotate addresses cheaply. Layer your keys:

  1. Authenticated identity first: API key, OAuth client or user ID.
  2. Tenant or organization, so one noisy user can't consume a whole customer's allowance unfairly, and one customer can't starve others.
  3. IP address as a coarse backstop for unauthenticated routes such as login and password reset.
  4. Endpoint cost, so a search that fans out to the database counts for more than a health check.

Weighted costs are underused. Say a list endpoint costs 1 unit and a bulk export costs 20. One quota then covers both without letting exports crowd out everything else.

How to choose your first limits

Don't guess numbers from a blog post. Measure, then enforce in stages.

  1. Log request counts per key per minute for two weeks, with no blocking.
  2. Find the busiest legitimate keys and set your default limit comfortably above them.
  3. Turn on enforcement in log-only mode: record what would have been blocked and read the list.
  4. Contact the keys that would have been blocked. If they're legitimate, raise their tier; if they're bugs, tell the owners.
  5. Enforce, and keep a per-key override so support can lift a limit without a deploy.

Set different tiers for free, paid and internal callers from the start. Retrofitting tiers later is harder than adding a column now.

Say your busiest legitimate key sends 40 requests a minute at peak, and a typical key sends 5. A default of 100 a minute leaves generous headroom without letting one runaway script flood you. When a customer asks for more, raise their limit after asking what they're building, since a real need for higher throughput is often a sign to offer a batch endpoint instead.

What should a 429 response include?

A good rejection tells the client what happened and when to retry. Return status 429, a Retry-After header in seconds, and a JSON body with a stable error code, not just prose. Many teams also expose remaining quota and reset time in response headers on every call, so well-behaved clients slow down before they're rejected.

Document the recommended client behavior: exponential backoff with random jitter, and a cap on retries. Without jitter, a hundred clients that were throttled together all retry together. Never return a 500 or a 503 for a limit breach, because clients and monitoring will treat it as your outage.

Where should enforcement live?

Enforce at more than one layer. The edge, whether a CDN or a gateway product like Kong, absorbs floods before they reach your servers and applies coarse limits by IP or key. The application layer knows about tenants, plans and endpoint cost, so it applies the precise rules.

In a multi-instance setup, the counters need shared storage, usually a fast in-memory store. Decide what happens when that store is down. Failing open keeps your API up but unprotected; failing closed protects the backend but causes an outage. Most teams fail open with a conservative local fallback limit, and alert loudly.

Executive Capability Standard

What Good Looks Like

Every public endpoint has a documented limit per authenticated identity, returns 429 with Retry-After, and has been tested against a client in a retry loop.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand fixed window, sliding window, token bucket and concurrency limits, and which fits your traffic shape.
2. Do Manually:Log per-key request rates for two weeks and write the first limits into a short internal policy.
3. Delegate:Have one backend engineer own the limiter, the override process and the documentation.
4. Automate:Enforce limits in the gateway or middleware, expose quota headers and alert when a key hits its ceiling repeatedly.
5. Buy:Use a managed gateway or edge service for coarse limits once your own limiter becomes a maintenance burden.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Kong

Fits when you run several services and want consistent per-key limits enforced at one gateway layer.

Visit Kong→
Cloudflare

Fits when you want to absorb floods and apply coarse IP or path limits at the edge before traffic reaches your servers.

Visit Cloudflare→

Frequently Asked Questions

What HTTP status code should rate limiting return?

Return 429 Too Many Requests with a Retry-After header. That is the standard signal clients and libraries understand. Avoid 500 or 503, which look like server failures and trigger the wrong alerts and retry behavior.

Should you rate limit by IP address or API key?

Prefer the API key or user identity, and use IP only as a backstop. Shared corporate networks put many legitimate users behind one IP, while attackers can rotate IPs easily. Keying on identity is fairer and harder to evade.

How do you pick the right rate limit number?

Measure real usage per key for a couple of weeks, then set the default above your busiest legitimate callers. Run in log-only mode before enforcing so you can see who would be blocked and why.

Do internal services need rate limits too?

Yes, at a looser level. A buggy internal caller in a retry loop can take down a shared dependency just as easily as an outside client. Concurrency limits and timeouts are often more useful there than request quotas.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides