Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

Picking the Right Fix When You Hit an API's Rate Limit

When an upstream API starts returning 429s, the instinct is to add a retry loop and move on. That works for a traffic spike that passes in a minute. It doesn't work when the real problem is that your steady-state call volume has grown past what the vendor's plan allows, in which case retries just turn a clear error into a slow, silent backlog.

The right fix depends on which kind of limit you're hitting and how your load is shaped. This guide walks through diagnosing that, then through four common fixes, so you can pick the cheapest one that actually solves your case instead of reaching for the biggest hammer first.

How do you diagnose which API rate limit you're hitting?

Most vendors expose more than one limit, and the fix that works for one makes another worse. A requests-per-second cap responds well to backoff and queuing. A daily or monthly quota doesn't care how you pace requests, only how many you send in total, so smoothing traffic just delays hitting the same wall. A concurrent-connection limit is different again: you can be well under your rate cap and still get throttled because too many calls are open at once, usually from a connection pool that isn't closing cleanly.

Check the response headers before building anything. Most APIs that rate limit at all return a remaining-quota count and a reset time, either on the 429 itself or on every response. Log those for a day before deciding which fix to build; guessing wrong means solving a problem you don't have.

Exponential Backoff: Right for Bursts, Wrong for Sustained Load

Backoff with jitter, waiting a random amount of time that grows with each retry, is the correct first response to a rate-per-second limit hit during a traffic burst. It buys the upstream service time to recover and spreads your retries so they don't all land in the same second and trigger the limit again.

It stops being the right answer once your baseline call volume is the problem rather than a spike. If requests are getting throttled throughout normal business hours, backoff just adds latency to every user-facing action without ever reducing total volume. That's the signal to move to smoothing or caching instead of tuning the retry curve further.

Queuing and Rate Smoothing for Steady, Too-High Load

When the issue is a flat rate that's simply above the ceiling, put a queue in front of the calls and drain it at the rate the API allows, using a token bucket or leaky bucket limiter rather than a fixed sleep between calls. A token bucket handles bursts within your allowance gracefully while still capping the long-run average, which a naive fixed-delay loop can't do.

This adds latency for the caller, so it only works for requests that can tolerate a delay: background sync jobs, webhook processing, batch enrichment. For anything on a user-facing request path, queuing usually means you've picked the wrong architecture and should look at caching or reducing call volume instead.

Does sharding across API keys help, and when does it backfire?

Splitting traffic across multiple API keys or accounts raises your effective limit, and some vendors support it explicitly with team or organization plans built for exactly this. It backfires when the vendor's terms treat multiple keys from one customer as a violation, which some do specifically to prevent this workaround, or when the underlying limit is tied to your account or IP rather than the key itself.

Read the vendor's rate-limit documentation for language about fair use or prohibited key sharing before you build this. If it's not explicitly allowed, ask your account contact rather than finding out during an incident that access has been suspended.

Caching to Avoid Needing the Call at All

The cheapest fix is the one that removes the call entirely. Good caching candidates share a few traits:

  • Data that changes slower than you're currently polling it, like a vendor's product catalog or a partner's rate card.
  • Responses that are identical across many of your users, such as a shared reference lookup rather than per-user account data.
  • Calls made speculatively, just in case, rather than because a user action needs the answer right now.
  • Idempotent GET requests, which are safe to cache without worrying about side effects, unlike a POST that changes state.

Set a time-to-live that matches how often the underlying data actually changes, not how often you're currently calling it. Most teams find that alone cuts call volume enough that the rate limit stops being an active problem.

Executive Capability Standard

What Good Looks Like

Good rate-limit handling means your team knows which specific limit it's hitting before choosing backoff, queuing, sharding, or caching, instead of guessing.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Log the rate-limit and quota-remaining headers on every call to the upstream API for a few days to see the actual shape of your traffic against its limit.
2. Do Manually:Add exponential backoff with jitter to the calls that are failing, and confirm it actually clears the errors before building anything more complex.
3. Delegate:Have one engineer own the rate-limit strategy for each major upstream integration, including what happens when the vendor changes the limit.
4. Automate:Put a token bucket limiter in front of any integration with steady, high-volume calls, and cache anything that doesn't need to be fetched live.
5. Buy:If a core product feature depends on a vendor API near its limit, talk to the vendor about a higher tier built for your volume rather than engineering around a limit meant for smaller usage.

How to Get Started

Frequently Asked Questions

Is exponential backoff enough on its own?

Only if you're hitting a per-second or per-minute limit during short bursts. If throttling happens continuously during normal business hours, backoff just adds latency without reducing total call volume, and you need queuing, caching, or a higher plan tier instead.

Is it against the rules to use multiple API keys for a higher limit?

It depends on the vendor. Some explicitly offer higher tiers built for exactly this, others treat multiple keys from one customer as a terms of service violation. Check the documentation for language about fair use before building around it, and ask your account contact if it's unclear.

What's the fastest way to find out which limit we're actually hitting?

Log the rate-limit headers most APIs return on every response, not just on the 429, for about a day. That tells you whether you're near a per-second cap, a daily quota, or a concurrent-connection limit, and each needs a different fix.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides