Picking the Right Fix When You Hit an API's Rate Limit
When an upstream API starts returning 429s, the instinct is to add a retry loop and move on. That works for a traffic spike that passes in a minute. It doesn't work when the real problem is that your steady-state call volume has grown past what the vendor's plan allows, in which case retries just turn a clear error into a slow, silent backlog.
The right fix depends on which kind of limit you're hitting and how your load is shaped. This guide walks through diagnosing that, then through four common fixes, so you can pick the cheapest one that actually solves your case instead of reaching for the biggest hammer first.
How do you diagnose which API rate limit you're hitting?
Most vendors expose more than one limit, and the fix that works for one makes another worse. A requests-per-second cap responds well to backoff and queuing. A daily or monthly quota doesn't care how you pace requests, only how many you send in total, so smoothing traffic just delays hitting the same wall. A concurrent-connection limit is different again: you can be well under your rate cap and still get throttled because too many calls are open at once, usually from a connection pool that isn't closing cleanly.
Check the response headers before building anything. Most APIs that rate limit at all return a remaining-quota count and a reset time, either on the 429 itself or on every response. Log those for a day before deciding which fix to build; guessing wrong means solving a problem you don't have.
Exponential Backoff: Right for Bursts, Wrong for Sustained Load
Backoff with jitter, waiting a random amount of time that grows with each retry, is the correct first response to a rate-per-second limit hit during a traffic burst. It buys the upstream service time to recover and spreads your retries so they don't all land in the same second and trigger the limit again.
It stops being the right answer once your baseline call volume is the problem rather than a spike. If requests are getting throttled throughout normal business hours, backoff just adds latency to every user-facing action without ever reducing total volume. That's the signal to move to smoothing or caching instead of tuning the retry curve further.
Queuing and Rate Smoothing for Steady, Too-High Load
When the issue is a flat rate that's simply above the ceiling, put a queue in front of the calls and drain it at the rate the API allows, using a token bucket or leaky bucket limiter rather than a fixed sleep between calls. A token bucket handles bursts within your allowance gracefully while still capping the long-run average, which a naive fixed-delay loop can't do.
This adds latency for the caller, so it only works for requests that can tolerate a delay: background sync jobs, webhook processing, batch enrichment. For anything on a user-facing request path, queuing usually means you've picked the wrong architecture and should look at caching or reducing call volume instead.
Does sharding across API keys help, and when does it backfire?
Splitting traffic across multiple API keys or accounts raises your effective limit, and some vendors support it explicitly with team or organization plans built for exactly this. It backfires when the vendor's terms treat multiple keys from one customer as a violation, which some do specifically to prevent this workaround, or when the underlying limit is tied to your account or IP rather than the key itself.
Read the vendor's rate-limit documentation for language about fair use or prohibited key sharing before you build this. If it's not explicitly allowed, ask your account contact rather than finding out during an incident that access has been suspended.
Caching to Avoid Needing the Call at All
The cheapest fix is the one that removes the call entirely. Good caching candidates share a few traits:
- Data that changes slower than you're currently polling it, like a vendor's product catalog or a partner's rate card.
- Responses that are identical across many of your users, such as a shared reference lookup rather than per-user account data.
- Calls made speculatively, just in case, rather than because a user action needs the answer right now.
- Idempotent GET requests, which are safe to cache without worrying about side effects, unlike a POST that changes state.
Set a time-to-live that matches how often the underlying data actually changes, not how often you're currently calling it. Most teams find that alone cuts call volume enough that the rate limit stops being an active problem.
What Good Looks Like
Good rate-limit handling means your team knows which specific limit it's hitting before choosing backoff, queuing, sharding, or caching, instead of guessing.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is exponential backoff enough on its own?
Only if you're hitting a per-second or per-minute limit during short bursts. If throttling happens continuously during normal business hours, backoff just adds latency without reducing total call volume, and you need queuing, caching, or a higher plan tier instead.
Is it against the rules to use multiple API keys for a higher limit?
It depends on the vendor. Some explicitly offer higher tiers built for exactly this, others treat multiple keys from one customer as a terms of service violation. Check the documentation for language about fair use before building around it, and ask your account contact if it's unclear.
What's the fastest way to find out which limit we're actually hitting?
Log the rate-limit headers most APIs return on every response, not just on the 429, for about a day. That tells you whether you're near a per-second cap, a daily quota, or a concurrent-connection limit, and each needs a different fix.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
How to Actually Compare API Gateways on Latency
Why most API gateway latency comparisons are misleading, and a more honest way to benchmark the tradeoffs that actually matter for your own traffic.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Why Your Agent Loop Feels Slow, and How to Fix It
A diagnostic guide to finding where latency actually comes from in an agentic system, and which fixes help each cause instead of masking it.
Scanning Your Agent Stack for the Vulnerabilities That Matter
Benchmarking what continuous vulnerability scanning should actually cover for an agentic system, including the MCP server surface most scanners miss.
Setting API Standards So MCP Integrations Don't Break
How to compare approaches to building and standardizing MCP tools so a new integration doesn't quietly break every agent that depends on it.