What to Do When an Upstream API Starts Rate Limiting You
Every service you don't own eventually says no. A payment processor, a mapping API, a model provider, all of them cap how fast you can call them, and the cap usually gets hit for the first time in production, at the worst possible moment.
The difference between an outage and a non-event is almost entirely in how your code reacts to that first 429, not in whether you saw it coming. None of the fixes below require the provider's cooperation, which matters, because most vendors are slow to raise a limit and fast to enforce one.
How Should You Read a 429 Response Before Retrying?
Most providers tell you exactly what to do next in the response itself: a Retry-After header, or rate limit remaining and reset headers. Retrying blind, on a fixed delay, ignores information the provider is handing you for free. Parse the header, wait at least that long, and you'll clear most rate limit incidents without a human ever getting paged.
Log the headers too, not just the retry outcome. A provider that quietly lowers your limit, or changes how it counts requests, shows up in those numbers weeks before anyone notices the errors.
How Do You Back Off Exponentially and Cap the Wait?
For the providers that don't tell you when to come back, exponential backoff with jitter is the standard answer: double the wait after each failure, add a small random offset so a fleet of retrying clients doesn't resynchronize into a new wave of requests, and cap the maximum wait so a single stuck job doesn't silently retry for hours.
Set a maximum number of attempts too, and decide up front what happens when that ceiling is hit: does the request go to a dead letter queue for a person to look at, or does the feature just fail closed for that user. Deciding this during an incident is worse than deciding it now.
Cache or batch what you can
The cheapest fix for a rate limit is calling the API less. Cache responses that don't change often, batch requests the provider's API supports batching for, and de-duplicate calls that fire twice because of a retry somewhere upstream in your own system. Say your checkout flow calls a tax API on every cart update instead of once at checkout, that's demand you can cut before you ever talk to the vendor about a bigger plan.
This is usually where the biggest win hides. Teams tend to assume their call volume is fixed and go straight to asking for more quota, when a surprising share of that volume turns out to be avoidable once someone actually traces where the calls are coming from.
Buy headroom before you buy a bigger plan
A dedicated rate limiter and queue in front of the outbound call gives you a place to smooth bursts before they ever reach the provider. That's often cheaper and faster to ship than a support ticket asking for a higher tier, and it protects you against the next provider's limit too, not just this one.
It also gives you a single place to add priority: a paying customer's request can jump ahead of a background sync job when the queue is under pressure, something a raw retry loop can't do for you.
Common mistakes that make a quota problem worse
Retrying immediately on every failure turns a brief limit into a sustained one, because your retries become part of the load that's keeping you rate limited. Failing to distinguish a rate limit from a real error means you might drop a request you should have retried, or hammer an endpoint that's actually down for another reason.
Treating the limit as an emergency to escalate every time, instead of building the backoff logic once, means the same incident repeats indefinitely. And sharing one API key across every environment, including staging and a developer's laptop, means a load test can burn through the quota your production traffic needed.
Avoid these patterns, which turn a brief limit into a sustained one:
- Retrying immediately on every failure, so your retries become part of the load that keeps you rate limited.
- Treating a rate limit like any other error, which can drop a request you should have retried or hammer an endpoint that is down.
- Retrying blind on a fixed delay while ignoring the Retry-After information the provider already sent.
- Handling each incident as a one-off instead of building the backoff and caching logic once.
What Good Looks Like
Good rate limit handling means a provider's 429 response degrades your feature gracefully instead of paging anyone or dropping data.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should we always retry a 429 response?
Only with backoff, and only up to a sane limit. Retrying immediately just adds to the load causing the limit in the first place. Respect any Retry-After header the provider sends, and give up gracefully after a handful of attempts rather than retrying forever.
How do we know if we need a bigger API plan or just better code?
Look at your actual call pattern first. If most of your volume is redundant or cacheable, fix that before paying for more headroom. If your traffic is genuinely growing and there's no waste left to cut, that's when a higher tier or a partner conversation makes sense.
What's the biggest mistake teams make with third-party rate limits?
Treating each incident as a one-off instead of building the backoff and caching logic once. The same provider will rate limit you again, and the fix should already be in place the second time it happens.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
How to Version an API Your Model-Serving Clients Depend On
How to design and version an AI model-serving API contract so a model swap never silently breaks a client, including streaming and deprecation windows.
How Much Latency Your Gateway Adds to an Inference Call
A simple way to measure how much delay your API gateway adds on top of raw inference time, and what to check before blaming the model for a slow response.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
Setting a Scanning Cadence for Your Model-Serving Stack
A scanning cadence for AI model serving covering the inference server, container images, GPU drivers, and remediation timelines.