AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

What to Do When an Upstream API Starts Rate Limiting You

Every service you don't own eventually says no. A payment processor, a mapping API, a model provider, all of them cap how fast you can call them, and the cap usually gets hit for the first time in production, at the worst possible moment.

The difference between an outage and a non-event is almost entirely in how your code reacts to that first 429, not in whether you saw it coming. None of the fixes below require the provider's cooperation, which matters, because most vendors are slow to raise a limit and fast to enforce one.

How Should You Read a 429 Response Before Retrying?

Most providers tell you exactly what to do next in the response itself: a Retry-After header, or rate limit remaining and reset headers. Retrying blind, on a fixed delay, ignores information the provider is handing you for free. Parse the header, wait at least that long, and you'll clear most rate limit incidents without a human ever getting paged.

Log the headers too, not just the retry outcome. A provider that quietly lowers your limit, or changes how it counts requests, shows up in those numbers weeks before anyone notices the errors.

How Do You Back Off Exponentially and Cap the Wait?

For the providers that don't tell you when to come back, exponential backoff with jitter is the standard answer: double the wait after each failure, add a small random offset so a fleet of retrying clients doesn't resynchronize into a new wave of requests, and cap the maximum wait so a single stuck job doesn't silently retry for hours.

Set a maximum number of attempts too, and decide up front what happens when that ceiling is hit: does the request go to a dead letter queue for a person to look at, or does the feature just fail closed for that user. Deciding this during an incident is worse than deciding it now.

Cache or batch what you can

The cheapest fix for a rate limit is calling the API less. Cache responses that don't change often, batch requests the provider's API supports batching for, and de-duplicate calls that fire twice because of a retry somewhere upstream in your own system. Say your checkout flow calls a tax API on every cart update instead of once at checkout, that's demand you can cut before you ever talk to the vendor about a bigger plan.

This is usually where the biggest win hides. Teams tend to assume their call volume is fixed and go straight to asking for more quota, when a surprising share of that volume turns out to be avoidable once someone actually traces where the calls are coming from.

Buy headroom before you buy a bigger plan

A dedicated rate limiter and queue in front of the outbound call gives you a place to smooth bursts before they ever reach the provider. That's often cheaper and faster to ship than a support ticket asking for a higher tier, and it protects you against the next provider's limit too, not just this one.

It also gives you a single place to add priority: a paying customer's request can jump ahead of a background sync job when the queue is under pressure, something a raw retry loop can't do for you.

Common mistakes that make a quota problem worse

Retrying immediately on every failure turns a brief limit into a sustained one, because your retries become part of the load that's keeping you rate limited. Failing to distinguish a rate limit from a real error means you might drop a request you should have retried, or hammer an endpoint that's actually down for another reason.

Treating the limit as an emergency to escalate every time, instead of building the backoff logic once, means the same incident repeats indefinitely. And sharing one API key across every environment, including staging and a developer's laptop, means a load test can burn through the quota your production traffic needed.

Avoid these patterns, which turn a brief limit into a sustained one:

  • Retrying immediately on every failure, so your retries become part of the load that keeps you rate limited.
  • Treating a rate limit like any other error, which can drop a request you should have retried or hammer an endpoint that is down.
  • Retrying blind on a fixed delay while ignoring the Retry-After information the provider already sent.
  • Handling each incident as a one-off instead of building the backoff and caching logic once.
Executive Capability Standard

What Good Looks Like

Good rate limit handling means a provider's 429 response degrades your feature gracefully instead of paging anyone or dropping data.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read each critical provider's documented rate limits and response headers before you're surprised by them in production.
2. Do Manually:Add basic retry-with-backoff around the handful of outbound calls that see the most traffic, starting with your busiest integration.
3. Delegate:Give one engineer ownership of an outbound request wrapper that every integration uses, so backoff logic isn't reinvented per call site.
4. Automate:Standardize on a shared client library with built-in backoff, jitter, and caching for every outbound integration in the codebase.
5. Buy:Add a queue or gateway in front of outbound calls when volume grows past what in-process backoff can smooth on its own.

How to Get Started

Frequently Asked Questions

Should we always retry a 429 response?

Only with backoff, and only up to a sane limit. Retrying immediately just adds to the load causing the limit in the first place. Respect any Retry-After header the provider sends, and give up gracefully after a handful of attempts rather than retrying forever.

How do we know if we need a bigger API plan or just better code?

Look at your actual call pattern first. If most of your volume is redundant or cacheable, fix that before paying for more headroom. If your traffic is genuinely growing and there's no waste left to cut, that's when a higher tier or a partner conversation makes sense.

What's the biggest mistake teams make with third-party rate limits?

Treating each incident as a one-off instead of building the backoff and caching logic once. The same provider will rate limit you again, and the fix should already be in place the second time it happens.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides