A Runbook for When an Upstream API Starts Throttling You
Rate limiting from a third-party API doesn't look like an outage at first. Response times stay normal, most requests still succeed, and then a growing slice of calls start coming back with a 429 while your own error dashboards, tuned for 5xx errors, stay quiet. By the time someone notices, a queue has usually backed up behind the throttled dependency.
This is a runbook for the four things that need to happen, in order: detect the throttling specifically, keep it from cascading into the rest of your system, protect the request path while it's happening, and then fix whatever's causing you to hit the limit in the first place.
How do you detect upstream 429s specifically?
A generic error-rate alert usually won't distinguish a provider's rate limiting from any other failure, and the response is different for each. Alert on 429 responses from each upstream dependency as their own signal, and if the provider publishes quota headers (remaining calls, reset time), track those directly instead of waiting for the first rejection. Say a dashboard shows you're most of the way through your quota with an hour left in the window; that gives you time to react before the first request actually gets rejected.
Absorb it at the client without retry-storming
Match your client's outbound rate to the provider's published limit with a token bucket or similar limiter, so you're shaping traffic before it hits the wall instead of reacting after. When a request does get throttled, retry with exponential backoff and jitter rather than a fixed delay; a fixed delay means every client that got throttled at the same moment also retries at the same moment, which just recreates the spike. If calls are bursty by nature (a batch job, a webhook fan-out), put a queue in front of the caller instead of letting it fire requests as fast as the work arrives.
Worker count matters here too. Adding more workers to drain a queue faster feels like the obvious fix for a backlog, but against a fixed upstream quota it just means more callers competing for the same limited pool of calls, so throughput doesn't actually improve; the limiter, not the worker pool, sets the ceiling. Size the limiter to the quota first, and only then decide how many workers you need to keep it saturated.
Protect the rest of the system while it's happening
Wrap the dependency in a circuit breaker so a throttled provider doesn't back up worker threads or connection pools that other, unrelated requests also need. Where the feature can tolerate it, degrade gracefully instead of blocking: serve cached or slightly stale data, disable a non-critical feature temporarily, or queue the work for later instead of making the user's request wait on a dependency that's currently refusing calls. The goal is to contain the blast radius to the one feature that depends on the throttled API, not let it take down request handling for everything else.
How do you fix the root cause of the throttling?
Batch calls where the provider supports it instead of making one request per record; a single request that fetches or updates fifty records instead of fifty separate calls buys you the same headroom as a fifty-fold increase in your limit, without asking anyone for anything. Cache aggressively for data that doesn't need to be fetched fresh on every request, especially reference data that changes rarely but gets requested on every page load. If usage is genuinely growing, ask the provider for a higher limit or a different pricing tier rather than engineering around a ceiling that's meant to be negotiated. Spreading load across multiple API keys is sometimes technically possible, but check the provider's terms first; several treat that as a violation and will revoke access for all of them at once, which turns a throttling problem into an outage.
A mistake worth naming: fixing the retry logic, not the origin
A common response to hitting a rate limit is to tune the retry count and backoff on the calling code and stop there. That reduces failed requests without reducing the number of calls being made, so the same quota gets consumed just as fast, and the throttling recurs at the same volume threshold. The retry logic controls how gracefully you fail; it doesn't touch how often you're asking. Fixing the volume, through batching, caching, or a higher limit, is the part that actually moves the ceiling.
When an upstream API starts throttling you, work through the runbook in order:
- Alert on 429 responses from each upstream dependency as their own signal, and track quota headers such as remaining calls and reset time if the provider publishes them.
- Match your outbound rate to the provider's limit with a token bucket, and retry throttled calls with exponential backoff and jitter.
- Wrap the dependency in a circuit breaker, and degrade gracefully with cached data or a disabled non-critical feature where you can.
- Cut the number of calls at the source by batching and caching, rather than only tuning retry counts.
What Good Looks Like
Good handling of upstream rate limits means the throttling shows up as a contained, specifically-alerted condition on one dependency, not a generic error spike that takes debugging time to trace back to its source.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should we build our own client-side rate limiter or just handle 429s as they come?
Build a limiter matched to the provider's published quota. Reacting only after a 429 means you've already wasted a call and added latency to that request; shaping traffic to stay under the limit avoids both, and it's usually a small amount of code compared with debugging cascading failures later.
Why does exponential backoff with jitter matter more than just retrying quickly?
Without jitter, every client throttled in the same window retries on the same schedule, recreating the exact spike that caused the throttling. Jitter spreads those retries out in time, so the provider sees a smoother trickle of retried requests instead of a second synchronized burst.
Is it reasonable to ask a vendor for a higher rate limit?
Yes, and it's usually faster than engineering around the ceiling. Bring usage data showing the growth trend and where the current limit is actually constraining you; most providers have a path to a higher tier or a custom limit for accounts with a clear, legitimate need.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Benchmarking API Gateway Latency the Right Way
A methodology for benchmarking API gateway latency that reflects real traffic, the mistakes that produce misleading numbers, and what to test beyond raw speed.
Keeping API Contracts From Breaking Between Services
A practical standard for versioning, owning, and validating API contracts so one team's change doesn't quietly break three other services.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.
Catching a Breaking API Change Before It Ships, Not After
How consumer-driven contract testing catches breaking changes between services before deploy, and how to set it up without slowing every release down.