Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

How to Stop Getting Rate Limited by Your Own Vendors

A background job starts failing at 2am because a vendor API returns 429s. Someone bumps the retry timeout, the alerts go quiet, and the same failure comes back next month with a bigger blast radius, because nothing about the actual demand on that vendor changed.

Most of these outages are self-inflicted: your own concurrency spiked past what you'd planned for, not past what the vendor actually allows. Fixing that takes quota planning, not a longer retry loop.

Finding the Real Ceiling: Vendor Quota vs Your Own Concurrency

Every vendor publishes some form of rate limit, usually requests per minute per API key, sometimes a shared ceiling across your whole account. That published number is rarely what actually breaks you. What breaks you is your own code opening more concurrent calls than you intended, often because an autoscaler added instances or a batch job and a user-facing request happened to overlap.

Before you touch retry logic, separate the two questions: what does the vendor actually allow, and what does your system actually attempt at peak. If your peak concurrency is well under the documented limit and you're still getting throttled, you're likely hitting an undocumented per-second burst limit rather than the per-minute number in the docs.

Reading a 429 Correctly

A 429 response is not one thing. Check for a Retry-After header first; if the vendor tells you exactly how long to wait, honor that instead of guessing. If there's no header, exponential backoff with jitter avoids every failed caller retrying at the same instant and reproducing the same spike a second later.

Also distinguish a hard quota rejection, you've used your entire monthly allotment, from a transient throttle, you sent too many requests in one second. The first needs a business conversation with the vendor or your own usage; the second just needs better pacing.

Splitting Quota Across Teams and Jobs

When more than one team or job calls the same vendor, an unmanaged shared quota turns into a race: whichever job runs first that hour eats the budget. A token bucket per consumer, with a priority tier for user-facing requests over a nightly batch job, keeps one team's backfill from starving another team's live traffic.

Have batch jobs back off automatically during business hours and take the leftover quota overnight. That single rule fixes more vendor throttling incidents than any amount of retry tuning.

When to Ask for a Higher Limit vs Redesign the Job

If your steady-state usage is regularly close to the documented limit and the traffic driving it is genuine user activity, that's a real case for asking the vendor for a limit increase, and bring usage data to that conversation. If the pressure only shows up during a nightly batch window, redesigning the batch job to spread its calls over a longer window usually solves it for free, without waiting on a vendor's account team.

A Checklist Before You Add a New Vendor Dependency

  • Confirm the published per-key and per-account limits, and whether they're shared across your whole organization or scoped to one API key.
  • Build backoff and jitter into the integration before its first production call, not after the first outage.
  • Set an internal alert at a share of the documented limit so you see the trend coming instead of finding out from a 429.
  • Decide up front which of your own jobs are allowed to be delayed if quota gets tight, and which are not.

What to Bring to the Vendor Instead of Just Asking for More

A limit increase request lands very differently with real numbers attached than as a vague complaint about getting throttled. Pull your actual peak requests per minute over the last month, the growth trend behind it, and which endpoints are driving the volume, and bring that to the conversation instead of asking for a round number that sounds safe.

That same data tells you something a vendor's response time never will: whether the growth is coming from genuine usage you want to keep serving, or from an inefficient integration calling the same endpoint repeatedly when it could cache the result or batch several lookups into one request. Fixing the inefficient call is usually cheaper than negotiating a bigger limit, and it's worth ruling out first.

Executive Capability Standard

What Good Looks Like

Good quota management means you can see how close you are to every vendor's rate limit before a request gets rejected, and your own concurrency never spikes past what you've planned to queue.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull the documented per-key and per-account limits for every vendor API you call in production and compare them against your actual peak call volume.
2. Do Manually:Add exponential backoff with jitter by hand to the two or three integrations that fail most often today.
3. Delegate:Assign one engineer to own a shared quota budget across every job and service that calls the same vendor.
4. Automate:Build a token bucket limiter in front of each vendor integration so your own code enforces the ceiling before the vendor has to.
5. Buy:Bring in a platform engineer to redesign batch scheduling around vendor quota windows instead of repeatedly asking for higher limits.

How to Get Started

Frequently Asked Questions

Should every failed request just retry automatically?

No. Blind automatic retries during a throttle amplify the exact problem you're trying to fix, because every failed caller tries again at roughly the same moment. Queue the request instead, back off with jitter, and only retry a bounded number of times before surfacing the failure.

How do I know I'm approaching a vendor's rate limit before it becomes an outage?

Watch the rate limit headers most APIs return on every response, not just failed ones, and alert when usage crosses a share of the documented ceiling. Waiting for the first 429 to notice a problem means you already failed a real request before you had any warning.

Is paying for a higher tier the fix?

Sometimes, but only after you've ruled out a self-inflicted concurrency spike. If your actual steady demand is genuinely near the limit, a higher tier is the right call. If the spike only happens because a batch job and live traffic overlap, redesigning the schedule fixes it without a new bill.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides