Four Rules for API Integrations That Survive Production
The first API integration a team builds usually works fine, because it gets someone's full attention for a week. The third one breaks a customer's workflow six months later, because by then nobody remembers which service owns retries, what happens when a partner changes their schema without warning, or who gets paged when a webhook silently stops arriving.
These are four standards worth setting before you have that third integration, so the pattern is already in place instead of being reverse-engineered from an incident.
Version every contract, even the ones you control
An internal API between two of your own services still needs a version in the path or header, because internal does not mean stable. The service that seemed safe to change without a version bump is the one that breaks a downstream consumer nobody remembered existed, usually a batch job that runs once a month.
Treat a breaking schema change as a new version, not an in-place edit, and give consumers a deprecation window before the old version disappears. Six weeks is a reasonable default unless a contract says otherwise.
How should you handle retries and idempotency across integrations?
Every new integration reinvents retry logic unless there is already a shared pattern to copy. Without one, some services retry three times with no delay, others retry forever, and a few do not retry at all, and the difference only shows up during a partner's outage, at the worst possible time.
Write down one retry policy, exponential backoff with a capped number of attempts, and one rule for idempotency: every write operation accepts an idempotency key so a retried request cannot double-charge a customer or create a duplicate record. Apply both to every new integration by default instead of deciding fresh each time.
Write these rules down once and apply them to every integration:
- Use exponential backoff with a capped number of attempts for every retry, so no service retries forever or retries instantly.
- Require every write operation to accept an idempotency key, so a retried request cannot double-charge a customer or create a duplicate record.
- Make these rules the default for each new integration, instead of deciding fresh every time.
- Publish the policy in a short internal reference, so the second integration reuses it instead of rebuilding it.
What happens when a partner's API changes without warning?
A schema validation step on every inbound webhook and API response catches a silent partner change before it corrupts your data, instead of after a support ticket arrives asking why a report looks wrong. Log the raw payload before you parse it, so when a field disappears or changes type, you have the evidence to show the partner instead of a guess.
A partner's status page is not a substitute for your own monitoring. Alert on payload shape changes and unexpected null fields directly, since a partner's outage dashboard will not tell you about a quiet breaking change that technically still returns a 200.
Name an owner before the integration ships, not after it breaks
Every integration needs a named engineer or team who gets paged when it fails, documented somewhere searchable, not just known by whoever happened to build it. When that person is unavailable and the integration breaks, the on-call engineer needs to find the owner, the partner's support contact, and the last known-good behavior in minutes, not by asking around in chat.
A short runbook per integration, what it does, who owns it, how to tell it is broken, and how to roll back, turns a three-hour incident into a twenty-minute one the first time it actually matters.
For example, a billing partner's webhook stops arriving late on a Friday night. If the integration has a named owner and a one-page runbook, the on-call engineer can see what the integration does, who to contact at the partner, how to confirm the webhook has stopped, and how to roll back to the last known-good behavior. Without that page, the same engineer spends the evening asking around in chat. The common mistake is writing the runbook after the first incident. Write it before launch, store it where search will find it, and review it whenever the owner changes teams so the named person is always current.
The second integration is where the pattern actually gets tested
The first integration a team builds gets careful attention because it is new and nobody has a shortcut to reach for yet. The second one is where the real test happens: does the team reuse the retry policy, the validation approach, and the runbook template from the first one, or does each engineer quietly rebuild their own version because the pattern was never written down anywhere.
Write the pattern down after the first integration ships, while the decisions are still fresh, rather than waiting until the third one forces the conversation. A short internal reference document that says exactly how retries, versioning, and ownership work saves more time on the second integration than any tooling choice does, and it turns a growing pile of one-off integrations into something that actually looks and behaves like a system.
What Good Looks Like
Good here means every integration has a documented owner, a written retry and idempotency policy, and payload validation on inbound data, and a new engineer can find all three without asking someone who was there when it was built.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How long should a deprecation window be for an internal API change?
Six weeks is a reasonable default for most internal services, long enough for a downstream team to notice and migrate without treating it as an emergency. Extend it for anything touching batch jobs that run monthly or quarterly, since those consumers may not exercise the old path again before the window closes.
Do we really need idempotency keys for every write, even low-risk ones?
Apply the rule consistently rather than deciding case by case, since the retries that cause real damage are rarely the ones anyone predicted in advance. The overhead of adding an idempotency key is small compared to the cost of debugging a duplicate charge or a duplicated record weeks after the fact.
What is the most common cause of integration outages that isn't the partner's fault?
A schema change on your own side, shipped without checking who else reads that data, is more common than most teams expect. The fix is the same discipline you would apply to a partner's API: version it, validate it, and know who consumes it before you change it.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Setting API Integration Standards Before a Postmortem Forces Them
The versioning, error shape, and idempotency decisions worth making before your API has enough integrations that changing them breaks someone.
Why Your API Gateway Load Test Doesn't Match Production
Most gateway benchmarks measure the wrong thing: raw throughput on a synthetic route. Here is how to test what actually matters for your traffic.
What to Build Before Your Next Vendor API Throttles You
Backoff, jitter, circuit breakers, and quota tracking: the pieces every team needs before an upstream API's rate limit turns into a production incident.
Where Production Deployment Budgets Quietly Leak
The recurring places engineering teams overspend on production deployment architecture, and a practical order for fixing them without a full rebuild.
The API Integration Standards Partners Actually Need From You
Answers to the questions partners and internal teams actually ask when integrating with your APIs under a zero trust model, from auth method to versioning.
Four Safeguards Before You Ship a New API Integration
The four checks that catch most API integration failures before they reach production: contracts, auth boundaries, error handling and versioning.