Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Keeping API Contracts From Breaking Between Services

A distributed system runs on contracts between services, and most of the incidents that look like a code bug are actually a contract nobody agreed on breaking. One team changes a field's type, another team's parser chokes on it three deploys later, and by then nobody remembers the connection.

This is a working standard for keeping those contracts stable without freezing every API in place.

Version your API before you're forced to

Put a version in the URL path or a header from the very first release, even when there's only one consumer. Retrofitting versioning after three teams depend on an unversioned endpoint is far more painful than starting with it, because now every change has to be backward compatible by accident instead of by design.

A simple rule works for most teams: additive changes, new optional fields, new endpoints, ship without a version bump. Anything that removes a field, changes a type, or changes required behavior gets a new version, with the old one supported until every known consumer has moved.

Decide who owns a contract before two teams disagree

The team that owns the service publishing the API owns the contract, full stop, but that only works if consumers have a clear, low-friction way to request a change instead of just calling the owning team's Slack channel and hoping.

A short RFC process, even a one-page doc and a 48-hour comment window, forces the conversation to happen before the breaking change ships instead of after it's already in production and someone's dashboard is red.

Validate the schema in CI, not in someone's head

Manual contract review catches maybe half of what breaks. Automated schema validation, run against every pull request that touches a public interface, catches the rest: the field that changed type, the required property that quietly became optional, the enum value that got removed.

Consumer-driven contract tests take this further, letting each consuming service assert what it actually depends on, so a producer's CI run fails immediately if a change would break a real consumer instead of an assumption about what might be out there.

A worked example: a webhook integration that broke silently

Say a billing service adds a new required field to its webhook payload to support a new plan type. Internal consumers get updated in the same release. An external partner integration, built against the old shape, starts failing validation on every event and silently drops them, since webhook failures rarely page anyone the way an API error would.

Nobody notices until the partner asks why their records stopped updating two weeks later. Making the new field optional with a sensible default, or versioning the webhook payload the same way as the API, would have kept old consumers working while the new field rolled out.

Where API governance breaks down in practice

  • A breaking change shipped because the team assumed nobody else was calling that endpoint
  • Documentation that describes the API as it was six months ago, not as it is today
  • No deprecation timeline, so an old version lingers forever because removing it feels risky
  • Internal and external consumers held to different standards, so external partners break more often

Most of these come from treating internal APIs as informal, when the actual blast radius of breaking one is often larger than an external partner integration.

Keep documentation close enough to the code that it stays true

Documentation that lives in a separate wiki drifts from reality within a few releases. Generating docs from the schema itself, an OpenAPI spec, protobuf definitions, whatever your stack uses, means the documentation can't drift further than the schema does, because it's the same artifact.

This also gives you a free source of truth for the contract tests described above: the schema becomes both what's documented and what's enforced, instead of two things maintained separately that quietly disagree.

When it's fine to break a contract on purpose

Not every breaking change is a mistake. Sometimes the old shape is actively wrong, a field encodes the wrong unit, an enum is missing a value the business now needs, and holding onto it forever just preserves the bug for new consumers too.

The difference between a planned break and an accidental one is whether it goes through the same version bump and deprecation window as everything else. Give the old version a real end date, tell every known consumer directly instead of relying on a changelog nobody reads, and confirm the last of them has moved before you delete it. A break that follows the process is a normal part of running a system for years. A break that skips it is the incident this whole standard exists to prevent.

Executive Capability Standard

What Good Looks Like

A well-governed API surface means every contract is versioned from the start, has a named owner, and is validated automatically in CI rather than relying on a reviewer noticing a breaking change.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pick your three most-called internal APIs and list every known consumer for each one, including the ones nobody wrote down.
2. Do Manually:Write a one-page contract doc for your riskiest shared API and circulate it for a 48-hour comment window before your next breaking change.
3. Delegate:Give each service's owning team explicit authority over its contract, with a documented process for consumers to request changes.
4. Automate:Add schema validation and consumer-driven contract tests to CI so a breaking change fails the build instead of shipping.
5. Buy:Bring in a fractional CTO or API platform specialist once you're coordinating contract changes across more teams than a lightweight RFC process can handle.

How to Get Started

Frequently Asked Questions

How long should we support an old API version after deprecating it?

Long enough for every known consumer to migrate, with a hard deadline communicated up front rather than an open-ended promise. For internal services, 60 to 90 days is usually enough once you've confirmed who's actually still calling the old version.

Do internal-only APIs need the same rigor as public ones?

Yes, and often more. A public API has one clearly defined boundary and usually a smaller number of consumers than an internal API that's been quietly adopted by six other teams over two years without anyone tracking it.

Is GraphQL easier to version than REST?

It shifts the problem rather than solving it. GraphQL's schema evolution model handles additive changes gracefully, but removing or renaming a field still breaks any client selecting it, so you still need deprecation tracking and consumer awareness, just expressed differently.

What's a reasonable way to find out who's actually calling an internal API before we change it?

Check access logs or request tracing for the endpoint over the last 30 to 60 days rather than trusting a wiki page. Undocumented consumers show up here far more often than teams expect, and this is the check that catches them before a change ships.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides