How to Catch Breaking API Changes Before They Reach Production
A service passes its own test suite. The service that depends on it passes its own test suite too. The two still break when deployed together, because nothing tested the actual boundary between them, only each side's own assumptions about what that boundary looks like.
This is a runbook for testing that boundary directly: catching a breaking API change before it reaches a service that depends on it, not after a deploy already shipped it.
Why do unit tests miss breaking API changes?
Unit tests are written against a service's own understanding of its dependencies, usually a mock built by whoever wrote the test. If the real dependency's behavior drifts (a field gets renamed, an error code changes, a previously optional field becomes required) the mock keeps agreeing with the code that built it, and the test suite stays green while the actual integration is broken.
This gap is invisible until a deploy, because it only shows up when both sides run against each other for real. A contract test closes that gap by checking a service's actual behavior against a shared, explicit definition of the contract, not against a mock either side controls on its own.
Start with the boundaries that would hurt the most if they broke
Not every service boundary needs the same investment. Rank your internal API boundaries by what breaks downstream if the contract drifts silently: a boundary that feeds a customer-facing checkout flow deserves contract tests before an internal reporting endpoint that a nightly batch job reads.
Write that list down explicitly rather than testing whichever boundary happens to be easiest to reach first. Teams that start with the boundaries most convenient to test often end up with strong coverage on low-risk paths and none on the ones that would actually cause an incident.
For example, suppose your checkout service calls a pricing service, and a nightly reporting job reads from an analytics endpoint. If the pricing service renames a field, customers see failed orders within minutes. If the analytics endpoint renames one, a report is late the next morning. Both boundaries matter, but only one of them justifies a contract test this week. A simple decision rule: if a silent break would reach a customer before anyone on your team notices, the boundary goes to the top of the list. Everything else can wait for the second round, once the first contracts are running in the pipeline.
Test the contract, not just each side's own assumptions
The core idea of a consumer-driven contract test is simple: the consumer of an API states what it actually depends on (specific fields, specific status codes, specific error shapes), and that expectation becomes a test the provider runs against its own code. Instead of the provider guessing what consumers need, the provider can see explicitly what would break them.
This catches a category of bug that neither side's own test suite can see alone: the provider's tests confirm it does what the provider intended, and the consumer's tests confirm it handles what the consumer expects, but neither one confirms those two things still actually match after a change.
How do you run contract tests so they never get skipped?
A contract test that lives in a wiki page or a manual pre-release checklist gets skipped the first time a release is rushed, and it stops protecting anything at exactly the moment it matters most. Wire it into the same pipeline that runs your other automated tests, so a broken contract fails the build the same way a failing unit test would.
Teams with the fastest deploy cadence ship multiple releases a day rather than batching changes into infrequent, larger drops1, and automated contract tests are what make that pace safe instead of reckless, since there is no time for a person to manually verify every service boundary before each one ships.
A common mistake: testing the happy path and skipping error responses
It's tempting to write a contract test that confirms the success case, a 200 response with the expected fields, and call the boundary covered. Consumers depend on error behavior just as much as success behavior: what status code a timeout returns, what shape a validation error takes, whether a missing optional field comes back as null or gets omitted entirely.
A provider that quietly changes its error format, while still passing every happy-path contract test, can break every consumer's error handling in production without a single test catching it beforehand. Include at least one failure case in every contract you write, not only the success case, since that's usually where undetected drift actually lives.
A workable path from no contract tests to protected boundaries:
- Rank your internal API boundaries by what breaks downstream if the contract drifts, and write that ranked list down.
- Have each consumer state exactly which fields, status codes and error shapes it depends on.
- Run those expectations as a test against the provider's own code before every change ships.
- Include at least one failure case, such as a timeout or validation error, in every contract.
- Wire the contract tests into the same pipeline as your other automated tests so a broken contract fails the build.
What Good Looks Like
Good here means every boundary that would hurt production if it broke has an automated contract test running in the pipeline, and a provider change that breaks a consumer's stated expectations fails the build before it can merge.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the difference between contract testing and integration testing?
Integration testing runs two real services together end to end, and tends to be slow and brittle since it depends on both being deployed and configured correctly. Contract testing checks one side's behavior against a shared, explicit definition of what the other side expects, without needing both running together, so it stays fast enough to run in every pipeline.
Which side should own writing the contract test, the consumer or the provider?
The consumer should define what it actually depends on, since that's the part that would break if it changed. The provider then runs that expectation as a test against its own code before every change ships, which is what catches a break before it reaches production instead of after.
How do we handle a breaking change when we genuinely need to change a contract?
Version the contract or add the new behavior alongside the old one for a transition window, giving every consumer time to migrate before the old behavior goes away. Update the contract test to cover both versions during that window, then retire the old version's test once every consumer has confirmed it migrated.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Build or Buy: Deciding on an Evaluation Framework
A decision guide for choosing between a custom evaluation framework and an off-the-shelf one, based on what actually differs about your testing needs.
Why Synthetic Load Tests Miss the Failures That Actually Happen
The specific ways a synthetic load test differs from a real traffic spike, and what to build into the test so it catches what actually breaks.
Giving Every Pull Request Its Own Disposable Environment
A worked example of moving from one shared staging environment to per-PR ephemeral environments, including safe seed data and teardown cost control.
Where Production Deployment Budgets Quietly Leak
The recurring places engineering teams overspend on production deployment architecture, and a practical order for fixing them without a full rebuild.
What "Zero Trust" Actually Means for Device Verification
Zero trust device verification means a device is trusted continuously, based on its current state, not once at login. Here is what that actually requires.
How to Benchmark Your System Before It Has to Scale
A practical runbook for benchmarking throughput and capacity before you actually need the headroom, so scaling decisions are based on data, not guesses.