Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Why Your Test Suite Passes and Your Deploys Still Break Things

A green test suite feels like permission to deploy, and in a distributed system that feeling is often wrong. Unit tests pass because each service works in isolation, and the actual failure only shows up once two services that were tested separately talk to each other in production.

This is what a pipeline needs beyond passing tests to actually catch that class of failure before it ships.

Unit tests prove a service works alone, not with its neighbors

Mocked dependencies in a unit test suite are only as accurate as whoever wrote the mock, and mocks drift from the real service's behavior the moment that service changes without the mock being updated to match. A test suite that's fully green can still be testing against a contract that no longer exists.

Integration tests against a real, or realistically faked, version of the dependency catch this class of drift. They're slower and more expensive to run, which is exactly why most teams under-invest in them relative to unit tests, but they catch the failures unit tests structurally can't.

Stage deploys so a bad change reaches a fraction of traffic first

Deploying a new version to one hundred percent of traffic at once means a bad change reaches every user simultaneously, and rollback happens after the damage, not before it. A canary or staged rollout, five percent of traffic, then twenty five, then everyone, catches most regressions while they're still affecting a small slice.

This needs real metrics watching the canary automatically, error rate, latency, not just a person eyeballing a dashboard for ten minutes before moving on. Staged rollouts are also what make a high deploy frequency safe in the first place: benchmark data on deploy frequency shows the fastest-moving engineering organizations shipping on demand, essentially daily, versus roughly monthly at the slowest end, and getting there depends on each individual release being small and cheap to verify1.

Make rollback as tested as the deploy itself

A rollback path that's never been exercised outside an emergency is a rollback path you can't trust during an emergency. Practicing a rollback on a low-stakes service, on a schedule, not just when something's actually broken, is what turns 'we can roll back' from a hope into a verified fact.

This matters more in a distributed system than a monolith, because rolling back one service while its dependencies have already moved forward can itself cause a compatibility break. Rollback plans need to account for what else changed since the version being restored.

A worked example: a passing pipeline that still broke production

Say a service's unit tests all pass after a schema change to an event it publishes, because the tests only check the publisher's own behavior. A downstream consumer, tested separately with a mock of the old event shape, breaks the moment the new version reaches production and the real event doesn't match what the mock promised.

A contract test between publisher and consumer, run in the pipeline before either side deploys, would have caught this before it shipped. This is the specific gap between 'tests pass' and 'the system works together' that services tested in isolation can't close on their own.

Where CI/CD pipelines have blind spots

  • Integration tests skipped because they're slow, so only unit tests gate the deploy
  • A staged rollout with no automated rollback trigger, relying on someone noticing the dashboard
  • Rollback tested for the first time during an actual incident
  • Contract tests that exist for some service pairs and not others, usually the newest ones

Treat pipeline gaps the same way you'd treat a code review gap

A pipeline that only checks unit tests is doing part of the job and giving false confidence about the rest. Closing that gap doesn't mean adding every possible check at once, it means identifying which failure modes have actually hit you in production and building the specific check that would have caught each one.

The pipeline earns trust the same way a person does: by catching real problems over time, not by being exhaustive on paper. Start with the last two or three incidents that a green pipeline didn't prevent, and build backward from there.

Executive Capability Standard

What Good Looks Like

A CI/CD pipeline that actually prevents incidents combines integration tests that catch cross-service drift, staged rollouts with automated rollback triggers, and contract tests between real dependencies.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Review your last three production incidents and identify exactly what kind of check, if it had existed, would have caught each one.
2. Do Manually:Add the highest-value missing check from that review to your pipeline by hand for your most critical service pair.
3. Delegate:Give one team ownership of pipeline reliability, distinct from feature delivery, so pipeline gaps get fixed instead of worked around.
4. Automate:Add automated canary analysis and rollback triggers so a bad deploy is caught and reverted without waiting for a person to notice.
5. Buy:Bring in a platform engineering specialist once you're maintaining pipeline infrastructure across more services than one team can own well.

How to Get Started

Frequently Asked Questions

How slow is too slow for a CI pipeline before people start skipping it?

Past roughly ten to fifteen minutes, engineers start batching changes or finding workarounds to avoid waiting, which defeats the purpose. Parallelizing tests and running the slowest integration suites only on the services actually affected by a change usually gets you back under that threshold.

Do we need canary deploys for a small team with low traffic?

A simpler staged rollout, even just one instance before the rest, catches a surprising share of obvious regressions without the full complexity of automated canary analysis. Save the more sophisticated tooling for when traffic volume makes a small canary slice statistically meaningful.

What's the fastest way to find where our pipeline has gaps?

Look at your last several production incidents and ask, honestly, whether a green pipeline would have caught each one. The pattern in what it missed tells you exactly what to build next, rather than guessing at hypothetical gaps.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides