Why Your Test Suite Passes and Your Deploys Still Break Things
A green test suite feels like permission to deploy, and in a distributed system that feeling is often wrong. Unit tests pass because each service works in isolation, and the actual failure only shows up once two services that were tested separately talk to each other in production.
This is what a pipeline needs beyond passing tests to actually catch that class of failure before it ships.
Unit tests prove a service works alone, not with its neighbors
Mocked dependencies in a unit test suite are only as accurate as whoever wrote the mock, and mocks drift from the real service's behavior the moment that service changes without the mock being updated to match. A test suite that's fully green can still be testing against a contract that no longer exists.
Integration tests against a real, or realistically faked, version of the dependency catch this class of drift. They're slower and more expensive to run, which is exactly why most teams under-invest in them relative to unit tests, but they catch the failures unit tests structurally can't.
Stage deploys so a bad change reaches a fraction of traffic first
Deploying a new version to one hundred percent of traffic at once means a bad change reaches every user simultaneously, and rollback happens after the damage, not before it. A canary or staged rollout, five percent of traffic, then twenty five, then everyone, catches most regressions while they're still affecting a small slice.
This needs real metrics watching the canary automatically, error rate, latency, not just a person eyeballing a dashboard for ten minutes before moving on. Staged rollouts are also what make a high deploy frequency safe in the first place: benchmark data on deploy frequency shows the fastest-moving engineering organizations shipping on demand, essentially daily, versus roughly monthly at the slowest end, and getting there depends on each individual release being small and cheap to verify1.
Make rollback as tested as the deploy itself
A rollback path that's never been exercised outside an emergency is a rollback path you can't trust during an emergency. Practicing a rollback on a low-stakes service, on a schedule, not just when something's actually broken, is what turns 'we can roll back' from a hope into a verified fact.
This matters more in a distributed system than a monolith, because rolling back one service while its dependencies have already moved forward can itself cause a compatibility break. Rollback plans need to account for what else changed since the version being restored.
A worked example: a passing pipeline that still broke production
Say a service's unit tests all pass after a schema change to an event it publishes, because the tests only check the publisher's own behavior. A downstream consumer, tested separately with a mock of the old event shape, breaks the moment the new version reaches production and the real event doesn't match what the mock promised.
A contract test between publisher and consumer, run in the pipeline before either side deploys, would have caught this before it shipped. This is the specific gap between 'tests pass' and 'the system works together' that services tested in isolation can't close on their own.
Where CI/CD pipelines have blind spots
- Integration tests skipped because they're slow, so only unit tests gate the deploy
- A staged rollout with no automated rollback trigger, relying on someone noticing the dashboard
- Rollback tested for the first time during an actual incident
- Contract tests that exist for some service pairs and not others, usually the newest ones
Treat pipeline gaps the same way you'd treat a code review gap
A pipeline that only checks unit tests is doing part of the job and giving false confidence about the rest. Closing that gap doesn't mean adding every possible check at once, it means identifying which failure modes have actually hit you in production and building the specific check that would have caught each one.
The pipeline earns trust the same way a person does: by catching real problems over time, not by being exhaustive on paper. Start with the last two or three incidents that a green pipeline didn't prevent, and build backward from there.
What Good Looks Like
A CI/CD pipeline that actually prevents incidents combines integration tests that catch cross-service drift, staged rollouts with automated rollback triggers, and contract tests between real dependencies.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How slow is too slow for a CI pipeline before people start skipping it?
Past roughly ten to fifteen minutes, engineers start batching changes or finding workarounds to avoid waiting, which defeats the purpose. Parallelizing tests and running the slowest integration suites only on the services actually affected by a change usually gets you back under that threshold.
Do we need canary deploys for a small team with low traffic?
A simpler staged rollout, even just one instance before the rest, catches a surprising share of obvious regressions without the full complexity of automated canary analysis. Save the more sophisticated tooling for when traffic volume makes a small canary slice statistically meaningful.
What's the fastest way to find where our pipeline has gaps?
Look at your last several production incidents and ask, honestly, whether a green pipeline would have caught each one. The pattern in what it missed tells you exactly what to build next, rather than guessing at hypothetical gaps.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Catching a Breaking API Change Before It Ships, Not After
How consumer-driven contract testing catches breaking changes between services before deploy, and how to set it up without slowing every release down.
Testing an AI Feature When 'Correct' Isn't a Fixed Answer
How to build an evaluation framework for AI-backed features in a distributed system, where a unit test can't tell you if the output is actually good.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
What an AI Code Reviewer Catches in a Distributed System, and What It Misses
Which distributed-systems failure modes AI code review catches well, which still need a senior engineer, and how to configure and roll out the tool.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.