Building a CI/CD Pipeline That Actually Catches Bugs
A lot of CI pipelines are green on every merge and still let real bugs through, because the checks that run are fast and easy to write, linting, formatting, a handful of unit tests, rather than the checks that would actually catch the bug that shipped last month.
A pipeline that's earned the team's trust isn't the one with the most steps. It's the one where a green check mark genuinely means someone doesn't need to manually re-verify the change before it ships.
How do you pick CI tests from your last five real bugs?
Look at the last five bugs that reached production and ask, honestly, whether any reasonable test would have caught each one. Some will be genuinely hard to test for; most will reveal a specific gap, no integration test covering that code path, no check for that particular edge case, that's worth closing directly rather than adding tests in the abstract.
This produces a much more useful test suite than starting from a textbook ratio of unit to integration to end-to-end tests, because it's anchored to failures you've actually had instead of failures you're guessing you might have.
Repeat this exercise every quarter, not just once when you're first designing the pipeline. The category of bug that slips through changes as the codebase and team change, and a test suite built entirely around last year's failure patterns can develop blind spots of its own if nobody revisits it against what's actually been breaking more recently.
Which CI checks should block a merge?
Not every check needs to be a hard gate. Style and formatting can auto-fix or warn without blocking; tests covering core business logic should block without exception; a flaky end-to-end test that fails one time in twenty shouldn't block merges at all until it's fixed, because a gate everyone learns to ignore or re-run until it passes stops being a gate.
A pipeline that's earned the team's trust isn't the one with the most steps, it's the one where a green check mark genuinely means someone doesn't need to manually re-verify the change before it ships.
A workable split of checks looks like this:
- Block merges on tests that cover core business logic and known regression risks, without exception.
- Let style and formatting checks auto-fix or warn instead of blocking the merge.
- Keep a flaky end-to-end test out of the blocking path until it is fixed, because a gate people re-run until it passes stops being a gate.
Keep the feedback loop fast enough that people wait for it
A ten minute pipeline gets watched. A forty minute pipeline gets context switched away from, and the person merges based on a green check they glanced at twenty minutes ago rather than the actual current state at the moment they hit merge. Parallelize test suites, cache dependencies aggressively, and split slow end-to-end suites to run only on the paths they actually cover instead of on every single change regardless of relevance to that change.
A pipeline that reliably catches a fewer number of things but runs fast enough that people actually wait for and trust it beats a slower, more thorough one that quietly trains people to merge before it finishes running.
Deploy cadence follows pipeline trust, not the other way around
A pipeline that actually blocks bad merges is what separates teams with strong deploy frequency, DORA's top performing band, from teams stuck at the 30 day medium tier because every release still needs a manual regression pass before anyone trusts it1. Investing in pipeline reliability is usually a more direct path to shipping more often than any process change aimed at deploy frequency on its own, since the process change without the underlying trust just adds pressure without removing risk.
Teams rarely decide to "deploy more often" and succeed by willpower alone; they build a pipeline trustworthy enough that deploying often stops feeling risky, and the frequency follows naturally from that confidence.
A common mistake: adding tests only for the bug that just happened
Reactive test writing, add a test for exactly the bug that just shipped, is better than nothing but tends to produce a suite full of narrow regression tests and gaps everywhere else. When a bug reveals a gap, ask what category of bug it represents, not just the specific case, and write a test broad enough to catch the category, not only the exact input that triggered it last time.
Over time this produces a suite that's actually resilient to new bugs in the same area, rather than one that only ever catches the exact same mistake twice.
What Good Looks Like
CI/CD is working when a green pipeline genuinely means a change is safe to ship without manual re-verification, and the pipeline runs fast enough that the team actually waits for and trusts it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should every failing check block a merge?
No. Reserve hard blocking for checks covering core business logic and known regression risks. A check that's frequently flaky or low-value as a blocker should warn instead of block, or be fixed until it's reliable enough to trust as a gate again.
How fast should a CI pipeline run to keep people engaged with it?
Under ten minutes is a good target for the checks that run on every merge; anything longer and people start merging based on a check they glanced at a while ago rather than the current result. Push slower, more thorough suites to run less frequently rather than on every single change.
Is full test coverage a reasonable goal for a pipeline?
Not on its own. Coverage percentage tells you what code ran during tests, not whether the tests actually verify correct behavior. A team chasing a coverage number can hit it while still missing the exact bugs that matter; anchor test priorities to real past failures instead.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
What Your CI/CD Pipeline Actually Costs You
A way to think about CI/CD pipeline cost beyond the compute bill, including engineer waiting time, flaky test triage, and what to fix first.
A Worksheet for Sizing Your CI Pipeline's Real Cost
A step-by-step worksheet for pricing out what your automated test pipeline actually costs in compute and engineering wait time, and where to trim it.
Building a CI/CD Pipeline That Understands Streaming Code
A step by step way to build CI/CD around stream processing code, so topic changes, schema checks, and consumer deploys are automated, not manual steps.
Building a CI/CD Pipeline That Doesn't Slow You Down
How to build a CI/CD pipeline engineers actually trust: what belongs in it, why speed matters more than coverage, and how deploy frequency really changes.
CI/CD Stages That Actually Catch RAG Pipeline Regressions
A standard test suite misses RAG failure modes. Here are the CI/CD stages worth adding: retrieval gates, model version checks, and a real test index.
Building a CI/CD Pipeline for Agent and Tool Code
What changes about continuous integration once prompts and tool definitions ship alongside code, and how to test both before they reach production.