Circuit Breakers and Bulkheads: Stopping a Failure From Spreading
A single slow or failing dependency can take down an entire service that's otherwise healthy, if every request thread ends up blocked waiting on that one dependency to respond. Circuit breakers and bulkheads are two different answers to that same problem, and most resilient systems need both, not one or the other.
Getting either pattern's thresholds wrong is worse than not having them at all, since a misconfigured breaker either trips on normal traffic or never trips when it should.
What a circuit breaker actually does
A circuit breaker watches the failure rate of calls to a specific dependency, and once that rate crosses a threshold, it stops sending new calls entirely for a cooldown period, failing fast instead of letting every request queue up waiting on a dependency that's already struggling.
After the cooldown, it allows a small number of test requests through; if those succeed, it closes again and traffic resumes normally, and if they fail, it stays open and waits longer. The point isn't preventing the failure, it's preventing the failure from cascading into every other part of the system that's waiting on a response.
Bulkheads: isolating so one failure can't starve everything else
A bulkhead limits how much of a shared resource, a thread pool, a connection pool, concurrent request slots, any single dependency or feature can consume, so a slow or failing one can't starve every other feature of the same resource.
Without bulkheads, a circuit breaker on one dependency can still fail to help if every thread in your shared pool is already blocked waiting on that same slow dependency by the time the breaker trips. Partition your resource pools per dependency, or at least per criticality tier, so a checkout-blocking failure and a recommendations-widget failure aren't drawing from the same limited pool of threads.
Protecting the uptime budget these patterns exist for
A circuit breaker's whole job is making sure one flaky dependency doesn't spend your entire uptime budget at once, and at 99.95 percent availability that budget is only 4.38 hours a year1.
A dependency failure that's contained by a breaker and a bulkhead costs you a degraded feature for its cooldown period; the same failure without containment can cost you the whole service for as long as the dependency stays down, which is a very different draw against the same annual budget.
Tuning mistakes that make these patterns worse than nothing
- A failure threshold set so low that normal, brief error-rate blips trip the breaker constantly, training engineers to ignore its alerts or disable it outright.
- A cooldown period so short that the breaker flaps open and closed rapidly against a dependency that's genuinely struggling, adding overhead without actually protecting anything.
- Sharing one thread pool across every dependency despite having circuit breakers configured per dependency, which defeats the isolation a bulkhead is supposed to provide.
- No fallback behavior defined for when a breaker is open, so users get a hard error instead of a degraded but functional experience, like cached data or a simplified response.
A worked example: one dependency, contained
Say a third-party shipping-rate API starts timing out under its own load. Without a circuit breaker, every checkout request waits the full timeout duration for a response that never comes, threads pile up waiting, and the connection pool serving that call exhausts, which starts affecting unrelated requests sharing the same pool. Checkout, and possibly the whole service, goes down with a dependency that was never actually part of its core function.
With a circuit breaker and a bulkhead isolating the shipping-rate call to its own resource pool, the breaker trips after a handful of timeouts, checkout falls back to a flat-rate shipping estimate for the cooldown period, and every other part of the service, including checkout's core purchase flow, keeps working normally. The dependency failure cost a slightly less accurate shipping quote for a few minutes instead of a full outage.
Reviewing thresholds after every real trip
A circuit breaker's configuration shouldn't be a one-time setup task. Every time a breaker actually trips in production, that's real data about how that dependency fails, how long it typically takes to recover, and whether the current threshold and cooldown period matched what actually happened. Review the trip afterward the same way you'd review any incident: did the breaker trip too early, too late, or about right, and does the fallback behavior that activated actually serve users well, or just technically avoid an error.
What Good Looks Like
A good resilience setup pairs a tuned circuit breaker with bulkhead resource isolation per dependency, and defines a real fallback for when the breaker is open, so a contained failure degrades gracefully instead of just failing fast.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the difference between a circuit breaker and a bulkhead?
A circuit breaker stops sending requests to a dependency once its failure rate crosses a threshold, failing fast instead of queuing. A bulkhead limits how much of a shared resource, like a thread pool, any single dependency can consume, so one failing dependency can't starve every other feature of the same resource. Most resilient systems need both.
How do we pick the right failure threshold for a circuit breaker?
Start conservative, higher than your normal baseline error rate with margin, and tighten it based on real incident data rather than guessing upfront. A threshold set too low trips on ordinary blips and trains engineers to ignore or disable the breaker, which defeats the purpose.
What should happen when a circuit breaker is open?
Define a fallback, not just a hard error: cached data, a simplified response, or a clear degraded-mode message, depending on what the dependency provides. A breaker with no fallback behavior just moves the failure from slow and cascading to fast and total, which is better but still not the goal.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Circuit Breakers and Bulkheads, Explained With Checkout
A slow payment provider times out, threads pile up waiting, and the whole service goes unresponsive. How circuit breakers and bulkheads contain that.
The Circuit Breaker Checklist Most Teams Skip Half Of
A checklist for implementing circuit breakers and bulkheads correctly: the failure thresholds, half-open behavior, and isolation mistakes teams miss.
Circuit Breakers and Bulkheads: Stopping One Bad Dependency
How circuit breakers and bulkheads stop one slow dependency from cascading into a full outage, plus the thresholds and pitfalls that make them work.
Circuit Breakers and Bulkheads: Configuring Them So They Help
How to set trip thresholds against your real availability target, contain failures with bulkheads, and avoid the mistake of one setting for every call.
When a Circuit Breaker Helps, and When It Just Hides a Bug
A decision guide for using circuit breakers and bulkheads to stop one failing dependency from cascading, and where a circuit breaker can mask a real problem.
Circuit Breakers and Bulkheads: What Each One Actually Prevents
A comparison of circuit breaker and bulkhead resilience patterns: what failure each one actually prevents, and how to set thresholds that work.