Circuit Breakers and Bulkheads: Stopping One Bad Dependency
A single slow dependency, a third-party API taking ten seconds to respond instead of two hundred milliseconds, shouldn't be able to take down your entire application. In practice, it often does, because requests pile up waiting on that one slow call, threads or connections get exhausted, and the failure spreads to parts of your system that have nothing to do with the original dependency.
Circuit breakers and bulkheads are the two patterns built specifically to stop that spread, and they solve slightly different parts of the problem.
What a Circuit Breaker Actually Does
A circuit breaker watches the failure rate or latency of calls to a specific dependency, and once that rate crosses a threshold, it stops making new calls entirely for a cooldown period, returning a fast failure instead of waiting on a call that's likely to fail or time out anyway. This is the same idea as an electrical circuit breaker: rather than letting a fault keep drawing power and causing more damage, it trips and cuts the connection until conditions are safe to try again. The value is in the fast failure: a caller that fails in ten milliseconds can degrade gracefully, while a caller stuck waiting the full timeout on every request has no room to do anything but wait.
What a Bulkhead Does That a Circuit Breaker Doesn't
A circuit breaker protects against a dependency that's actively failing. A bulkhead protects against a dependency that's just slow, by isolating the resources, threads, connection pool slots, used to call it so that dependency can't exhaust resources shared with everything else. Named after the watertight compartments in a ship's hull that stop one breach from flooding the whole vessel, a bulkhead means a slow third-party API can only consume its own dedicated slice of your capacity, not your entire connection pool. Without this isolation, a circuit breaker still helps once it trips, but the damage during the time it takes to detect the problem can already be done.
Setting Thresholds That Actually Match Reality
A circuit breaker's failure-rate threshold and cooldown period need to be based on the dependency's actual normal behavior, not a generic default. A threshold set too sensitive trips on ordinary transient blips and adds unnecessary fast-failures during totally normal operation. A threshold set too loose lets real cascading failures run for too long before tripping. Pull a few weeks of real latency and error data for the dependency in question, and set the threshold a deliberate distance above that baseline, the same principle that applies to alert thresholds generally.
The Half-Open State Is Where Most Implementations Get Sloppy
After a cooldown period, a circuit breaker should move to a half-open state, allowing a small number of test requests through to check whether the dependency has recovered, rather than immediately resuming full traffic. Skipping this step and just reopening the circuit fully the moment the cooldown ends risks slamming a barely-recovered dependency with your full traffic volume again immediately, which can retrigger the same failure. Allow a handful of test requests through first, and only fully reopen once those succeed.
What Happens on the Caller's Side When the Breaker Trips
A tripped circuit breaker needs a fallback behavior defined ahead of time, not improvised in the moment. Depending on the dependency, that might mean serving cached or default data, degrading a feature gracefully instead of failing the whole request, or queuing the request for later instead of dropping it. A circuit breaker with no fallback plan just turns a slow failure into a fast one, which is real progress, but pairing it with an actual fallback is what turns a dependency outage into a degraded experience instead of a full one for your users.
A resilient call to an external dependency has these parts:
- A circuit breaker that trips on failure rate or latency and returns a fast failure during a cooldown period.
- A bulkhead that isolates threads or connection pool slots so a slow dependency cannot exhaust resources shared with everything else.
- A half-open state that lets a few test requests through before full traffic resumes against a barely recovered dependency.
- A fallback defined ahead of time, such as cached or default data, a gracefully degraded feature, or a queued request.
- Thresholds and cooldowns based on the dependency's normal behavior, not generic defaults.
A Worked Example: A Shipping-Rate API Going Down
Say your checkout flow calls a third-party shipping-rate API to calculate delivery cost, and that API starts timing out under its own load. Without a circuit breaker, every checkout request waits the full timeout before failing, and if the rate lookup shares a thread pool with the rest of checkout, unrelated requests start queuing behind it too. With a circuit breaker and a bulkhead in place, the breaker trips after a handful of failures, checkout requests get a fast fallback, a flat estimated rate instead of a live one, and the isolated thread pool means the slow API never touches capacity the rest of checkout needs. Customers still complete their purchase, just with a slightly less precise shipping estimate, instead of not completing it at all.
What Good Looks Like
A resilient dependency setup pairs a circuit breaker tuned to the dependency's real baseline latency and failure rate with resource isolation so one slow call can't exhaust capacity shared by everything else, plus a defined fallback for when the breaker trips.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need both a circuit breaker and a bulkhead, or does one cover the other?
They cover different failure modes and work best together. A bulkhead limits how much of your resources a slow dependency can consume while it's still being called; a circuit breaker stops calling it altogether once it's clearly failing. Using only one leaves a gap the other was built to close.
How do we pick a cooldown period for a circuit breaker?
Base it on how long the dependency typically takes to recover from the kind of failure you're protecting against, often somewhere between a few seconds and a minute for most transient issues. Too short and you're retrying against a dependency that hasn't actually recovered; too long and you're needlessly failing requests after the dependency is already healthy again.
Should every external dependency have a circuit breaker?
It's worth it for any dependency whose failure could cascade into unrelated parts of your system, which is most external calls in a request path that other features depend on. It matters less for a background job with generous retry tolerance and no direct user-facing impact if it's delayed.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Circuit Breakers and Bulkheads, Explained With Checkout
A slow payment provider times out, threads pile up waiting, and the whole service goes unresponsive. How circuit breakers and bulkheads contain that.
The Circuit Breaker Checklist Most Teams Skip Half Of
A checklist for implementing circuit breakers and bulkheads correctly: the failure thresholds, half-open behavior, and isolation mistakes teams miss.
Circuit Breakers and Bulkheads: Configuring Them So They Help
How to set trip thresholds against your real availability target, contain failures with bulkheads, and avoid the mistake of one setting for every call.
When a Circuit Breaker Helps, and When It Just Hides a Bug
A decision guide for using circuit breakers and bulkheads to stop one failing dependency from cascading, and where a circuit breaker can mask a real problem.
Circuit Breakers and Bulkheads: Stopping a Failure From Spreading
How circuit breakers and bulkhead isolation stop one failing dependency from taking down a whole service, and the tuning mistakes that hurt.
Circuit Breakers and Bulkheads: What Each One Actually Prevents
A comparison of circuit breaker and bulkhead resilience patterns: what failure each one actually prevents, and how to set thresholds that work.