Circuit Breakers and Bulkheads: Stopping One Outage Becoming Three
A dependency that's completely down is, in a strange way, the easy case: calls fail immediately, and your app can react. A dependency that's slow, timing out inconsistently, or failing intermittently is the dangerous one, because every caller keeps waiting on it, and that waiting is what turns one struggling service into a cascading outage across everything that calls it.
Circuit breakers and bulkheads exist for exactly that failure mode: stop calling a dependency that's clearly in trouble, and keep one bad dependency from starving the resources the rest of your app needs.
What a circuit breaker actually prevents
Without a breaker, every caller to a slow dependency waits out its own timeout before giving up, and while it waits, it's holding a connection, a thread, or a resource from whatever pool it borrowed from. Enough concurrent callers stuck waiting exhausts that pool, and now requests that have nothing to do with the failing dependency start failing too, because there's nothing left in the pool to serve them.
A tight uptime target leaves very little room for that kind of cascade to play out before it's a real incident, since the downtime budget behind a high availability target is measured in minutes a year, not hours1. A circuit breaker interrupts the pattern by failing fast once a dependency is clearly unhealthy, instead of letting every caller queue up to discover that themselves.
The three states, and the one people get wrong
A breaker has three states: closed, calls go through normally while the dependency looks healthy; open, calls fail immediately without even attempting the dependency, once failures cross a threshold; and half-open, a trial state that lets a small number of calls through to test whether the dependency has recovered.
Half-open is the state teams most often misconfigure. Let too much traffic through during the test and a dependency that's only partially recovered gets hit with a full load again, tripping the breaker straight back to open before it had a real chance to stabilize. Keep the trial volume small and the threshold to fully close deliberate, not a single lucky successful call.
Bulkheads: isolating one bad dependency from the rest
A bulkhead gives each downstream dependency its own connection or thread pool instead of sharing one pool across all of them. When one dependency slows down, it exhausts its own allocation and its own alone, leaving the resources the rest of your application needs to talk to healthy dependencies untouched.
Without that isolation, a single struggling dependency can starve calls to completely unrelated, perfectly healthy services simply because they were all drawing from the same shared pool. Bulkheads and circuit breakers solve related but different problems, and most resilient systems need both: the breaker decides whether to call a dependency at all, the bulkhead limits the blast radius when something goes wrong anyway.
Timeouts and retries: tune them together
A retry policy without a sane timeout underneath it just multiplies load on a dependency that's already struggling, each failed attempt retried adds more requests to something that couldn't handle the load it already had. Set a timeout that's meaningfully shorter than whatever is calling you is willing to wait, and use exponential backoff with jitter between retry attempts so retries from many callers don't all land at once and create a new spike.
These settings interact with each other and with the breaker's own thresholds, so tune them as a set, not independently: a generous timeout paired with aggressive retries can keep a breaker from ever tripping even while the dependency is clearly unhealthy, because from the breaker's view, calls are eventually succeeding, just very slowly.
Testing a breaker before a real outage does it for you
The only way to know a breaker actually works is to watch it trip. In a lower environment, deliberately fail a dependency, block its port, return errors from a test double, kill the process, and confirm the breaker opens, calls fail fast instead of hanging, and the application degrades gracefully rather than cascading.
Watch what "degrades gracefully" actually looks like for your users during that test, a fallback response, a cached value, a clearly disabled feature, rather than assuming the breaker tripping automatically means the user experience is acceptable. Those are two different things, and only one of them is guaranteed by the breaker configuration alone.
A simple drill to prove the breaker works:
- In a lower environment, deliberately fail a dependency by blocking its port, returning errors from a test double, or killing the process.
- Confirm the breaker opens once failures cross its threshold, so calls fail immediately instead of hanging.
- Check that half-open lets only a small number of trial calls through, and that the breaker closes again only when they succeed.
- Watch what graceful degradation looks like for your application, and confirm calls to healthy dependencies keep working through their own bulkhead pools.
What Good Looks Like
Good here means a slow or failing dependency degrades one feature instead of taking down requests that never touched it, because failures are isolated and calls fail fast instead of piling up.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Does a circuit breaker replace retries?
No, they work together. The breaker decides whether it's even worth attempting a call to a dependency that's clearly unhealthy. Retries handle transient failures once the breaker has decided the call is worth attempting in the first place.
What should happen when a breaker is open?
Return a fast, honest failure or a fallback immediately, a cached value, a degraded version of the feature, rather than letting the caller hang waiting for something that's been failing. A quick, clear failure is almost always better than a slow, uncertain one.
How do I pick a timeout value?
Base it on what's actually calling you and how long that caller is willing to wait, then set your timeout meaningfully shorter than that. An arbitrary round number that ignores the caller's own patience just moves the problem one layer up the stack.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Circuit Breakers and Bulkheads, Explained With Checkout
A slow payment provider times out, threads pile up waiting, and the whole service goes unresponsive. How circuit breakers and bulkheads contain that.
Circuit Breakers and Bulkheads: What Each One Actually Prevents
A comparison of circuit breaker and bulkhead resilience patterns: what failure each one actually prevents, and how to set thresholds that work.
The Circuit Breaker Checklist Most Teams Skip Half Of
A checklist for implementing circuit breakers and bulkheads correctly: the failure thresholds, half-open behavior, and isolation mistakes teams miss.
Redis, Postgres, or etcd: Choosing a Distributed Lock
A comparison of Redis locks, Postgres advisory locks, and etcd or ZooKeeper for coordinating work across multiple instances of a service.
Circuit Breakers and Bulkheads: Stopping One Bad Dependency
How circuit breakers and bulkheads stop one slow dependency from cascading into a full outage, plus the thresholds and pitfalls that make them work.
Circuit Breakers and Bulkheads: Configuring Them So They Help
How to set trip thresholds against your real availability target, contain failures with bulkheads, and avoid the mistake of one setting for every call.