API Security, Identity & Zero-TrustPlaybook3 min readUpdated September 2026

Circuit Breakers and Bulkheads: Stopping a Failure From Spreading

A single slow or failing dependency can take down an entire service that's otherwise healthy, if every request thread ends up blocked waiting on that one dependency to respond. Circuit breakers and bulkheads are two different answers to that same problem, and most resilient systems need both, not one or the other.

Getting either pattern's thresholds wrong is worse than not having them at all, since a misconfigured breaker either trips on normal traffic or never trips when it should.

What a circuit breaker actually does

A circuit breaker watches the failure rate of calls to a specific dependency, and once that rate crosses a threshold, it stops sending new calls entirely for a cooldown period, failing fast instead of letting every request queue up waiting on a dependency that's already struggling.

After the cooldown, it allows a small number of test requests through; if those succeed, it closes again and traffic resumes normally, and if they fail, it stays open and waits longer. The point isn't preventing the failure, it's preventing the failure from cascading into every other part of the system that's waiting on a response.

Bulkheads: isolating so one failure can't starve everything else

A bulkhead limits how much of a shared resource, a thread pool, a connection pool, concurrent request slots, any single dependency or feature can consume, so a slow or failing one can't starve every other feature of the same resource.

Without bulkheads, a circuit breaker on one dependency can still fail to help if every thread in your shared pool is already blocked waiting on that same slow dependency by the time the breaker trips. Partition your resource pools per dependency, or at least per criticality tier, so a checkout-blocking failure and a recommendations-widget failure aren't drawing from the same limited pool of threads.

Protecting the uptime budget these patterns exist for

A circuit breaker's whole job is making sure one flaky dependency doesn't spend your entire uptime budget at once, and at 99.95 percent availability that budget is only 4.38 hours a year1.

A dependency failure that's contained by a breaker and a bulkhead costs you a degraded feature for its cooldown period; the same failure without containment can cost you the whole service for as long as the dependency stays down, which is a very different draw against the same annual budget.

Tuning mistakes that make these patterns worse than nothing

  • A failure threshold set so low that normal, brief error-rate blips trip the breaker constantly, training engineers to ignore its alerts or disable it outright.
  • A cooldown period so short that the breaker flaps open and closed rapidly against a dependency that's genuinely struggling, adding overhead without actually protecting anything.
  • Sharing one thread pool across every dependency despite having circuit breakers configured per dependency, which defeats the isolation a bulkhead is supposed to provide.
  • No fallback behavior defined for when a breaker is open, so users get a hard error instead of a degraded but functional experience, like cached data or a simplified response.

A worked example: one dependency, contained

Say a third-party shipping-rate API starts timing out under its own load. Without a circuit breaker, every checkout request waits the full timeout duration for a response that never comes, threads pile up waiting, and the connection pool serving that call exhausts, which starts affecting unrelated requests sharing the same pool. Checkout, and possibly the whole service, goes down with a dependency that was never actually part of its core function.

With a circuit breaker and a bulkhead isolating the shipping-rate call to its own resource pool, the breaker trips after a handful of timeouts, checkout falls back to a flat-rate shipping estimate for the cooldown period, and every other part of the service, including checkout's core purchase flow, keeps working normally. The dependency failure cost a slightly less accurate shipping quote for a few minutes instead of a full outage.

Reviewing thresholds after every real trip

A circuit breaker's configuration shouldn't be a one-time setup task. Every time a breaker actually trips in production, that's real data about how that dependency fails, how long it typically takes to recover, and whether the current threshold and cooldown period matched what actually happened. Review the trip afterward the same way you'd review any incident: did the breaker trip too early, too late, or about right, and does the fallback behavior that activated actually serve users well, or just technically avoid an error.

Executive Capability Standard

What Good Looks Like

A good resilience setup pairs a tuned circuit breaker with bulkhead resource isolation per dependency, and defines a real fallback for when the breaker is open, so a contained failure degrades gracefully instead of just failing fast.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map your service's external and internal dependencies and identify which ones currently have no circuit breaker or resource isolation at all.
2. Do Manually:Manually add a circuit breaker library to your highest-risk dependency first, tuned conservatively, and review its trip history after a few weeks of real traffic.
3. Delegate:Assign a platform or reliability engineer to own resilience pattern standards across services, including reviewing threshold tuning after incidents.
4. Automate:Standardize on a shared circuit breaker and bulkhead library internally so every new dependency integration gets consistent, pre-tuned protection instead of ad hoc handling.
5. Buy:A reliability engineering consultant or fractional CTO is worth bringing in when a cascading failure has already caused a real incident and you want the resilience patterns designed properly the first time.

How to Get Started

Frequently Asked Questions

What's the difference between a circuit breaker and a bulkhead?

A circuit breaker stops sending requests to a dependency once its failure rate crosses a threshold, failing fast instead of queuing. A bulkhead limits how much of a shared resource, like a thread pool, any single dependency can consume, so one failing dependency can't starve every other feature of the same resource. Most resilient systems need both.

How do we pick the right failure threshold for a circuit breaker?

Start conservative, higher than your normal baseline error rate with margin, and tighten it based on real incident data rather than guessing upfront. A threshold set too low trips on ordinary blips and trains engineers to ignore or disable the breaker, which defeats the purpose.

What should happen when a circuit breaker is open?

Define a fallback, not just a hard error: cached data, a simplified response, or a clear degraded-mode message, depending on what the dependency provides. A breaker with no fallback behavior just moves the failure from slow and cascading to fast and total, which is better but still not the goal.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides