AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

Circuit Breakers and Bulkheads: Stopping One Outage Becoming Three

A dependency that's completely down is, in a strange way, the easy case: calls fail immediately, and your app can react. A dependency that's slow, timing out inconsistently, or failing intermittently is the dangerous one, because every caller keeps waiting on it, and that waiting is what turns one struggling service into a cascading outage across everything that calls it.

Circuit breakers and bulkheads exist for exactly that failure mode: stop calling a dependency that's clearly in trouble, and keep one bad dependency from starving the resources the rest of your app needs.

What a circuit breaker actually prevents

Without a breaker, every caller to a slow dependency waits out its own timeout before giving up, and while it waits, it's holding a connection, a thread, or a resource from whatever pool it borrowed from. Enough concurrent callers stuck waiting exhausts that pool, and now requests that have nothing to do with the failing dependency start failing too, because there's nothing left in the pool to serve them.

A tight uptime target leaves very little room for that kind of cascade to play out before it's a real incident, since the downtime budget behind a high availability target is measured in minutes a year, not hours1. A circuit breaker interrupts the pattern by failing fast once a dependency is clearly unhealthy, instead of letting every caller queue up to discover that themselves.

The three states, and the one people get wrong

A breaker has three states: closed, calls go through normally while the dependency looks healthy; open, calls fail immediately without even attempting the dependency, once failures cross a threshold; and half-open, a trial state that lets a small number of calls through to test whether the dependency has recovered.

Half-open is the state teams most often misconfigure. Let too much traffic through during the test and a dependency that's only partially recovered gets hit with a full load again, tripping the breaker straight back to open before it had a real chance to stabilize. Keep the trial volume small and the threshold to fully close deliberate, not a single lucky successful call.

Bulkheads: isolating one bad dependency from the rest

A bulkhead gives each downstream dependency its own connection or thread pool instead of sharing one pool across all of them. When one dependency slows down, it exhausts its own allocation and its own alone, leaving the resources the rest of your application needs to talk to healthy dependencies untouched.

Without that isolation, a single struggling dependency can starve calls to completely unrelated, perfectly healthy services simply because they were all drawing from the same shared pool. Bulkheads and circuit breakers solve related but different problems, and most resilient systems need both: the breaker decides whether to call a dependency at all, the bulkhead limits the blast radius when something goes wrong anyway.

Timeouts and retries: tune them together

A retry policy without a sane timeout underneath it just multiplies load on a dependency that's already struggling, each failed attempt retried adds more requests to something that couldn't handle the load it already had. Set a timeout that's meaningfully shorter than whatever is calling you is willing to wait, and use exponential backoff with jitter between retry attempts so retries from many callers don't all land at once and create a new spike.

These settings interact with each other and with the breaker's own thresholds, so tune them as a set, not independently: a generous timeout paired with aggressive retries can keep a breaker from ever tripping even while the dependency is clearly unhealthy, because from the breaker's view, calls are eventually succeeding, just very slowly.

Testing a breaker before a real outage does it for you

The only way to know a breaker actually works is to watch it trip. In a lower environment, deliberately fail a dependency, block its port, return errors from a test double, kill the process, and confirm the breaker opens, calls fail fast instead of hanging, and the application degrades gracefully rather than cascading.

Watch what "degrades gracefully" actually looks like for your users during that test, a fallback response, a cached value, a clearly disabled feature, rather than assuming the breaker tripping automatically means the user experience is acceptable. Those are two different things, and only one of them is guaranteed by the breaker configuration alone.

A simple drill to prove the breaker works:

  1. In a lower environment, deliberately fail a dependency by blocking its port, returning errors from a test double, or killing the process.
  2. Confirm the breaker opens once failures cross its threshold, so calls fail immediately instead of hanging.
  3. Check that half-open lets only a small number of trial calls through, and that the breaker closes again only when they succeed.
  4. Watch what graceful degradation looks like for your application, and confirm calls to healthy dependencies keep working through their own bulkhead pools.
Executive Capability Standard

What Good Looks Like

Good here means a slow or failing dependency degrades one feature instead of taking down requests that never touched it, because failures are isolated and calls fail fast instead of piling up.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map which of your downstream dependencies currently have no timeout, retry, or fallback, since that's where a breaker or bulkhead is worth adding first.
2. Do Manually:Add a hard timeout by hand to your riskiest downstream call this week, even before building a full circuit breaker around it.
3. Delegate:Have someone own resilience patterns for the two or three dependencies most likely to cause a cascading outage, since they matter more than the rest combined.
4. Automate:Add circuit breakers and separate connection pools per dependency in your service framework, so new code gets the protection by default.
5. Buy:A resilience library or service mesh feature that implements circuit breaking and bulkheads out of the box can save you from maintaining that logic yourself across every service.

How to Get Started

Frequently Asked Questions

Does a circuit breaker replace retries?

No, they work together. The breaker decides whether it's even worth attempting a call to a dependency that's clearly unhealthy. Retries handle transient failures once the breaker has decided the call is worth attempting in the first place.

What should happen when a breaker is open?

Return a fast, honest failure or a fallback immediately, a cached value, a degraded version of the feature, rather than letting the caller hang waiting for something that's been failing. A quick, clear failure is almost always better than a slow, uncertain one.

How do I pick a timeout value?

Base it on what's actually calling you and how long that caller is willing to wait, then set your timeout meaningfully shorter than that. An arbitrary round number that ignores the caller's own patience just moves the problem one layer up the stack.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides