Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

Circuit Breakers and Bulkheads: Stopping One Bad Dependency

A single slow dependency, a third-party API taking ten seconds to respond instead of two hundred milliseconds, shouldn't be able to take down your entire application. In practice, it often does, because requests pile up waiting on that one slow call, threads or connections get exhausted, and the failure spreads to parts of your system that have nothing to do with the original dependency.

Circuit breakers and bulkheads are the two patterns built specifically to stop that spread, and they solve slightly different parts of the problem.

What a Circuit Breaker Actually Does

A circuit breaker watches the failure rate or latency of calls to a specific dependency, and once that rate crosses a threshold, it stops making new calls entirely for a cooldown period, returning a fast failure instead of waiting on a call that's likely to fail or time out anyway. This is the same idea as an electrical circuit breaker: rather than letting a fault keep drawing power and causing more damage, it trips and cuts the connection until conditions are safe to try again. The value is in the fast failure: a caller that fails in ten milliseconds can degrade gracefully, while a caller stuck waiting the full timeout on every request has no room to do anything but wait.

What a Bulkhead Does That a Circuit Breaker Doesn't

A circuit breaker protects against a dependency that's actively failing. A bulkhead protects against a dependency that's just slow, by isolating the resources, threads, connection pool slots, used to call it so that dependency can't exhaust resources shared with everything else. Named after the watertight compartments in a ship's hull that stop one breach from flooding the whole vessel, a bulkhead means a slow third-party API can only consume its own dedicated slice of your capacity, not your entire connection pool. Without this isolation, a circuit breaker still helps once it trips, but the damage during the time it takes to detect the problem can already be done.

Setting Thresholds That Actually Match Reality

A circuit breaker's failure-rate threshold and cooldown period need to be based on the dependency's actual normal behavior, not a generic default. A threshold set too sensitive trips on ordinary transient blips and adds unnecessary fast-failures during totally normal operation. A threshold set too loose lets real cascading failures run for too long before tripping. Pull a few weeks of real latency and error data for the dependency in question, and set the threshold a deliberate distance above that baseline, the same principle that applies to alert thresholds generally.

The Half-Open State Is Where Most Implementations Get Sloppy

After a cooldown period, a circuit breaker should move to a half-open state, allowing a small number of test requests through to check whether the dependency has recovered, rather than immediately resuming full traffic. Skipping this step and just reopening the circuit fully the moment the cooldown ends risks slamming a barely-recovered dependency with your full traffic volume again immediately, which can retrigger the same failure. Allow a handful of test requests through first, and only fully reopen once those succeed.

What Happens on the Caller's Side When the Breaker Trips

A tripped circuit breaker needs a fallback behavior defined ahead of time, not improvised in the moment. Depending on the dependency, that might mean serving cached or default data, degrading a feature gracefully instead of failing the whole request, or queuing the request for later instead of dropping it. A circuit breaker with no fallback plan just turns a slow failure into a fast one, which is real progress, but pairing it with an actual fallback is what turns a dependency outage into a degraded experience instead of a full one for your users.

A resilient call to an external dependency has these parts:

  • A circuit breaker that trips on failure rate or latency and returns a fast failure during a cooldown period.
  • A bulkhead that isolates threads or connection pool slots so a slow dependency cannot exhaust resources shared with everything else.
  • A half-open state that lets a few test requests through before full traffic resumes against a barely recovered dependency.
  • A fallback defined ahead of time, such as cached or default data, a gracefully degraded feature, or a queued request.
  • Thresholds and cooldowns based on the dependency's normal behavior, not generic defaults.

A Worked Example: A Shipping-Rate API Going Down

Say your checkout flow calls a third-party shipping-rate API to calculate delivery cost, and that API starts timing out under its own load. Without a circuit breaker, every checkout request waits the full timeout before failing, and if the rate lookup shares a thread pool with the rest of checkout, unrelated requests start queuing behind it too. With a circuit breaker and a bulkhead in place, the breaker trips after a handful of failures, checkout requests get a fast fallback, a flat estimated rate instead of a live one, and the isolated thread pool means the slow API never touches capacity the rest of checkout needs. Customers still complete their purchase, just with a slightly less precise shipping estimate, instead of not completing it at all.

Executive Capability Standard

What Good Looks Like

A resilient dependency setup pairs a circuit breaker tuned to the dependency's real baseline latency and failure rate with resource isolation so one slow call can't exhaust capacity shared by everything else, plus a defined fallback for when the breaker trips.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map your critical request paths and identify which external dependencies currently have no circuit breaker or resource isolation at all.
2. Do Manually:Add a basic circuit breaker with a conservative threshold to your highest-risk dependency as a first, manual pass.
3. Delegate:Have a backend engineer own tuning breaker thresholds from real latency and failure data, and defining a fallback behavior for each protected dependency.
4. Automate:Standardize circuit breaker and bulkhead configuration through a shared library so every new external call gets consistent protection by default.
5. Buy:Bring in an infrastructure specialist if you're designing resilience patterns across a large number of services and want a consistent approach applied from the start.

How to Get Started

Frequently Asked Questions

Do we need both a circuit breaker and a bulkhead, or does one cover the other?

They cover different failure modes and work best together. A bulkhead limits how much of your resources a slow dependency can consume while it's still being called; a circuit breaker stops calling it altogether once it's clearly failing. Using only one leaves a gap the other was built to close.

How do we pick a cooldown period for a circuit breaker?

Base it on how long the dependency typically takes to recover from the kind of failure you're protecting against, often somewhere between a few seconds and a minute for most transient issues. Too short and you're retrying against a dependency that hasn't actually recovered; too long and you're needlessly failing requests after the dependency is already healthy again.

Should every external dependency have a circuit breaker?

It's worth it for any dependency whose failure could cascade into unrelated parts of your system, which is most external calls in a request path that other features depend on. It matters less for a background job with generous retry tolerance and no direct user-facing impact if it's delayed.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides