Model Context Protocol & Agentic ArchitecturePlaybook4 min readUpdated September 2026

Circuit Breakers and Bulkheads: What Each One Actually Prevents

A circuit breaker stops calls to a dependency that is already failing, while a bulkhead keeps one dependency from using up the capacity every other call needs. A single slow or failing downstream call can exhaust a healthy service's threads or connections, and implementing only one pattern leaves you exposed to the other half of that failure.

This compares what each pattern actually prevents, how to set a circuit breaker's thresholds so it trips at the right time, and how the two work together without adding more complexity than the failure mode actually justifies.

The Failure Both Patterns Are Trying to Stop

Cascading failure starts small: one dependency gets slow or starts erroring, and calls to it start taking longer to fail or time out than they normally would. If nothing intervenes, every request that depends on that call eventually piles up waiting on it, consuming threads, connections, or memory until the calling service itself becomes unresponsive, even though the calling service's own code has no bug in it.

Both circuit breakers and bulkheads exist to contain this before it spreads. The difference is what each one actually isolates: a circuit breaker stops sending calls to a dependency that's already shown it's failing, while a bulkhead limits how much of your service's own capacity any single dependency can consume, whether or not that dependency has failed yet.

Circuit Breakers: Stopping Calls to a Failing Dependency

A circuit breaker tracks the failure rate of calls to a specific dependency, and once that rate crosses a threshold, it stops sending new calls entirely for a cooldown period, failing fast instead of letting new requests queue up waiting on a dependency that's already known to be failing. After the cooldown, it allows a small number of test calls through to check whether the dependency has recovered, and only fully reopens if those succeed.

The real value is failing fast rather than failing slow. A dependency that times out slowly on every call is far more damaging to your own service's capacity than one that fails instantly, since the timeout duration is exactly how long each of your own threads or connections stays tied up waiting. A circuit breaker converts that slow failure into a fast one once it's tripped.

How do you set circuit breaker thresholds that trip on time?

A threshold that's too sensitive trips on ordinary, brief blips, adding unnecessary failures to requests that would have succeeded if they'd just been retried normally. A threshold that's too lenient lets a real, sustained failure keep consuming capacity for too long before the breaker finally trips and starts protecting the rest of the service.

Base the threshold on your dependency's normal baseline error rate, measured over a real observation period, not a round number picked without data. Require a minimum volume of calls before the breaker can trip at all, so a handful of failures during genuinely low traffic doesn't trip the breaker on a sample too small to mean anything. And tune the cooldown period against how long the dependency actually tends to take to recover from a real outage, not an arbitrary short window that reopens the breaker before the dependency has actually stabilized.

Bulkheads: Stopping One Dependency From Starving the Rest

A bulkhead limits the resources, connection pool slots, thread pool capacity, concurrent request limits, that any single downstream dependency can consume, so a problem calling one dependency can't exhaust the resources your service needs to keep calling its other, healthy dependencies. Without a bulkhead, a single slow dependency can starve calls to every other dependency your service depends on, even ones that are working fine.

This matters even when a circuit breaker is also in place, because the circuit breaker only trips after enough failures accumulate; during the window before it trips, an unbounded number of concurrent calls to the failing dependency can still exhaust shared resources. A bulkhead caps that concurrent call count directly, containing the damage during exactly the window a circuit breaker hasn't reacted yet.

How do you use both patterns without overcomplicating things?

Most services don't need every dependency wrapped in both patterns with heavily customized thresholds. Apply them where the failure mode actually justifies the complexity:

  • Add a circuit breaker to any dependency whose failure could otherwise generate a large volume of slow, timing-out calls, especially anything without a short, reliable timeout of its own.
  • Add a bulkhead specifically for dependencies that share a resource pool with other, unrelated calls, so one dependency's problems can't starve capacity meant for something else.
  • Start with sensible defaults for both patterns rather than hand-tuning every dependency individually; refine thresholds only for the dependencies that actually show problems in practice.
  • Don't wrap a dependency in both patterns if it already has a short, reliable timeout and doesn't share resources with anything else; the added complexity isn't buying you protection against a failure mode that can't really happen there.

Match the pattern to the actual failure mode each dependency can cause, rather than applying both everywhere by default, which adds configuration surface area without a proportional reduction in real risk.

Executive Capability Standard

What Good Looks Like

Good resilience patterning means circuit breaker thresholds and bulkhead limits are set from a dependency's real observed behavior, and applied only where the specific failure mode each pattern prevents can actually happen.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map out which of your service's downstream dependencies lack a short, reliable timeout, since those are the strongest candidates for a circuit breaker.
2. Do Manually:Manually set a conservative circuit breaker threshold on your riskiest dependency first, based on its observed baseline error rate, and watch it in production before adding more.
3. Delegate:Have whoever owns each service decide which of its dependencies need a circuit breaker, a bulkhead, both, or neither, rather than applying one blanket policy everywhere.
4. Automate:Add alerting on circuit breaker state changes specifically, so a breaker tripping in production surfaces as a signal, not just a silent fallback that masks a real dependency outage.
5. Buy:Use a resilience library or service mesh feature that implements both patterns rather than hand-rolling the state machine and thresholds yourself, unless your failure handling needs are unusual enough to require custom logic.

How to Get Started

Frequently Asked Questions

Do circuit breakers and bulkheads solve the same problem?

No, they defend against different parts of a cascading failure. A circuit breaker stops sending calls to a dependency that's already shown it's failing. A bulkhead limits how much of your own capacity any single dependency can consume, whether or not it has failed yet, which matters during the window before a circuit breaker trips.

How do we avoid a circuit breaker that trips too easily?

Base the failure rate threshold on your dependency's real baseline error rate instead of a round number. Also require a minimum call volume before the breaker can trip, so a handful of failures during low traffic doesn't open it on too small a sample to mean anything.

Do we need both patterns on every downstream dependency?

No. Apply a circuit breaker to dependencies that could otherwise generate a lot of slow, timing-out calls, and a bulkhead to dependencies that share a resource pool with other calls. A dependency with a short, reliable timeout and no shared resources often doesn't need either.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides