Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Circuit Breakers and Bulkheads, Explained With Checkout

A downstream payment provider starts timing out. Every thread handling a request that touches payments piles up waiting on it, and within minutes the entire service is unresponsive, even to requests that have nothing to do with payments at all, because nothing isolated the failure to just that one dependency.

A circuit breaker and a bulkhead solve two related but different problems, and most incidents like this one are missing at least one of them.

What a Circuit Breaker Actually Does

After a threshold of failures or timeouts against a dependency, the breaker opens and starts failing fast on calls to it for a cooldown period, instead of letting every caller wait out the full timeout on a dependency that's already known to be failing. This protects your own service's limited resources, threads and connections, from being consumed waiting on something that isn't going to answer anyway.

The Half-Open State Most Implementations Get Wrong

After the cooldown period, a well-behaved breaker lets a small number of trial requests through to test whether the dependency has actually recovered, rather than either staying open forever or slamming straight back to fully closed. A common bug is a breaker that reopens the instant that very first trial request fails, which under any residual flakiness in the dependency means it effectively never recovers on its own.

Bulkheads: Isolating Resources So One Dependency Can't Starve Everything

A bulkhead gives each downstream dependency its own bounded pool of threads or connections, so a stuck call to one dependency can exhaust its own pool without touching the pool serving unrelated requests. Without that isolation, threads blocked waiting on a slow payment provider are the exact same threads needed to serve a completely unrelated product page request, which is how one dependency's problem becomes everyone's problem.

Setting Thresholds That Reflect Real Failure, Not Normal Variance

A threshold that's too sensitive trips the breaker on ordinary latency variance, opening it for a dependency that was actually fine and unnecessarily degrading a feature that didn't need to degrade. A threshold that's too loose never trips before your own service is already in real trouble. Base it on the dependency's own observed failure rate and latency distribution from real traffic, not a number that sounded reasonable in a planning meeting.

A useful decision rule is to set each breaker from that dependency's own history. Look at how often it times out during a normal week and how slow its responses get at busy hours, then set the trip point clearly beyond that normal variation. For example, a tax calculation service that regularly slows down at month end should not share a threshold with a payment provider that is normally steady. A common mistake is copying one threshold across every dependency. Review each threshold after any incident where a breaker tripped late, or tripped for no good reason.

A Worked Example

A checkout service calls three downstream dependencies: inventory, payments, and tax calculation. Each gets its own breaker and its own bulkhead. When payments starts timing out, its breaker opens and its bulkhead's threads stay exhausted only within that isolated pool. Inventory and tax calls, along with the parts of checkout that don't depend on payment confirmation, keep serving normally while payments degrades gracefully behind a clear fallback message instead of taking the entire checkout flow down with it.

Where This Pattern Gets Skipped

  • Applying a circuit breaker without a bulkhead, so failures get detected correctly but shared resources can still be starved by the failing dependency.
  • Setting one global timeout for every dependency instead of one per dependency, so a fast dependency's timeout ends up either too aggressive or too lax for a genuinely slower one.
  • Never deliberately testing the open and half-open states before you need them, so the first real test of the whole mechanism happens live, during an actual incident.

What the Caller Should See When a Breaker Is Open

A raw error thrown all the way up to the end user is rarely the right fallback. For checkout specifically, a clear message that payment processing is temporarily unavailable, with an obvious way to retry shortly, treats the customer better than a generic failure and buys you the cooldown window without losing the sale outright.

Design the fallback deliberately per dependency instead of letting whatever exception happens to surface become the user-facing message by accident. A payments outage and an inventory-lookup outage warrant genuinely different fallback behavior, and that difference is worth designing on purpose rather than discovering live during the next incident.

Executive Capability Standard

What Good Looks Like

Good resilience patterns mean one failing dependency degrades only the feature that depends on it, with its own isolated resource pool and its own breaker, while the rest of the service keeps serving requests normally.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map which of your downstream dependencies currently share a thread or connection pool with unrelated request paths.
2. Do Manually:Manually add a bulkhead to your highest-risk shared dependency call before touching breaker thresholds anywhere else.
3. Delegate:Assign one engineer to own breaker and bulkhead configuration for each major downstream dependency.
4. Automate:Automate a chaos test that deliberately fails one dependency and confirms the breaker and bulkhead actually contain the blast radius.
5. Buy:Bring in a reliability specialist to review resource isolation across your critical path before your next major dependency integration.

How to Get Started

Frequently Asked Questions

What's the difference between a circuit breaker and just setting a timeout?

A timeout limits how long one call waits before giving up, but every subsequent call still tries and waits out the same timeout again. A circuit breaker remembers that the dependency is currently failing and skips the wait entirely for a cooldown period, protecting your own threads and connections from repeatedly waiting on something already known to be down.

Why do I need a bulkhead if I already have a circuit breaker?

A circuit breaker stops new calls to a failing dependency, but calls already in flight when it opens can still hold onto shared resources until they time out. A bulkhead ensures those resources are scoped to that one dependency in the first place, so even before the breaker trips, a stuck call can't starve requests that don't depend on it.

How do you pick a failure threshold that trips the breaker at the right time?

Base it on the dependency's own observed failure rate and latency distribution under real traffic, not a round number chosen in advance. Too sensitive and you'll trip the breaker on normal variance; too loose and your own service is already struggling before it ever opens. Tune it against real data, then revisit it periodically.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides