Circuit Breakers and Bulkheads, Explained With Checkout
A downstream payment provider starts timing out. Every thread handling a request that touches payments piles up waiting on it, and within minutes the entire service is unresponsive, even to requests that have nothing to do with payments at all, because nothing isolated the failure to just that one dependency.
A circuit breaker and a bulkhead solve two related but different problems, and most incidents like this one are missing at least one of them.
What a Circuit Breaker Actually Does
After a threshold of failures or timeouts against a dependency, the breaker opens and starts failing fast on calls to it for a cooldown period, instead of letting every caller wait out the full timeout on a dependency that's already known to be failing. This protects your own service's limited resources, threads and connections, from being consumed waiting on something that isn't going to answer anyway.
The Half-Open State Most Implementations Get Wrong
After the cooldown period, a well-behaved breaker lets a small number of trial requests through to test whether the dependency has actually recovered, rather than either staying open forever or slamming straight back to fully closed. A common bug is a breaker that reopens the instant that very first trial request fails, which under any residual flakiness in the dependency means it effectively never recovers on its own.
Bulkheads: Isolating Resources So One Dependency Can't Starve Everything
A bulkhead gives each downstream dependency its own bounded pool of threads or connections, so a stuck call to one dependency can exhaust its own pool without touching the pool serving unrelated requests. Without that isolation, threads blocked waiting on a slow payment provider are the exact same threads needed to serve a completely unrelated product page request, which is how one dependency's problem becomes everyone's problem.
Setting Thresholds That Reflect Real Failure, Not Normal Variance
A threshold that's too sensitive trips the breaker on ordinary latency variance, opening it for a dependency that was actually fine and unnecessarily degrading a feature that didn't need to degrade. A threshold that's too loose never trips before your own service is already in real trouble. Base it on the dependency's own observed failure rate and latency distribution from real traffic, not a number that sounded reasonable in a planning meeting.
A useful decision rule is to set each breaker from that dependency's own history. Look at how often it times out during a normal week and how slow its responses get at busy hours, then set the trip point clearly beyond that normal variation. For example, a tax calculation service that regularly slows down at month end should not share a threshold with a payment provider that is normally steady. A common mistake is copying one threshold across every dependency. Review each threshold after any incident where a breaker tripped late, or tripped for no good reason.
A Worked Example
A checkout service calls three downstream dependencies: inventory, payments, and tax calculation. Each gets its own breaker and its own bulkhead. When payments starts timing out, its breaker opens and its bulkhead's threads stay exhausted only within that isolated pool. Inventory and tax calls, along with the parts of checkout that don't depend on payment confirmation, keep serving normally while payments degrades gracefully behind a clear fallback message instead of taking the entire checkout flow down with it.
Where This Pattern Gets Skipped
- Applying a circuit breaker without a bulkhead, so failures get detected correctly but shared resources can still be starved by the failing dependency.
- Setting one global timeout for every dependency instead of one per dependency, so a fast dependency's timeout ends up either too aggressive or too lax for a genuinely slower one.
- Never deliberately testing the open and half-open states before you need them, so the first real test of the whole mechanism happens live, during an actual incident.
What the Caller Should See When a Breaker Is Open
A raw error thrown all the way up to the end user is rarely the right fallback. For checkout specifically, a clear message that payment processing is temporarily unavailable, with an obvious way to retry shortly, treats the customer better than a generic failure and buys you the cooldown window without losing the sale outright.
Design the fallback deliberately per dependency instead of letting whatever exception happens to surface become the user-facing message by accident. A payments outage and an inventory-lookup outage warrant genuinely different fallback behavior, and that difference is worth designing on purpose rather than discovering live during the next incident.
What Good Looks Like
Good resilience patterns mean one failing dependency degrades only the feature that depends on it, with its own isolated resource pool and its own breaker, while the rest of the service keeps serving requests normally.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the difference between a circuit breaker and just setting a timeout?
A timeout limits how long one call waits before giving up, but every subsequent call still tries and waits out the same timeout again. A circuit breaker remembers that the dependency is currently failing and skips the wait entirely for a cooldown period, protecting your own threads and connections from repeatedly waiting on something already known to be down.
Why do I need a bulkhead if I already have a circuit breaker?
A circuit breaker stops new calls to a failing dependency, but calls already in flight when it opens can still hold onto shared resources until they time out. A bulkhead ensures those resources are scoped to that one dependency in the first place, so even before the breaker trips, a stuck call can't starve requests that don't depend on it.
How do you pick a failure threshold that trips the breaker at the right time?
Base it on the dependency's own observed failure rate and latency distribution under real traffic, not a round number chosen in advance. Too sensitive and you'll trip the breaker on normal variance; too loose and your own service is already struggling before it ever opens. Tune it against real data, then revisit it periodically.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
The Circuit Breaker Checklist Most Teams Skip Half Of
A checklist for implementing circuit breakers and bulkheads correctly: the failure thresholds, half-open behavior, and isolation mistakes teams miss.
Deciding Where a Circuit Breaker Actually Belongs in Your Pipeline
A decision guide for where circuit breakers and bulkhead isolation genuinely prevent cascading failure, and where they just add complexity without benefit.
Circuit Breakers and Bulkheads: Configuring Them So They Help
How to set trip thresholds against your real availability target, contain failures with bulkheads, and avoid the mistake of one setting for every call.
Circuit Breakers and Bulkheads: Stopping One Bad Dependency
How circuit breakers and bulkheads stop one slow dependency from cascading into a full outage, plus the thresholds and pitfalls that make them work.
When a Circuit Breaker Helps, and When It Just Hides a Bug
A decision guide for using circuit breakers and bulkheads to stop one failing dependency from cascading, and where a circuit breaker can mask a real problem.
Circuit Breakers and Bulkheads: Stopping a Failure From Spreading
How circuit breakers and bulkhead isolation stop one failing dependency from taking down a whole service, and the tuning mistakes that hurt.