Circuit Breakers and Bulkheads: Configuring Them So They Help
A circuit breaker that trips too eagerly turns a brief blip into an outage of its own, and one that never trips lets a struggling downstream dependency drag your whole service down with it. Getting the configuration right matters more than the pattern itself, and it starts with being specific about what each breaker is actually protecting.
What a circuit breaker is actually deciding
A breaker trips based on an error rate measured over a rolling window, not a single failed request, and moves into an open state where it fails fast instead of calling the downstream dependency at all. After a cooldown period it moves to a half-open state, letting a small number of probe requests through to check whether the dependency has recovered before fully closing again.
The rolling window size matters as much as the error rate threshold. Too short a window and normal noise trips it constantly; too long a window and it's slow to react to a dependency that's actually degrading.
Pick a window long enough to smooth over a single bad second of traffic but short enough that a genuinely failing dependency trips the breaker before its errors pile up into a bigger incident. There's no universal number here; it depends on your request volume against that specific dependency.
Bulkheads: containing the blast radius the breaker can't
A circuit breaker stops you from calling a failing dependency, but if that dependency shares a connection pool or thread pool with everything else your service does, a slow dependency can exhaust the shared pool before the breaker's error-rate threshold even trips, since slow isn't the same as failing.
Separate pools per downstream dependency, the bulkhead pattern, keep one slow dependency from starving requests that have nothing to do with it. Size each pool to what that specific dependency actually needs, not to a single shared number that assumes every dependency behaves the same way.
Tuning thresholds against your actual availability target
If you're holding yourself to three nines of availability, work backward from that downtime budget when deciding how long a breaker stays open before it probes again1.
A cooldown period that's too long keeps failing fast on a dependency that already recovered, quietly extending an outage past when it needed to end. A cooldown that's too short sends real traffic back at a dependency that hasn't actually stabilized, and can trip the breaker right back open.
The failure mode a breaker can hide: cascading timeouts upstream
Tripping a breaker doesn't make the underlying problem disappear; it moves the failure one hop up the call chain wearing a different name. If the caller has no fallback behavior for an open breaker, it just times out or errors instead, and whatever called it faces the exact same choice.
Design an explicit fallback for every breaker you add, whether that's cached data, a degraded response, or a clear error the caller can act on, rather than assuming an open breaker alone counts as handling the failure.
The mistake: setting the same threshold for every downstream call
A call to a non-critical enrichment service, one that makes a response nicer but not wrong if it's missing, tolerates a much more aggressive breaker than a call to your payments processor, where you'd rather wait a little longer for a real answer than fail fast into a degraded checkout flow.
Configure breakers per dependency based on what failing fast actually costs you there, not a single default copied across every integration in the codebase.
Testing a breaker without waiting for a real outage
A breaker that's never actually tripped in a rehearsed test is an untested piece of failure handling, no different from any other code path nobody has exercised. Fault injection, deliberately forcing a downstream call to fail or time out in a controlled environment, confirms the breaker trips at the threshold you configured and that the fallback you wrote actually runs.
Run this check whenever the threshold or cooldown changes, not only when the breaker is first added. A configuration change that looks harmless on paper can shift when it trips in ways that only show up once you force it.
A simple test routine covers both the breaker and its fallback:
- Use fault injection in a controlled environment to force a downstream call to fail or time out.
- Confirm the breaker trips at the error-rate threshold you configured, not sooner or later.
- Check that the fallback you wrote actually runs while the breaker is open, whether that is cached data or a degraded response.
- Watch the half-open probes after the cooldown and confirm the breaker closes once the dependency has recovered.
What Good Looks Like
Mature circuit breaker configuration means every downstream dependency has a threshold and cooldown set for what failing fast actually costs there, plus an explicit fallback, not a single default copied everywhere.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the difference between a circuit breaker and a simple retry with backoff?
A retry assumes the next attempt might succeed and keeps trying against the same dependency. A circuit breaker assumes a struggling dependency needs a break from traffic entirely, and stops calling it for a cooldown period. The two are often used together: retry a few times, then let the breaker trip if the error rate stays high.
How do I choose the error-rate threshold that trips the breaker?
Start from what a failure actually costs downstream of this specific call, not a single default across your codebase. A non-critical dependency can tolerate a more aggressive threshold than one your core flow depends on, so set thresholds per dependency rather than picking one number for everything.
Do circuit breakers help with slow responses, not just failures?
Only if slowness is counted toward the trip condition, since a breaker configured purely on hard failures won't react to a dependency that's technically succeeding but taking far too long. Many implementations let you treat a timeout as a failure for this reason, which is worth enabling explicitly.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Circuit Breakers and Bulkheads, Explained With Checkout
A slow payment provider times out, threads pile up waiting, and the whole service goes unresponsive. How circuit breakers and bulkheads contain that.
The Circuit Breaker Checklist Most Teams Skip Half Of
A checklist for implementing circuit breakers and bulkheads correctly: the failure thresholds, half-open behavior, and isolation mistakes teams miss.
Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)
A walkthrough of a real double-charge bug caused by a Redis lock's TTL expiring mid-job, and the fencing-token pattern that actually fixes it.
Circuit Breakers and Bulkheads: Stopping One Bad Dependency
How circuit breakers and bulkheads stop one slow dependency from cascading into a full outage, plus the thresholds and pitfalls that make them work.
When a Circuit Breaker Helps, and When It Just Hides a Bug
A decision guide for using circuit breakers and bulkheads to stop one failing dependency from cascading, and where a circuit breaker can mask a real problem.
Circuit Breakers and Bulkheads: Stopping a Failure From Spreading
How circuit breakers and bulkhead isolation stop one failing dependency from taking down a whole service, and the tuning mistakes that hurt.