The Circuit Breaker Checklist Most Teams Skip Half Of
A circuit breaker sounds simple: stop calling a failing dependency, let it recover, resume calling it. In practice, most homegrown implementations get the easy majority of that right and skip the details that determine whether the breaker actually protects the system during a real cascading failure or just adds complexity that never quite works when it matters.
This is a checklist of the parts teams most commonly skip, organized by how much damage skipping each one actually causes.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Threshold tuning: too sensitive is almost as bad as too loose
A circuit breaker that trips on the first failure treats a single transient blip the same as a genuine outage, and will flap open and closed constantly against any dependency with even modest normal error noise, training the team to distrust or disable it. One that requires an unrealistically high failure count before tripping delays protection long enough that the cascading failure it exists to prevent has already started elsewhere in the system. Base the threshold on your dependency's actual observed error rate under normal conditions, not a default copied from a library's documentation example. Revisit that threshold periodically too, since a dependency's normal error rate drifts as its own traffic and infrastructure change over time.
How should a circuit breaker behave in the half-open state?
After a breaker trips and its cooldown period passes, it should move to a half-open state that allows a small number of test requests through to check whether the dependency has recovered, rather than either staying fully closed to all traffic or snapping immediately back to fully open. The most common mistake is letting all traffic flood back the instant the breaker reopens, which can immediately overwhelm a dependency that had only just started to recover, tripping the breaker again in a loop that never actually resolves. Limit half-open traffic to a small, fixed number of probe requests, and only fully reopen once a clear majority of those probes succeed.
Bulkheads: isolating one failure from taking down everything else
A circuit breaker protects against a single failing dependency; a bulkhead protects against that failing dependency consuming so many resources, threads, connections, memory, that it starves every other part of the system too, even parts with no relationship to the failure at all. Without bulkhead isolation, a slow or hanging call to one downstream service can exhaust your entire connection pool, and a customer trying to do something completely unrelated to the failing dependency gets a timeout anyway. Give each external dependency its own dedicated resource pool, sized deliberately, rather than sharing one pool across everything.
For example, a checkout service calls a tax provider, an address validator, and an email service through one shared pool of connections. When the email service hangs, requests pile up waiting on it and the pool empties, so tax calculation times out even though the tax provider is healthy. Splitting the pool into one per dependency, each sized to that dependency's importance, contains the damage: email calls back up in their own pool while checkout continues. Pair each pool with its own timeout so a slow call releases its slot quickly instead of holding it until the caller gives up.
What should your fallback be when a circuit breaker opens?
- Return cached or default data when the failing dependency is non-critical to the immediate request, a recommendation widget, a secondary data enrichment call, rather than failing the whole request over one optional piece.
- Queue for later when the call is a write that can tolerate delay, rather than blocking the user-facing request on it succeeding synchronously.
- Fail the request clearly when there's genuinely no safe fallback, a payment authorization, for instance, rather than silently proceeding with incomplete or stale data that could mislead a downstream decision.
Deciding this per dependency in advance, not improvising during an incident, is what actually makes a circuit breaker rollout pay off.
Observability: knowing the breaker tripped before a customer tells you
A circuit breaker with no dashboard or alert tied to its state changes protects the system silently, which sounds fine until a dependency has been degraded for hours and nobody on the team noticed because nothing ever paged. Emit a metric on every state transition, closed to open, open to half-open, half-open back to closed or back to open, and alert specifically on a breaker staying open past a reasonable window, since that's a strong signal of a real, ongoing outage worth a human's attention rather than a transient blip the breaker absorbed on its own.
What Good Looks Like
A resilient circuit breaker setup tunes thresholds against real observed error rates, limits half-open traffic to a small probe count, isolates each dependency with its own resource pool, and defines fallback behavior per dependency in advance rather than during an incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How many consecutive failures should trip a circuit breaker?
Base it on the dependency's actual normal error rate, not a fixed number copied from an example. Say a dependency normally runs a small baseline error rate day to day: it needs a meaningfully higher threshold than one that's normally near flawless, or the breaker will trip constantly on ordinary noise.
Do we need bulkheads if we already have circuit breakers?
Yes, they solve different problems. A circuit breaker stops calling a failing dependency; a bulkhead prevents that dependency from exhausting shared resources, like a connection pool, in a way that degrades unrelated parts of the system even before the breaker has a chance to trip.
What's the safest default fallback when a non-critical dependency fails?
Cached or default data for anything the user doesn't strictly need in real time, like a recommendation panel, so the core request still succeeds. Reserve hard failure for cases where proceeding with stale or missing data could actually mislead the user or a downstream system.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Circuit Breakers and Bulkheads, Explained With Checkout
A slow payment provider times out, threads pile up waiting, and the whole service goes unresponsive. How circuit breakers and bulkheads contain that.
Circuit Breakers and Bulkheads: Configuring Them So They Help
How to set trip thresholds against your real availability target, contain failures with bulkheads, and avoid the mistake of one setting for every call.
Circuit Breakers and Bulkheads: Stopping One Bad Dependency
How circuit breakers and bulkheads stop one slow dependency from cascading into a full outage, plus the thresholds and pitfalls that make them work.
When a Circuit Breaker Helps, and When It Just Hides a Bug
A decision guide for using circuit breakers and bulkheads to stop one failing dependency from cascading, and where a circuit breaker can mask a real problem.
Circuit Breakers and Bulkheads: Stopping a Failure From Spreading
How circuit breakers and bulkhead isolation stop one failing dependency from taking down a whole service, and the tuning mistakes that hurt.
Deciding Where a Circuit Breaker Actually Belongs in Your Pipeline
A decision guide for where circuit breakers and bulkhead isolation genuinely prevent cascading failure, and where they just add complexity without benefit.