Retry, Circuit Break, or Dead-Letter: Handling a Failing Consumer
Handle a failing consumer with three layers in order: retry with backoff for transient failures, a circuit breaker when a dependency is consistently down, and a dead-letter queue for messages that still won't process. Most production pipelines need all three rather than picking one and hoping it covers every failure mode.
Here's how each one actually behaves, and how they fit together instead of competing with each other.
When should a failing consumer retry a message?
A brief network blip or a downstream service's momentary slowdown is exactly what retries are for: wait a short, increasing interval and try again, and most transient failures resolve within a few attempts. Without a cap, a retry loop against a genuinely down dependency just hammers it harder while burning through your consumer's capacity on messages that were never going to succeed this cycle.
Use exponential backoff with a maximum retry count, and make sure retries are visible in your monitoring; a consumer quietly retrying the same message hundreds of times looks like healthy processing on a naive lag dashboard while actually making no real progress.
Circuit breakers: stop hammering a dependency that's actually down
A circuit breaker trips after enough consecutive failures against a specific dependency and stops sending new requests to it for a cooldown period, checking back periodically rather than continuously. This protects a struggling downstream service from a consumer that would otherwise keep retrying against it at full volume, and it protects the consumer from burning all its capacity on calls that are very likely to fail anyway.
Set the failure threshold and cooldown based on the dependency's actual recovery pattern, not a generic default. A dependency that typically recovers within thirty seconds needs a much shorter cooldown than one that tends to stay down for several minutes.
What should happen to a message that won't process?
Once retries are exhausted, a message goes to a dead-letter queue instead of being dropped silently or blocking the whole partition behind it. This keeps the rest of the topic moving while preserving the failed message for investigation, which matters both for debugging and for reprocessing once the underlying issue is fixed.
A dead-letter queue that's never monitored is just a place where data quietly disappears with extra steps. Alert on its growth rate, and review what's actually landing there on a regular cadence, not only when someone happens to notice it's grown large.
Combining all three without one masking the others
In practice: retry a few times with backoff for transient failures, let a circuit breaker trip if a dependency is consistently failing so you're not retrying against a wall, and send anything that still fails to a dead-letter queue for later handling. The failure mode to watch for is a circuit breaker set so aggressively that everything ends up dead-lettered within seconds of a minor blip, which erases the distinction between a genuine outage and a normal transient hiccup.
Apply the three layers in this order:
- Retry a few times with exponential backoff and a maximum count for transient failures, and make retries visible in monitoring.
- Let a circuit breaker trip after consecutive failures against a dependency, then check back after a cooldown.
- Send anything that still fails to a dead-letter queue instead of dropping it or blocking the partition.
- Monitor the dead-letter queue and reprocess its messages once the underlying issue is fixed.
- After an incident, note which layer caught the failure and tune settings to match real failure patterns.
Tie your recovery strategy to an actual error budget
At 99.9 percent availability a system has roughly 8.76 hours of allowed downtime a year, and at 99.99 percent that shrinks to about 52.6 minutes1; your retry and circuit breaker settings are effectively spending or saving that budget every time a dependency has a bad few minutes. A cooldown tuned too long burns budget waiting to retry something that recovered already; one tuned too short burns it hammering something that hasn't.
Review incidents by which layer actually caught the failure
After an incident, ask specifically which of the three layers handled it: did retries quietly absorb a brief blip, did the circuit breaker trip and protect the dependency, or did messages end up dead-lettered because all three layers ran out of options. That answer tells you whether your current tuning matches the failure patterns you're actually seeing, rather than the ones you assumed when you first configured retries and circuit breakers.
A pattern of dead-lettered messages from the same dependency, incident after incident, usually means the circuit breaker's threshold or cooldown needs adjusting, not that the dead-letter queue is doing anything wrong. Fix the layer that's actually undertuned instead of just clearing the queue and moving on.
What Good Looks Like
Error recovery is solid when retries, circuit breakers, and dead-letter handling are each tuned to actual dependency behavior and combined so no single layer masks a real outage from the others.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many times should a consumer retry before giving up?
There's no universal number; it depends on how quickly your specific dependencies typically recover from transient failures. Three to five attempts with exponential backoff is a reasonable starting point for most cases, but tune it against your dependency's actual recovery pattern rather than treating it as fixed.
Do we need a circuit breaker if we already have retries with backoff?
Yes, for any dependency that can have extended outages. Retries alone still send a steady stream of doomed requests during a real outage, just spaced out. A circuit breaker stops that traffic entirely during the outage and checks back periodically, which protects both the dependency and your own consumer capacity.
What should happen to messages sitting in a dead-letter queue?
They need a defined review process, not just storage. Once the underlying issue is fixed, reprocess them deliberately rather than leaving them indefinitely, and if a specific message keeps failing even after reprocessing, that's usually a data or logic bug worth its own investigation rather than another retry.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Designing Retry and Fallback Logic That Doesn't Make Outages Worse
How to design retries, backoff, and circuit breakers so error recovery logic protects a struggling dependency instead of overwhelming it further.
Where Retry Logic Quietly Drains Your Infrastructure Budget
How poorly designed error handling and retry logic turns a small outage into a large bill, and the specific patterns worth fixing first.
The Retry Logic That Makes Outages Worse, Not Better
How naive retry and fallback logic can amplify an outage instead of recovering from it, and the specific patterns that actually help.
Designing Retry Logic That Doesn't Make Things Worse
How to decide when a retry actually helps, why fixed intervals cause outages, and how to design fallback logic that degrades instead of breaking.
Designing Retry and Fallback Logic That Doesn't Undermine Your Access Controls
A checklist for building retry, idempotency, and fallback logic for zero trust APIs, plus the specific pitfalls that quietly weaken access control.
What Should Happen When Your Vector Search Call Fails
Cached results, keyword fallback, or an honest error message: decide a RAG pipeline's failure behavior in advance, per feature, not during the outage.