What Happens When a Message Queue Backs Up, Walked Through Start to Finish
A message queue backs up when its consumer becomes slightly slower than its producer, often because a downstream dependency slows without failing. By the time a customer notices, the backlog can take hours to clear. This walkthrough shows how it happens end to end and which signals would have caught it sooner.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
The setup: a consumer that's slightly slower than the producer, most of the time
A service publishes events at a roughly steady rate. A downstream consumer processes them, calling an external API as part of each event's handling. Most of the time, the consumer keeps up: its processing rate is close enough to the publish rate that the queue depth stays low and nobody's watching it closely, because it's never been a problem.
Where it breaks: a dependency slows down, not fails
The external API the consumer calls doesn't go down; it gets slower, adding a few hundred milliseconds of latency per call during a period of its own load. That's not enough to trigger a timeout or an error, so nothing alerts. But it's enough to push the consumer's processing rate below the publish rate, and the queue starts accumulating a backlog that grows for as long as the slowdown lasts.
This is the part that makes queue backlogs sneaky: no single request fails, no error rate spikes, and the system looks healthy on every dashboard that watches for failures rather than watching for a growing gap between produce and consume rates.
What actually surfaced the problem
A customer noticed a delay between an action they took and the resulting notification arriving, and reported it as a bug. Tracing it back led to the queue, where the depth had grown for several hours before anyone looked at it, because nothing was actively monitoring queue depth or consumer lag as a first-class signal, only error rates and basic uptime.
The specific fixes that would have caught this sooner
Monitor consumer lag directly, meaning the gap between the newest published event and the oldest unprocessed one, and alert on it independent of error rate, since this failure mode produces no errors at all. Add a timeout and circuit breaker around the external API call, so a slow dependency degrades the consumer's throughput in a bounded, predictable way instead of an unbounded one. Autoscale consumer capacity based on queue depth rather than a fixed count, so a temporary slowdown gets absorbed by added capacity instead of accumulating. Set an explicit alert threshold on backlog age, not just depth, since a backlog of many small, fast-to-process events is a different problem from the same depth made up of larger, slower ones.
The fixes that would have caught the backlog sooner:
- Monitor consumer lag directly, meaning the gap between the newest published event and the oldest unprocessed one, and alert on it independently of error rate.
- Add a timeout and circuit breaker around the external API call, so a slow dependency degrades throughput in a bounded, predictable way.
- Autoscale consumer capacity on queue depth rather than a fixed count, so a temporary slowdown is absorbed by added capacity.
- Alert on backlog age as well as depth, since many small fast events and fewer slow ones are different problems at the same depth.
The harder fix: deciding what happens to a message that's too old to matter
Once the backlog cleared, the team faced a design question the incident exposed: some of the delayed notifications were no longer useful by the time they'd finally process, since the action they referred to was already stale. This is a decision worth making deliberately, not during the next incident: define a maximum useful age per event type, and either drop or specially handle events that exceed it, rather than processing every backlogged message as if it arrived on time.
For example, a queue that carries password reset emails has a short useful life: once the reset link has expired, sending the message only confuses the customer. A queue that carries audit records has the opposite profile, because a record delivered hours late is still valuable. Deciding this per event type means writing down a maximum useful age, then choosing what happens past it: drop the event, send it to a review queue, or process it with a note that it is late. Make that choice during design review, when it costs a conversation, not during the next incident.
Why nobody caught this in code review
The consumer code itself was correct. Every individual line did what it was supposed to do, including the call to the external dependency that eventually slowed down. What was missing wasn't a bug to catch in review, it was a monitoring gap: nobody had asked, at design time, what happens to this system's health signal if the consumer gets slower without ever actually failing.
That's a useful question to add to design review for any new consumer built on a queue: not just what happens if a call fails, but what happens if a call simply gets slower, and whether anything would notice before a customer does.
What Good Looks Like
Good event-driven architecture practice means consumer lag is monitored directly, slow dependencies are bounded with timeouts, consumer capacity scales with backlog, and stale events have a defined handling policy.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Industry-leading platform for Enterprise DevSecOps: Event-Driven Message Queuing.
A queue backlog with no obvious cause is also worth ruling out as a security event, not just a performance one; CrowdStrike's telemetry on the consumer host is a quick way to confirm the slowdown wasn't something worse than a slow dependency.
Frequently Asked Questions
Should every consumer have autoscaling based on queue depth?
For any consumer where processing rate can plausibly fall behind publish rate under real conditions, yes. For a consumer with generous, consistent headroom and low consequence if it lags briefly, a fixed capacity with good monitoring is a reasonable, simpler choice.
How do we choose a reasonable consumer lag alert threshold?
Base it on how quickly a delay becomes a real problem for whatever the events represent. A notification queue might tolerate a short lag before it matters to users; a payment processing queue likely tolerates almost none. Set the threshold from that impact, not from an arbitrary round number.
Is a dead letter queue enough to handle events that fail to process?
It handles events that error out, but not the scenario in this walkthrough, where events succeed but arrive too late to be useful. That's a separate design decision about staleness, worth handling explicitly rather than assuming a dead letter queue covers every failure mode.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Moving From Direct API Calls to an Event Queue Without Losing Messages
How to move one workflow from direct service calls to an event queue, covering delivery guarantees, dead letter queues, and idempotent consumers.
Decoupling Services With Events Without Losing Traceability
A worked example of decoupling two services with an event queue, and the specific traceability and ordering problems that show up once you do.
Event-Driven Architecture: The Questions to Answer Before You Adopt It
Message queues decouple services but trade synchronous simplicity for new failure modes. Here are the questions worth answering before you commit.
When Event-Driven Messaging Is Worth the Complexity
Where event-driven messaging genuinely earns its added complexity over direct calls, the debugging cost it adds, and a middle path that avoids both extremes.
The Message Queue Decision That Determines Your Failure Modes
Choosing between a queue and a stream for event-driven messaging sets your failure modes for years. What each actually guarantees, and where each breaks.
Moving From Request-Response to Event-Driven Without a Rewrite
A worked example of introducing event-driven messaging into an existing request-response API one workflow at a time, without a full rewrite.