The Message Queue Decision That Determines Your Failure Modes
Moving from direct service calls to event-driven messaging is usually framed as a scaling decision, and it is one, but the choice that matters most happens earlier: queue or stream, and what each guarantees about ordering, delivery, and replay. Get that choice wrong for your actual use case and you inherit the wrong failure modes for years, discovered one production incident at a time.
The two aren't interchangeable, even though they're often discussed as if picking between them is a minor implementation detail.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What is the difference between a queue and a stream?
A queue typically guarantees a message gets processed roughly once, by roughly one consumer, then it's gone. A stream keeps events around and lets multiple independent consumers each read through them at their own pace, replaying from any point if needed. Picking based on familiarity rather than which guarantee your actual use case needs is the root cause of a lot of downstream pain: a queue used where replay is needed, or a stream used where exactly-once processing matters, both create real problems later.
Do you need idempotent consumers with queues and streams?
Both queues and streams can, and eventually will, deliver a message more than once, whether from a retry after a timeout, a consumer crash before acknowledging, or a rebalance. Design every consumer to handle duplicate delivery safely from day one, not as a fix applied after the first incident where a duplicate charge or a duplicate email reveals the gap. A consumer that isn't idempotent is a production incident waiting for the right retry to trigger it. A simple, reliable pattern: generate a unique key for each logical operation, store it alongside the result once processed, and check that store before doing the work again on redelivery. It's a small amount of extra code per consumer, and it's far cheaper to write once upfront than to retrofit after a duplicate has already reached a customer.
Make a consumer safe against duplicate delivery like this:
- Generate a unique key for each logical operation when it is first created.
- Store that key alongside the result once the operation has been processed.
- Check the store for the key before doing any work whenever a message arrives again.
- Skip the work and acknowledge the message if the key is already there.
Dead letter handling determines whether failures are visible
A message that a consumer can't process, malformed data, a downstream dependency that's down, needs somewhere to go besides silently disappearing or blocking every message behind it in the same queue. A dead letter queue with actual alerting on it, not just a place messages quietly pile up, is what turns an invisible failure into one someone notices and fixes. Test the dead letter path deliberately: send a message you know will fail and confirm it surfaces somewhere a human will see it.
Ordering guarantees are easy to assume and easy to break
Some systems guarantee order within a single partition or key but not globally, which is fine until a consumer's logic quietly assumes global order and breaks the first time two related events land on different partitions. Document explicitly what ordering your system actually guarantees, at whatever granularity it guarantees it, and audit consumer logic against that documented guarantee rather than against what engineers assumed when they wrote it.
For example, if two events about the same order land on different partitions, a consumer may see the cancellation before the creation. A useful habit is to write down, next to each consumer, the ordering assumption its logic makes and the guarantee the system actually provides at that granularity. When the two don't match, either key the events so related ones share a partition or make the consumer tolerate out of order arrival. Catching that mismatch in a design review is far cheaper than finding it after a corrupted state reaches production.
Tie your recovery target to what the business can tolerate
How much message backlog or reprocessing time is acceptable during an outage depends on the actual cost of delay. A 99.95% uptime target leaves roughly 4.38 hours a year of downtime budget across every incident combined1, and a messaging outage that takes an hour to detect and another two to drain the backlog can consume most of that budget in one event. Write the acceptable backlog drain time down as an explicit target, not something discovered during the first real outage.
A worked example: the duplicate that wasn't caught until production
Say a payment confirmation consumer wasn't built to handle duplicate delivery, on the assumption the queue guaranteed exactly-once processing. During a network blip, a handful of messages got redelivered after an ack was lost in transit, and a small number of customers were charged twice. The fix, an idempotency key checked before processing, took an afternoon. Finding it took a support escalation and a manual refund process first, which is the more common order these bugs actually get discovered in.
What Good Looks Like
Good event-driven architecture means a queue or stream chosen against the actual guarantee needed, idempotent consumers by default, a tested and alerted dead letter path, and an explicit, documented backlog recovery target.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
tracking which consumers have been audited for idempotency in a tool like ClickUp turns a one-time review into a checklist that survives team turnover
Vanta can pull evidence of your incident response and monitoring controls once dead letter alerting and backlog tracking are in place
Frequently Asked Questions
How do we decide between a queue and a stream?
Start from what your use case actually needs: exactly-once processing by a single consumer points toward a queue, multiple independent consumers each needing their own pace and replay ability points toward a stream. Picking based on which one the team already has experience with, without checking the guarantee against the actual requirement, is the most common source of later problems.
Is idempotency really necessary if our queue promises exactly-once delivery?
Yes. Most exactly-once guarantees cover delivery mechanics, not the full path through your consumer logic, and retries, crashes, and rebalances can still produce duplicate processing in practice. Building consumers to handle duplicates safely is cheap insurance against a guarantee that's narrower than it sounds.
What's the fastest way to test our dead letter handling?
Deliberately send a message you know will fail processing and confirm where it ends up and whether anyone gets alerted. If the answer is 'it disappears' or 'it sits somewhere nobody's watching,' that's the gap to close before a real failure relies on that same path.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Moving From Direct API Calls to an Event Queue Without Losing Messages
How to move one workflow from direct service calls to an event queue, covering delivery guarantees, dead letter queues, and idempotent consumers.
What Happens When a Message Queue Backs Up, Walked Through Start to Finish
A walkthrough of a message queue backlog building up in production, what caused it, and the specific changes that would have caught it sooner.
Build or Buy for Verifying Every Device That Connects?
How to split device identity from device posture checking, what building either one in house actually costs, and where a platform earns its keep instead.
Decoupling Services With Events Without Losing Traceability
A worked example of decoupling two services with an event queue, and the specific traceability and ordering problems that show up once you do.
How to Ship a Risky Change Without a 2am Rollback
A concrete walkthrough of how to plan a risky production deployment: how to split it, what to watch, and when to decide the rollback trigger.
Event-Driven Architecture: The Questions to Answer Before You Adopt It
Message queues decouple services but trade synchronous simplicity for new failure modes. Here are the questions worth answering before you commit.